Topic 10.6
Collectors
In one line
A collector is a recipe that tells stream.collect(...) how to build a result: a list, a set, a joined string, a map, or a map of groups with counts and totals. Collectors provides ready-made ones like toMap, groupingBy, partitioningBy, joining and teeing, and they can be nested to answer questions like "total sales per city" in one pipeline.
Think of it like this
A post office sorting room. Letters arrive one at a time. The clerk has a rack of pigeonholes, drops each letter into the hole for its pin code, and at the end of the day counts how many are in each hole. A collector is the clerk's instructions: get an empty rack, put one letter in, (if two clerks worked in parallel) merge two racks, and finally turn the rack into the report.
Words you'll meet
New words in this topic, in plain English. Come back here whenever one feels fuzzy.
- Collector
- A recipe for building a result from stream elements: how to make an empty container, add an element, merge two containers and finish.
- Mutable reduction
- Combining elements by adding them into a changeable container (like a list or map), rather than creating a new value at every step.
- Classifier
- The function that decides which group an element belongs to in
groupingBy, such asSale::city. - Downstream collector
- A collector applied to the elements inside each group, such as
counting()ingroupingBy(Sale::city, counting()). - Merge function
- A function that decides the value to keep when two elements produce the same map key, like
(a, b) -> a + b. - Map factory
- A supplier such as
TreeMap::newthat chooses which kind of map a collector creates. - Finisher
- The last step of a collector, which turns the working container into the final result (for example a
StringBuilderinto aString). - Partition
- A split into exactly two groups: the elements that match a condition and those that don't.
Step by step
01collect versus reduce
reduce combines values into a new value at each step: 0 + 120, then 120 + 300. That's perfect for numbers, but for building a list it would mean copying the list at every step.
collect creates one container per thread and mutates it: list.add(x). It's still safe in parallel, because each thread gets its own container from the supplier and the combiner merges them at the end. That's why you should never forEach into a shared list: collect already does the job correctly.
02The four parts of a collector
Collectors.joining(", ") is roughly: supplier () -> new StringJoiner(", "), accumulator StringJoiner::add, combiner StringJoiner::merge, finisher StringJoiner::toString. Every collector, however fancy, is these four functions.
The characteristics are hints: IDENTITY_FINISH (the container is the result, skip the finisher), UNORDERED (order doesn't matter, as for toSet), and CONCURRENT (one shared container can be updated by many threads, as for groupingByConcurrent).
03Lists, sets and strings
toList() and toSet() give a list or set without promising the implementation; toCollection(TreeSet::new) lets you choose. toUnmodifiableList() (Java 10) returns a list that rejects changes and rejects null elements.
joining() concatenates strings; joining(", ") adds a separator between elements (not after the last); joining(", ", "[", "]") adds a prefix and suffix. Only CharSequence elements can be joined, so map(String::valueOf) first if you have numbers.
04toMap and its two traps
toMap(Product::sku, Product::price) maps each element to one entry. If two elements produce the same key, there's no sensible default (keep the first? the last? add them?), so it throws IllegalStateException with both values in the message. Give it a merge function as the third argument to choose.
The second trap is null values. toMap uses Map.merge-style insertion, which doesn't allow null values, so a null from your value function throws NullPointerException even though HashMap itself accepts nulls. Filter them out, map them to a default, or collect with a plain loop.
Map<String, Integer> priceBySku = products.stream()
.collect(Collectors.toMap(
Product::sku, // key
Product::price, // value
(a, b) -> a, // on duplicate key: keep the first
TreeMap::new)); // map type: sorted by key05groupingBy: buckets by key
groupingBy(Sale::city) asks the classifier for each element's key, finds (or creates) the bucket for that key, and adds the element to it. The default is HashMap<K, List<T>>.
The three-argument form groupingBy(classifier, mapFactory, downstream) lets you choose the map type and what each bucket should become. Instead of a list per city, you can have a count, a sum, a set of product names, or the biggest sale.
06Downstream collectors, and filtering versus filter
Any collector can be a downstream: counting(), summingInt, averagingInt, mapping(f, collector), filtering(p, collector) and flatMapping(f, collector) (both Java 9), maxBy/minBy (they return Optional), toSet(), joining, even another groupingBy for two-level grouping.
filter(...) before groupingBy removes elements entirely, so a city with no big sales disappears from the map. filtering(...) inside groupingBy keeps every city as a key and filters within each group, so that city appears with 0 or an empty list. Choose based on whether "no matches" should be visible.
// cities with at least one big sale only
sales.stream().filter(s -> s.amount() >= 200)
.collect(groupingBy(Sale::city, TreeMap::new, counting())); // {Delhi=1, Pune=2}
// every city, with a count of its big sales
sales.stream()
.collect(groupingBy(Sale::city, TreeMap::new,
filtering(s -> s.amount() >= 200, counting()))); // {Delhi=1, Mumbai=0, Pune=2}07partitioningBy, collectingAndThen and teeing
partitioningBy(s -> s.marks() >= 40) returns Map<Boolean, List<Student>> with both true and false keys, always. It also takes a downstream: partitioningBy(pred, counting()).
collectingAndThen(maxBy(...), opt -> opt.map(Student::name).orElse("nobody")) runs a final function on another collector's result. teeing(minBy(...), maxBy(...), (min, max) -> ...) (Java 12) sends each element to both collectors and merges the two results, so you get two answers from one pass over the data.
Try it yourself
- 1
Swap the map factory
In the groupingBy example, change one
TreeMap::newtoLinkedHashMap::new(and import it). Predict the city order now: it follows the first time each city appears in the list. - 2
filter versus filtering
In the same example, add a version that uses
.filter(s -> s.amount() >= 200)beforegroupingBy(..., counting()). Predict whetherMumbaiappears, then compare with thefilteringline. - 3
Two-level grouping
Group the sales by city, then by product inside each city, summing amounts:
groupingBy(Sale::city, TreeMap::new, groupingBy(Sale::product, TreeMap::new, summingInt(Sale::amount))). Predict the printed map for Pune before you run it.
Code & diagrams
Expected output
toList: [mango, apple, kiwi, apple, banana]
toCollection: [apple, banana, kiwi, mango]
joining: mango, apple, kiwi, banana
joining + affixes: [mango | apple | kiwi | banana]
unmodifiable: [MANGO, APPLE, KIWI, APPLE, BANANA]
counting: 5
averaging length: 5.0
summarizing: IntSummaryStatistics{count=5, sum=25, min=4, average=5.000000, max=6}Expected output
TreeMap by sku: {P1=250, P2=900, P3=20}
LinkedHashMap keys: [P3, P1, P2]
duplicate keys: Duplicate key a (attempted merging values apple and avocado)
merged on collision: {a=apple+avocado, b=banana+blueberry, c=cherry}
counted by letter: {a=2, b=2, c=1}TreeMap::new keeps the cities sorted. With the default HashMap, the key order isn't something to rely on.
Expected output
Delhi -> 2
Mumbai -> 1
Pune -> 3
count: {Delhi=2, Mumbai=1, Pune=3}
total: {Delhi=360, Mumbai=90, Pune=770}
mapping: {Delhi=[coffee, tea], Mumbai=[tea], Pune=[tea, coffee, cake]}
biggest in Delhi: coffee
biggest in Mumbai: tea
biggest in Pune: cake
filtering: {Delhi=1, Mumbai=0, Pune=2}Collectors.teeing arrived in Java 12. It reads the data once and feeds both collectors.
Expected output
passed: [Asha, Meera], failed: [Ravi, Kabir]
empty input still has both keys: {false=0, true=0}
topper: Asha
lowest 28, highest 92
average: 55.5In parallel, each thread builds its own StringBuilder and the combiner joins them in encounter order, so the answer is the same.
Expected output
ARM
parallel too: JVM
[2, 4, 6]
toList characteristics: [IDENTITY_FINISH]
toSet characteristics: [UNORDERED, IDENTITY_FINISH]
initials characteristics: []Break it on purpose
Errors are the best teachers. Make each change, read the error, guess what went wrong, then reveal the answer.
Break #1
Duplicate keys in toMap
Collect List.of("apple", "avocado", "banana") with Collectors.toMap(w -> w.charAt(0), w -> w).
Break #2
A null value in toMap
Collect users to toMap(User::name, User::email) where one user has a null email.
Break #3
A null key in groupingBy
Group users by User::city when one user's city is null.
Myth vs fact
Myth
groupingBy returns the groups in sorted order.
Fact
By default it returns a HashMap, whose order depends on hash codes and capacity. Pass TreeMap::new for sorted keys or LinkedHashMap::new for first-seen order.
Myth
toMap keeps the last value for duplicate keys, like Map.put.
Fact
The two-argument toMap throws IllegalStateException on a duplicate key. You must supply a merge function to choose.
Myth
Collecting into a list from a parallel stream needs a synchronized list.
Fact
No: collect gives each thread its own container and merges them with the combiner, preserving encounter order for ordered streams. Only manual forEach(list::add) needs (and still shouldn't use) synchronization.
Interview problem
The problem
Top K frequent words
Given a list of words from search queries, return the 3 most frequent words. Ties are broken alphabetically. Write it with streams, then discuss its cost.
You're given
- Input: a
List<String>that may have millions of words. - Output: a
List<String>of length at most 3, most frequent first. - Equal counts are ordered alphabetically.
The interviewer follows up
Why does .reversed() need the explicit type witness Map.Entry.<String, Long>comparingByValue()?
Can you make the counting step parallel?
Pro corner
Extra depth for experienced readers. New to this? Skip it for now and come back later.
- ▸
Collector<T, A, R>has three type parameters: elementT, mutable accumulation typeA(often hidden as?in signatures), and resultR.ReduceOps.makeRef(collector)drives it: one container per leaf task, accumulator per element, combiner up the fork/join tree, and finisher at the top (skipped whenIDENTITY_FINISHis set). - ▸
With a
CONCURRENTandUNORDEREDcollector (groupingByConcurrent,toConcurrentMap) on a parallel stream (or one whose source is unordered), the framework uses a single sharedConcurrentHashMapinstead of merging per-thread maps. That avoids expensive map merges but gives up encounter order inside groups. - ▸
Merging maps is the real cost of parallel
groupingBy: each leaf builds its ownHashMapof lists, then the combiner merges maps key by key and concatenates lists. For high-cardinality keys this can make the parallel version slower than the sequential one (Topic 10.8). - ▸
toList()onStream(Java 16) isn't a collector at all: it's a terminal operation that can build the list directly from the internal array (SharedSecretsaccess to an unmodifiable list wrapper), so it's slightly cheaper thancollect(Collectors.toUnmodifiableList())and, unlike it, allowsnullelements.
Remember this
- 1
collectis a mutable reduction: instead of combining immutable values likereducedoes (Topic 10.5), it creates a container (aList, aMap, aStringBuilder) and adds elements to it. Every collector is four functions: a supplier (new empty container), an accumulator (add one element), a combiner (merge two containers, used in parallel) and a finisher (turn the container into the final result), plus characteristics that describe it. - 2
Simple collectors:
toList(),toSet(),toCollection(TreeSet::new)for a specific collection type, the unmodifiabletoUnmodifiableList/Set/Map(Java 10),joining(", ")(with optional prefix and suffix) for strings, and the numericcounting(),summingInt,averagingDoubleandsummarizingInt. - 3
toMap(keyFn, valueFn)builds a map, but it is strict: two elements with the same key throwIllegalStateException: Duplicate key ..., and a **nullvalue** throwsNullPointerException. Pass a merge function as the third argument to decide what happens on a collision ((a, b) -> akeeps the first,Integer::sumadds), and a map factory as the fourth (TreeMap::new,LinkedHashMap::new) when you care about order. - 4
groupingBy(classifier)puts each element in a bucket by key, returningMap<K, List<T>>. Its power is the downstream collector:groupingBy(Sale::city, counting())gives counts per city,summingInt(Sale::amount)totals,mapping(Sale::product, toList())transformed values,maxBy(...)the largest, andfiltering/flatMapping(Java 9) filter or flatten inside each group. Nest them to any depth. - 5
partitioningBy(predicate)is grouping with exactly two keys,trueandfalse; both keys are always present, even when a side is empty.collectingAndThen(collector, finisher)post-processes a result (for example make it unmodifiable or unwrap anOptional), andteeing(c1, c2, merger)(Java 12) feeds every element to two collectors at once and merges their results, like computing the minimum and maximum in one pass. - 6
Order and types:
groupingByandtoMapreturn aHashMapby default, whose iteration order isn't sorted or stable to rely on, so passTreeMap::neworLinkedHashMap::newwhen order matters (for output, tests, or APIs).groupingByrejectsnullkeys withelement cannot be mapped to a null key.
Explain it without notes
What are the components of a Collector, and what does each do?
How does toMap behave with duplicate keys and null values, and how do you control it?
Explain groupingBy with a downstream collector, using an example.
What is the difference between filtering before groupingBy and using Collectors.filtering as a downstream?
When would you use partitioningBy, collectingAndThen and teeing?
Practice
Given words "apple", "bat", "cherry", "dog", "egg", group them by length into a TreeMap and print it.
Count how many times each character appears in "banana" and print the result as a TreeMap<Character, Long>.
From a list of (name, dept, salary) employees, print the highest-paid employee's name per department, using groupingBy with collectingAndThen and maxBy.
Trade-offs
- ↔
Nested collectors answer complex questions in one pass and one expression, but deeply nested
groupingBy(..., mapping(..., collectingAndThen(...)))becomes hard to read. Extract named collectors into variables or methods, or use a small loop when it's clearer. - ↔
toMapfailing on duplicates is safer than silently overwriting, but it means a data change in production (a new duplicate) can crash a job. Decide the merge rule explicitly for any key that isn't guaranteed unique. - ↔
Choosing
TreeMaporLinkedHashMapgives predictable order for output and tests at a small cost (TreeMapis O(log n) per insert).HashMapis fastest when order doesn't matter.
Done when you can
Done when you can describe the four functions of a collector and write one with
Collector.of.Done when you can use
toMapwith a merge function and a map factory, and explain its null and duplicate rules.Done when you can write
groupingBywithcounting,summingInt,mapping,maxByandfilteringdownstreams.Done when you can use
partitioningBy,collectingAndThenandteeingappropriately.Done when you always choose a map type when output order matters.