Command Palette

Search for a command to run...

PHASE 10Intermediate Java 8+ ~34 min· topic 6 of 8

Topic 10.6

Collectors

In one line

A collector is a recipe that tells stream.collect(...) how to build a result: a list, a set, a joined string, a map, or a map of groups with counts and totals. Collectors provides ready-made ones like toMap, groupingBy, partitioningBy, joining and teeing, and they can be nested to answer questions like "total sales per city" in one pipeline.

Think of it like this

A post office sorting room. Letters arrive one at a time. The clerk has a rack of pigeonholes, drops each letter into the hole for its pin code, and at the end of the day counts how many are in each hole. A collector is the clerk's instructions: get an empty rack, put one letter in, (if two clerks worked in parallel) merge two racks, and finally turn the rack into the report.

Words you'll meet

New words in this topic, in plain English. Come back here whenever one feels fuzzy.

Collector
A recipe for building a result from stream elements: how to make an empty container, add an element, merge two containers and finish.
Mutable reduction
Combining elements by adding them into a changeable container (like a list or map), rather than creating a new value at every step.
Classifier
The function that decides which group an element belongs to in groupingBy, such as Sale::city.
Downstream collector
A collector applied to the elements inside each group, such as counting() in groupingBy(Sale::city, counting()).
Merge function
A function that decides the value to keep when two elements produce the same map key, like (a, b) -> a + b.
Map factory
A supplier such as TreeMap::new that chooses which kind of map a collector creates.
Finisher
The last step of a collector, which turns the working container into the final result (for example a StringBuilder into a String).
Partition
A split into exactly two groups: the elements that match a condition and those that don't.

Step by step

01collect versus reduce

reduce combines values into a new value at each step: 0 + 120, then 120 + 300. That's perfect for numbers, but for building a list it would mean copying the list at every step.

collect creates one container per thread and mutates it: list.add(x). It's still safe in parallel, because each thread gets its own container from the supplier and the combiner merges them at the end. That's why you should never forEach into a shared list: collect already does the job correctly.

02The four parts of a collector

Collectors.joining(", ") is roughly: supplier () -> new StringJoiner(", "), accumulator StringJoiner::add, combiner StringJoiner::merge, finisher StringJoiner::toString. Every collector, however fancy, is these four functions.

The characteristics are hints: IDENTITY_FINISH (the container is the result, skip the finisher), UNORDERED (order doesn't matter, as for toSet), and CONCURRENT (one shared container can be updated by many threads, as for groupingByConcurrent).

The four parts of a collectordiagram
Rendering diagram…

03Lists, sets and strings

toList() and toSet() give a list or set without promising the implementation; toCollection(TreeSet::new) lets you choose. toUnmodifiableList() (Java 10) returns a list that rejects changes and rejects null elements.

joining() concatenates strings; joining(", ") adds a separator between elements (not after the last); joining(", ", "[", "]") adds a prefix and suffix. Only CharSequence elements can be joined, so map(String::valueOf) first if you have numbers.

04toMap and its two traps

toMap(Product::sku, Product::price) maps each element to one entry. If two elements produce the same key, there's no sensible default (keep the first? the last? add them?), so it throws IllegalStateException with both values in the message. Give it a merge function as the third argument to choose.

The second trap is null values. toMap uses Map.merge-style insertion, which doesn't allow null values, so a null from your value function throws NullPointerException even though HashMap itself accepts nulls. Filter them out, map them to a default, or collect with a plain loop.

Main.javawhole filejava
Map<String, Integer> priceBySku = products.stream()
        .collect(Collectors.toMap(
                Product::sku,          // key
                Product::price,        // value
                (a, b) -> a,           // on duplicate key: keep the first
                TreeMap::new));        // map type: sorted by key

05groupingBy: buckets by key

groupingBy(Sale::city) asks the classifier for each element's key, finds (or creates) the bucket for that key, and adds the element to it. The default is HashMap<K, List<T>>.

The three-argument form groupingBy(classifier, mapFactory, downstream) lets you choose the map type and what each bucket should become. Instead of a list per city, you can have a count, a sum, a set of product names, or the biggest sale.

groupingBy: buckets by keydiagram
Rendering diagram…

06Downstream collectors, and filtering versus filter

Any collector can be a downstream: counting(), summingInt, averagingInt, mapping(f, collector), filtering(p, collector) and flatMapping(f, collector) (both Java 9), maxBy/minBy (they return Optional), toSet(), joining, even another groupingBy for two-level grouping.

filter(...) before groupingBy removes elements entirely, so a city with no big sales disappears from the map. filtering(...) inside groupingBy keeps every city as a key and filters within each group, so that city appears with 0 or an empty list. Choose based on whether "no matches" should be visible.

Main.javawhole filejava
// cities with at least one big sale only
sales.stream().filter(s -> s.amount() >= 200)
        .collect(groupingBy(Sale::city, TreeMap::new, counting()));          // {Delhi=1, Pune=2}

// every city, with a count of its big sales
sales.stream()
        .collect(groupingBy(Sale::city, TreeMap::new,
                filtering(s -> s.amount() >= 200, counting())));             // {Delhi=1, Mumbai=0, Pune=2}

07partitioningBy, collectingAndThen and teeing

partitioningBy(s -> s.marks() >= 40) returns Map<Boolean, List<Student>> with both true and false keys, always. It also takes a downstream: partitioningBy(pred, counting()).

collectingAndThen(maxBy(...), opt -> opt.map(Student::name).orElse("nobody")) runs a final function on another collector's result. teeing(minBy(...), maxBy(...), (min, max) -> ...) (Java 12) sends each element to both collectors and merges the two results, so you get two answers from one pass over the data.

Try it yourself

  1. 1

    Swap the map factory

    In the groupingBy example, change one TreeMap::new to LinkedHashMap::new (and import it). Predict the city order now: it follows the first time each city appears in the list.

  2. 2

    filter versus filtering

    In the same example, add a version that uses .filter(s -> s.amount() >= 200) before groupingBy(..., counting()). Predict whether Mumbai appears, then compare with the filtering line.

  3. 3

    Two-level grouping

    Group the sales by city, then by product inside each city, summing amounts: groupingBy(Sale::city, TreeMap::new, groupingBy(Sale::product, TreeMap::new, summingInt(Sale::amount))). Predict the printed map for Pune before you run it.

Code & diagrams

Lists, sets, strings and numbers Java 10+ New tab
Sign in to run this example in your browser.

Expected output

toList:            [mango, apple, kiwi, apple, banana]
toCollection:      [apple, banana, kiwi, mango]
joining:           mango, apple, kiwi, banana
joining + affixes: [mango | apple | kiwi | banana]
unmodifiable:      [MANGO, APPLE, KIWI, APPLE, BANANA]
counting:          5
averaging length:  5.0
summarizing:       IntSummaryStatistics{count=5, sum=25, min=4, average=5.000000, max=6}
toMap: merge functions, map types and the duplicate-key error Java 16+ New tab
Sign in to run this example in your browser.

Expected output

TreeMap by sku:      {P1=250, P2=900, P3=20}
LinkedHashMap keys:  [P3, P1, P2]
duplicate keys:      Duplicate key a (attempted merging values apple and avocado)
merged on collision: {a=apple+avocado, b=banana+blueberry, c=cherry}
counted by letter:   {a=2, b=2, c=1}
groupingBy with downstream collectors Java 16+ New tab

TreeMap::new keeps the cities sorted. With the default HashMap, the key order isn't something to rely on.

Sign in to run this example in your browser.

Expected output

Delhi -> 2
Mumbai -> 1
Pune -> 3
count:     {Delhi=2, Mumbai=1, Pune=3}
total:     {Delhi=360, Mumbai=90, Pune=770}
mapping:   {Delhi=[coffee, tea], Mumbai=[tea], Pune=[tea, coffee, cake]}
biggest in Delhi: coffee
biggest in Mumbai: tea
biggest in Pune: cake
filtering: {Delhi=1, Mumbai=0, Pune=2}
partitioningBy, collectingAndThen and teeing Java 16+ New tab

Collectors.teeing arrived in Java 12. It reads the data once and feeds both collectors.

Sign in to run this example in your browser.

Expected output

passed: [Asha, Meera], failed: [Ravi, Kabir]
empty input still has both keys: {false=0, true=0}
topper: Asha
lowest 28, highest 92
average: 55.5
Writing your own collector Java 9+ New tab

In parallel, each thread builds its own StringBuilder and the combiner joins them in encounter order, so the answer is the same.

Sign in to run this example in your browser.

Expected output

ARM
parallel too: JVM
[2, 4, 6]
toList characteristics: [IDENTITY_FINISH]
toSet characteristics:  [UNORDERED, IDENTITY_FINISH]
initials characteristics: []

Break it on purpose

Errors are the best teachers. Make each change, read the error, guess what went wrong, then reveal the answer.

Break #1

Duplicate keys in toMap

Collect List.of("apple", "avocado", "banana") with Collectors.toMap(w -> w.charAt(0), w -> w).

terminal
$ java Main.java
── what you'll see ──
Exception in thread "main" java.lang.IllegalStateException: Duplicate key a (attempted merging values apple and avocado)
at java.base/java.util.stream.Collectors.duplicateKeyException(Collectors.java:135)
at java.base/java.util.stream.Collectors.lambda$uniqKeysMapAccumulator$1(Collectors.java:182)
at java.base/java.util.stream.ReduceOps$3ReducingSink.accept(ReduceOps.java:169)
at java.base/java.util.AbstractList$RandomAccessSpliterator.forEachRemaining(AbstractList.java:722)
at java.base/java.util.stream.AbstractPipeline.copyInto(AbstractPipeline.java:509)
at java.base/java.util.stream.AbstractPipeline.wrapAndCopyInto(AbstractPipeline.java:499)
at java.base/java.util.stream.ReduceOps$ReduceOp.evaluateSequential(ReduceOps.java:921)
at java.base/java.util.stream.AbstractPipeline.evaluate(AbstractPipeline.java:234)
at java.base/java.util.stream.ReferencePipeline.collect(ReferencePipeline.java:682)
at Main.main(Main.java:9)

Break #2

A null value in toMap

Collect users to toMap(User::name, User::email) where one user has a null email.

terminal
$ java Main.java
── what you'll see ──
Exception in thread "main" java.lang.NullPointerException
at java.base/java.util.Objects.requireNonNull(Objects.java:233)
at java.base/java.util.stream.Collectors.lambda$uniqKeysMapAccumulator$1(Collectors.java:180)
at java.base/java.util.stream.ReduceOps$3ReducingSink.accept(ReduceOps.java:169)
at java.base/java.util.AbstractList$RandomAccessSpliterator.forEachRemaining(AbstractList.java:722)
at java.base/java.util.stream.AbstractPipeline.copyInto(AbstractPipeline.java:509)
at java.base/java.util.stream.AbstractPipeline.wrapAndCopyInto(AbstractPipeline.java:499)
at java.base/java.util.stream.ReduceOps$ReduceOp.evaluateSequential(ReduceOps.java:921)
at java.base/java.util.stream.AbstractPipeline.evaluate(AbstractPipeline.java:234)
at java.base/java.util.stream.ReferencePipeline.collect(ReferencePipeline.java:682)
at Main.main(Main.java:11)

Break #3

A null key in groupingBy

Group users by User::city when one user's city is null.

terminal
$ java Main.java
── what you'll see ──
Exception in thread "main" java.lang.NullPointerException: element cannot be mapped to a null key
at java.base/java.util.Objects.requireNonNull(Objects.java:259)
at java.base/java.util.stream.Collectors.lambda$groupingBy$53(Collectors.java:1105)
at java.base/java.util.stream.ReduceOps$3ReducingSink.accept(ReduceOps.java:169)
at java.base/java.util.AbstractList$RandomAccessSpliterator.forEachRemaining(AbstractList.java:722)
at java.base/java.util.stream.AbstractPipeline.copyInto(AbstractPipeline.java:509)
at java.base/java.util.stream.AbstractPipeline.wrapAndCopyInto(AbstractPipeline.java:499)
at java.base/java.util.stream.ReduceOps$ReduceOp.evaluateSequential(ReduceOps.java:921)
at java.base/java.util.stream.AbstractPipeline.evaluate(AbstractPipeline.java:234)
at java.base/java.util.stream.ReferencePipeline.collect(ReferencePipeline.java:682)
at Main.main(Main.java:10)

Myth vs fact

Myth

groupingBy returns the groups in sorted order.

Fact

By default it returns a HashMap, whose order depends on hash codes and capacity. Pass TreeMap::new for sorted keys or LinkedHashMap::new for first-seen order.

Myth

toMap keeps the last value for duplicate keys, like Map.put.

Fact

The two-argument toMap throws IllegalStateException on a duplicate key. You must supply a merge function to choose.

Myth

Collecting into a list from a parallel stream needs a synchronized list.

Fact

No: collect gives each thread its own container and merges them with the combiner, preserving encounter order for ordered streams. Only manual forEach(list::add) needs (and still shouldn't use) synchronization.

Interview problem

The problem

Top K frequent words

Given a list of words from search queries, return the 3 most frequent words. Ties are broken alphabetically. Write it with streams, then discuss its cost.

You're given

  • Input: a List<String> that may have millions of words.
  • Output: a List<String> of length at most 3, most frequent first.
  • Equal counts are ordered alphabetically.

The interviewer follows up

01

Why does .reversed() need the explicit type witness Map.Entry.<String, Long>comparingByValue()?

02

Can you make the counting step parallel?

Pro corner

Extra depth for experienced readers. New to this? Skip it for now and come back later.

  • ▸

    Collector<T, A, R> has three type parameters: element T, mutable accumulation type A (often hidden as ? in signatures), and result R. ReduceOps.makeRef(collector) drives it: one container per leaf task, accumulator per element, combiner up the fork/join tree, and finisher at the top (skipped when IDENTITY_FINISH is set).

  • ▸

    With a CONCURRENT and UNORDERED collector (groupingByConcurrent, toConcurrentMap) on a parallel stream (or one whose source is unordered), the framework uses a single shared ConcurrentHashMap instead of merging per-thread maps. That avoids expensive map merges but gives up encounter order inside groups.

  • ▸

    Merging maps is the real cost of parallel groupingBy: each leaf builds its own HashMap of lists, then the combiner merges maps key by key and concatenates lists. For high-cardinality keys this can make the parallel version slower than the sequential one (Topic 10.8).

  • ▸

    toList() on Stream (Java 16) isn't a collector at all: it's a terminal operation that can build the list directly from the internal array (SharedSecrets access to an unmodifiable list wrapper), so it's slightly cheaper than collect(Collectors.toUnmodifiableList()) and, unlike it, allows null elements.

Remember this

  1. 1

    collect is a mutable reduction: instead of combining immutable values like reduce does (Topic 10.5), it creates a container (a List, a Map, a StringBuilder) and adds elements to it. Every collector is four functions: a supplier (new empty container), an accumulator (add one element), a combiner (merge two containers, used in parallel) and a finisher (turn the container into the final result), plus characteristics that describe it.

  2. 2

    Simple collectors: toList(), toSet(), toCollection(TreeSet::new) for a specific collection type, the unmodifiable toUnmodifiableList/Set/Map (Java 10), joining(", ") (with optional prefix and suffix) for strings, and the numeric counting(), summingInt, averagingDouble and summarizingInt.

  3. 3

    toMap(keyFn, valueFn) builds a map, but it is strict: two elements with the same key throw IllegalStateException: Duplicate key ..., and a **null value** throws NullPointerException. Pass a merge function as the third argument to decide what happens on a collision ((a, b) -> a keeps the first, Integer::sum adds), and a map factory as the fourth (TreeMap::new, LinkedHashMap::new) when you care about order.

  4. 4

    groupingBy(classifier) puts each element in a bucket by key, returning Map<K, List<T>>. Its power is the downstream collector: groupingBy(Sale::city, counting()) gives counts per city, summingInt(Sale::amount) totals, mapping(Sale::product, toList()) transformed values, maxBy(...) the largest, and filtering/flatMapping (Java 9) filter or flatten inside each group. Nest them to any depth.

  5. 5

    partitioningBy(predicate) is grouping with exactly two keys, true and false; both keys are always present, even when a side is empty. collectingAndThen(collector, finisher) post-processes a result (for example make it unmodifiable or unwrap an Optional), and teeing(c1, c2, merger) (Java 12) feeds every element to two collectors at once and merges their results, like computing the minimum and maximum in one pass.

  6. 6

    Order and types: groupingBy and toMap return a HashMap by default, whose iteration order isn't sorted or stable to rely on, so pass TreeMap::new or LinkedHashMap::new when order matters (for output, tests, or APIs). groupingBy rejects null keys with element cannot be mapped to a null key.

Explain it without notes

01

What are the components of a Collector, and what does each do?

02

How does toMap behave with duplicate keys and null values, and how do you control it?

03

Explain groupingBy with a downstream collector, using an example.

04

What is the difference between filtering before groupingBy and using Collectors.filtering as a downstream?

05

When would you use partitioningBy, collectingAndThen and teeing?

Practice

01

Given words "apple", "bat", "cherry", "dog", "egg", group them by length into a TreeMap and print it.

02

Count how many times each character appears in "banana" and print the result as a TreeMap<Character, Long>.

03

From a list of (name, dept, salary) employees, print the highest-paid employee's name per department, using groupingBy with collectingAndThen and maxBy.

Trade-offs

  • ↔

    Nested collectors answer complex questions in one pass and one expression, but deeply nested groupingBy(..., mapping(..., collectingAndThen(...))) becomes hard to read. Extract named collectors into variables or methods, or use a small loop when it's clearer.

  • ↔

    toMap failing on duplicates is safer than silently overwriting, but it means a data change in production (a new duplicate) can crash a job. Decide the merge rule explicitly for any key that isn't guaranteed unique.

  • ↔

    Choosing TreeMap or LinkedHashMap gives predictable order for output and tests at a small cost (TreeMap is O(log n) per insert). HashMap is fastest when order doesn't matter.

Done when you can

  • Done when you can describe the four functions of a collector and write one with Collector.of.

  • Done when you can use toMap with a merge function and a map factory, and explain its null and duplicate rules.

  • Done when you can write groupingBy with counting, summingInt, mapping, maxBy and filtering downstreams.

  • Done when you can use partitioningBy, collectingAndThen and teeing appropriately.

  • Done when you always choose a map type when output order matters.