Command Palette

Search for a command to run...

PHASE 10Intermediate Java 8+ ~34 min· topic 5 of 8

Topic 10.5

Stream Operations in Depth

In one line

Stream operations fall into clear groups: stateless intermediate operations (filter, map, flatMap), stateful ones that must remember or buffer elements (distinct, sorted, limit), short-circuiting ones that can stop early (limit, anyMatch, findFirst), and terminal operations like reduce, min and toArray. Knowing which is which tells you what a pipeline costs and whether it can finish.

Think of it like this

An airport security line. The bag scanner looks at each passenger on their own and lets them through or not: it doesn't care who came before (stateless). The boarding desk that lines people up by seat row has to wait until everyone has arrived before it can let the first one through (stateful, a barrier). And the officer looking for "any passenger carrying a guitar" can stop the moment one appears (short-circuiting).

Words you'll meet

New words in this topic, in plain English. Come back here whenever one feels fuzzy.

Stateless operation
An operation that decides about each element on its own, without remembering any other element. filter and map are stateless.
Stateful operation
An operation that must remember earlier elements, like distinct (what it has seen) or sorted (everything, to sort it).
Barrier
A point in the pipeline where all elements must arrive before any can continue. sorted is a barrier.
Short-circuiting
Able to produce its result without processing every element, like findFirst stopping at the first match.
flatMap
An operation that turns each element into a stream of zero or more elements and joins all of those streams into one.
Reduction
Combining all elements into a single result, like a sum or a maximum.
Identity
A starting value that doesn't change the result when combined: 0 for addition, 1 for multiplication.
Associative
An operation where grouping doesn't matter: (a + b) + c equals a + (b + c). Subtraction isn't associative.
Encounter order
The order in which a stream's source presents its elements, such as index order for a List.

Step by step

01A map of the operations

Before memorising methods, learn the two questions to ask of any operation: does it need to remember other elements (stateful or stateless)? Can it stop early (short-circuiting)? The answers predict memory use, whether the pipeline works on infinite input, and how well it parallelises.

A map of the operationsdiagram
Rendering diagram…

02filter, map and flatMap

filter(predicate) keeps elements for which the predicate is true. map(function) replaces each element with the function's result, so a Stream<Order> mapped with Order::customer becomes a Stream<String>.

flatMap is for one-to-many: each order has a list of items, and you want one stream of all items. map(o -> o.items()) would give a stream of lists; flatMap(o -> o.items().stream()) opens each inner stream and pours its elements into the outer one. An empty inner stream simply contributes nothing.

filter, map and flatMapdiagram
Rendering diagram…

03mapMulti: flatMap without creating streams (Java 16)

flatMap creates a small stream object for every element, which is wasteful when each element produces only zero, one or two results. mapMulti((element, sink) -> ...) hands you a Consumer called the sink; you call sink.accept(x) once per output element, or not at all.

It's also handy for type filtering: .<Circle>mapMulti((shape, sink) -> { if (shape instanceof Circle c) sink.accept(c); }) filters and casts in one step. You often need the explicit type witness <Circle> because the compiler can't infer the output type from the lambda.

Main.javawhole filejava
List<String> tagged = orders.stream()
        .<String>mapMulti((o, sink) -> {
            for (String item : o.items()) {
                sink.accept(o.id() + ":" + item);      // zero or more outputs per order
            }
        })
        .toList();

04Stateful operations: distinct, sorted, limit, skip

distinct() passes an element only if it hasn't seen an equal one before. It keeps a HashSet of seen elements, so memory grows with the number of distinct values, and it relies on correct equals and hashCode.

sorted() (natural order) and sorted(comparator) can't emit anything until the last element has arrived: the smallest might be the last one. So sorted collects everything into a buffer, sorts it, then pushes elements on. On an ordered stream the sort is stable. On an infinite stream it runs until memory runs out.

skip(n) drops the first n elements and limit(n) stops after n. Together they implement paging (skip(page * size).limit(size)). limit is short-circuiting: once n elements have passed, the source stops being read.

terminal
$ java -Xmx64m Main.java
── expected output ──
Exception in thread "main" java.lang.OutOfMemoryError: Java heap space
at java.base/java.util.Arrays.copyOf(Arrays.java:3513)
at java.base/java.util.Arrays.copyOf(Arrays.java:3482)
at java.base/java.util.ArrayList.grow(ArrayList.java:237)
at java.base/java.util.ArrayList.grow(ArrayList.java:244)
at java.base/java.util.ArrayList.add(ArrayList.java:483)
at java.base/java.util.ArrayList.add(ArrayList.java:496)
at java.base/java.util.stream.SortedOps$RefSortingSink.accept(SortedOps.java:409)

05takeWhile and dropWhile (Java 9)

takeWhile(p) passes elements while p is true and stops at the first failure, even if later elements would pass again. dropWhile(p) skips elements while p is true and passes everything after the first failure.

Compare with filter, which tests every element independently. For temperatures 18, 21, 25, 31, 24, takeWhile(t -> t < 30) gives 18, 21, 25, but filter(t -> t < 30) gives 18, 21, 25, 24. takeWhile is short-circuiting, so it also works on infinite sorted streams where filter would run forever.

06Matching and finding

anyMatch(p) stops at the first true; allMatch(p) stops at the first false; noneMatch(p) stops at the first true. On an empty stream they return false, true and true: that's standard logic ("all of nothing" is vacuously true), and a common source of bugs when an empty list means "no data" rather than "all good".

findFirst() returns the first element in encounter order as an Optional (Topic 10.7). findAny() may return any element; in a sequential stream it's usually the first, but in parallel it returns whichever is found first, which is faster when you don't care.

07reduce: three forms

reduce(0, Integer::sum): start from the identity and fold each element in. Always returns a value; for an empty stream, the identity.

reduce(Integer::sum): no identity, so the first element is the starting point, and the result is an Optional that's empty for an empty stream.

reduce(0, (sum, item) -> sum + item.price(), Integer::sum): the accumulator takes a result-so-far and an element of a different type, and the combiner merges two partial results. The combiner is only used in parallel, but it must still be correct. For most real reductions, mapToInt(...).sum() or a collector (Topic 10.6) is clearer.

Main.javawhole filejava
int total        = prices.stream().reduce(0, Integer::sum);       // 0 if empty
Optional<Integer> t2 = prices.stream().reduce(Integer::sum);       // Optional.empty if empty
int total3       = cart.stream().reduce(0,
        (sum, item) -> sum + item.price(),                       // accumulator: Integer, Item
        Integer::sum);                                           // combiner: Integer, Integer
int best         = prices.stream().reduce(Integer.MIN_VALUE, Math::max);

08peek is for debugging, and side effects aren't guaranteed

peek(action) runs an action as each element passes, without changing it. It's handy to see what flows through a pipeline while debugging. Don't put real logic in it.

The library is allowed to skip work whose result can't change the answer. Since Java 9, list.stream().peek(...).count() returns the list's size without running peek at all, because nothing in the pipeline can change the count. Add a filter and the pipeline runs again. Code that counted events or updated metrics inside peek silently stopped working when teams upgraded from Java 8.

Try it yourself

  1. 1

    Move the barrier

    In the barrier example, move .sorted() to after .map(n -> n * 10). Predict the print order, then run. Then add .limit(2) after sorted(): how many saw lines appear now, and why not fewer?

  2. 2

    takeWhile on unsorted data

    Change the temperatures to start with 35. Predict the result of takeWhile(t -> t < 30) and dropWhile(t -> t < 30), then run. This is why takeWhile is mostly used on sorted or naturally ordered data.

  3. 3

    Find the bug in a reduction

    Replace reduce(0, Integer::sum) with reduce(1, Integer::sum). Predict total1 (it's off by one). Then write reduce(1, (a, b) -> a * b) over the prices: is 1 a correct identity for multiplication?

Code & diagrams

map, flatMap, mapMulti, distinct and sorted Java 16+ New tab
Sign in to run this example in your browser.

Expected output

map:      [[tea, samosa], [], [coffee, tea, cake]]
flatMap:  [tea, samosa, coffee, tea, cake]
distinct + sorted: [cake, coffee, samosa, tea]
mapMulti: [A1:tea, A1:samosa, A3:coffee, A3:tea, A3:cake]
lines to words: [hello, world, lazy, streams]
sorted is a barrier: watch the flow change Java 9+ New tab

Before sorted, elements flow one by one. sorted must see all four before it can send out 1, so every 'saw' comes before any 'got'.

Sign in to run this example in your browser.

Expected output

stateless only:
  saw 5
  got 50
  saw 2
  got 20
  saw 8
  got 80
  saw 1
  got 10
with sorted in the middle:
  saw 5
  saw 2
  saw 8
  saw 1
  got 10
  got 20
  got 50
  got 80
Short-circuiting, takeWhile, dropWhile and matching Java 16+ New tab
Sign in to run this example in your browser.

Expected output

limit on infinite: [64, 81, 100]
takeWhile < 30:  [18, 21, 25]
dropWhile < 30:  [31, 24, 35, 19]
filter < 30:     [18, 21, 25, 24, 19]
skip 2, limit 3: [25, 31, 24]
anyMatch > 30:   true
allMatch > 10:   true
noneMatch > 40:  true
empty anyMatch:  false
empty allMatch:  true
anyMatch on an infinite stream finished: true
Terminal operations: reduce, min, max, count, findFirst, toArray Java 16+ New tab

toArray() without an argument returns Object[]; passing String[]::new gives a properly typed array.

Sign in to run this example in your browser.

Expected output

1170 Optional[1170] 1170
empty reduce: Optional.empty
cheapest: pen, priciest: bag
items over 100: 2
first starting with b: book
[pen, book, bag] is a String[]
peek isn't guaranteed to run Java 9+ New tab

Pipeline A can't change the size of a sized list, so since Java 9 count() returns 3 without running it. On Java 8 'peek A' lines would print.

Sign in to run this example in your browser.

Expected output

count A = 3
peek B 2
peek B 3
count B = 2

Break it on purpose

Errors are the best teachers. Make each change, read the error, guess what went wrong, then reveal the answer.

Break #1

Return the wrong type from IntStream.map

Write IntStream.range(0, 3).map(i -> "item " + i).forEach(System.out::println);.

terminal
$ javac Main.java
── what you'll see ──
Main.java:5: error: incompatible types: bad return type in lambda expression
IntStream.range(0, 3).map(i -> "item " + i).forEach(System.out::println);
^
String cannot be converted to int
Note: Some messages have been simplified; recompile with -Xdiags:verbose to get full output
1 error

Break #2

Sort an infinite stream

Write Stream.iterate(1, n -> n + 1).sorted().limit(3).toList() hoping for [1, 2, 3].

terminal
$ java -Xmx64m Main.java
── what you'll see ──
Exception in thread "main" java.lang.OutOfMemoryError: Java heap space
at java.base/java.util.Arrays.copyOf(Arrays.java:3513)
at java.base/java.util.Arrays.copyOf(Arrays.java:3482)
at java.base/java.util.ArrayList.grow(ArrayList.java:237)
at java.base/java.util.ArrayList.grow(ArrayList.java:244)
at java.base/java.util.ArrayList.add(ArrayList.java:483)
at java.base/java.util.ArrayList.add(ArrayList.java:496)
at java.base/java.util.stream.SortedOps$RefSortingSink.accept(SortedOps.java:409)

Break #3

Call get() on an empty findFirst

Write names.stream().filter(n -> n.startsWith("z")).findFirst().get() when no name starts with z.

terminal
$ java Main.java
── what you'll see ──
Exception in thread "main" java.util.NoSuchElementException: No value present
at java.base/java.util.Optional.get(Optional.java:143)
at Main.main(Main.java:6)

Myth vs fact

Myth

limit(3) makes any pipeline cheap.

Fact

Only if nothing before it is a full barrier. sorted().limit(3) still buffers and sorts the entire input; limit only saves work for operations that come before it and are stateless.

Myth

takeWhile is a faster filter.

Fact

They mean different things. takeWhile stops at the first element that fails; filter checks them all. On unsorted data they give different results.

Myth

allMatch on an empty list returns false.

Fact

It returns true (vacuous truth). If an empty input should count as failure, check for emptiness separately.

Myth

peek is a good place for logging or metrics.

Fact

peek is not guaranteed to run (count() may skip it), and in parallel it runs on many threads in any order. Use it only for temporary debugging.

When it breaks

Metrics counted inside peek stop increasing after a Java upgrade

What you see

A job that did records.stream().peek(r -> metrics.increment()).count() reports zero processed records on Java 9+, because count() on a sized source no longer runs the pipeline. Dashboards show the job doing nothing while it actually works.

Fix & prevent

Don't put required side effects in peek or in intermediate lambdas. Count from the terminal result (long n = ...count(); metrics.add(n);) or do the side effect in forEach. Add a test that asserts the metric, so behavioural changes in the library are caught.

A stream with sorted() over a huge export runs the service out of memory

What you see

An endpoint does repository.streamAll().sorted(byDate).limit(100) over millions of rows. Heap spikes, long GC pauses, and finally OutOfMemoryError, because sorted buffers every row before limit sees one.

Fix & prevent

Push sorting and limiting into the database (ORDER BY ... LIMIT 100). If it must happen in memory, keep only the top N with a bounded PriorityQueue. Review every sorted() on an unbounded source.

Pro corner

Extra depth for experienced readers. New to this? Skip it for now and come back later.

  • ▸

    Stateful operations are what make parallel streams expensive: sorted sorts each chunk and merges; distinct on an ordered parallel stream must keep the first occurrence, which needs coordination; limit on an ordered parallel stream must track encounter positions. Calling .unordered() relaxes this and can make distinct and limit much cheaper when order doesn't matter.

  • ▸

    sorted() on a sequential stream collects into an ArrayList (or an exactly sized array when the size is known) and sorts with Arrays.sort/TimSort, which is stable, then replays. For Stream<T> without a comparator, elements must be Comparable or you get a ClassCastException at sort time, not compile time.

  • ▸

    Short-circuiting changes how the source is driven: with no short-circuit op, the source's forEachRemaining pushes everything; with one, the pipeline uses tryAdvance in a loop and checks cancellationRequested() between elements. That's also why flatMap before Java 10 wasn't lazy with findFirst: the inner stream was fully pushed (JDK-8075939, fixed in Java 10).

  • ▸

    Gatherers (JEP 485, final in Java 24) generalise intermediate operations: a Gatherer<T, A, R> has an initializer, an integrator that can push zero or more results and signal "stop", a combiner for parallel use, and a finisher. Gatherers.windowFixed(3), windowSliding(2), fold, scan and mapConcurrent(n, f) are built in. Like collectors for terminals, they turn custom stateful steps into reusable pieces.

Remember this

  1. 1

    Stateless intermediate operations handle each element independently: filter (keep or drop), map (one in, one out), flatMap (one in, zero or more out, flattened), mapMulti (Java 16, the same idea pushing to a consumer), peek (look without changing), and the mapToInt/mapToObj conversions. They cost nothing extra and work on infinite streams.

  2. 2

    Stateful intermediate operations need to remember earlier elements. distinct() keeps a set of everything it has seen (using equals/hashCode, Topic 4.8). sorted() must buffer every element before emitting the first one, so it's a barrier: it breaks the element-by-element flow and never finishes on an infinite stream. limit(n) and skip(n) keep a counter, and takeWhile/dropWhile (Java 9) remember whether the condition has failed yet.

  3. 3

    Short-circuiting operations can finish without looking at every element: limit and takeWhile among intermediates; anyMatch, allMatch, noneMatch, findFirst and findAny among terminals. They're what let Stream.iterate(1, n -> n + 1) (infinite) produce an answer. On an empty stream, anyMatch is false and allMatch is true ("every one of zero elements matches").

  4. 4

    reduce combines all elements into one value. reduce(identity, accumulator) needs an identity (a value that changes nothing: 0 for +, 1 for *, "" for concatenation) and an associative accumulator ((a + b) + c == a + (b + c)). Without an identity, reduce(accumulator) returns an Optional, empty for an empty stream. The three-argument form adds a combiner for when the result type differs from the element type.

  5. 5

    Encounter order is the order a source defines: a List has one, a HashSet doesn't. Sequential streams keep it, sorted creates it, and findFirst and forEachOrdered respect it even in parallel. forEach doesn't promise order in parallel (Topic 10.8). peek exists for debugging; it isn't guaranteed to run (since Java 9, count() can skip the pipeline entirely when the size is known).

  6. 6

    Since Java 24 (JEP 485), stream gatherers let you write your own intermediate operations with stream.gather(...). Gatherers ships ready-made ones such as fixed and sliding windows, fold, scan and mapConcurrent. They fill the gap that used to force people back into loops for "group every 3 elements" or "running total".

Explain it without notes

01

What is the difference between stateless and stateful intermediate operations? Give examples and explain why it matters.

02

Which operations are short-circuiting, and why do they matter for infinite streams?

03

Explain map versus flatMap, and when you'd use mapMulti.

04

What are the rules for the identity and accumulator in reduce, and what happens if you break them?

05

Why might a side effect inside peek or map not run?

Practice

01

Given sentences "the cat sat", "the dog ran", print the distinct words in alphabetical order using flatMap.

02

From List.of(3, 8, 12, 5, 20, 7), print the first two numbers greater than 6, and then the sum of all numbers using reduce with an identity.

03

Print the first 5 multiples of 7 that are also odd, starting from an infinite IntStream.iterate(7, n -> n + 7).

Trade-offs

  • ↔

    flatMap is expressive but allocates a stream per element; mapMulti is faster for small fan-outs but more verbose. Prefer flatMap for clarity unless profiling says otherwise.

  • ↔

    distinct and sorted are convenient but hold data in memory. For very large or unbounded inputs, a database query, an external sort, or a bounded data structure is safer than a stream barrier.

  • ↔

    reduce is general but easy to get wrong (identity, associativity, combiner). Specialised terminals (sum, max, count) and collectors (joining, summingInt, Topic 10.6) are clearer and harder to misuse.

Done when you can

  • Done when you can classify any stream operation as stateless, stateful or short-circuiting.

  • Done when you can predict the print order of a pipeline that contains sorted.

  • Done when you can choose correctly between map, flatMap and mapMulti.

  • Done when you know the results of the match operations on an empty stream.

  • Done when you can write all three forms of reduce with a correct identity.

  • Done when you can explain why peek and other side effects may not run.