Topic 10.5
Stream Operations in Depth
In one line
Stream operations fall into clear groups: stateless intermediate operations (filter, map, flatMap), stateful ones that must remember or buffer elements (distinct, sorted, limit), short-circuiting ones that can stop early (limit, anyMatch, findFirst), and terminal operations like reduce, min and toArray. Knowing which is which tells you what a pipeline costs and whether it can finish.
Think of it like this
An airport security line. The bag scanner looks at each passenger on their own and lets them through or not: it doesn't care who came before (stateless). The boarding desk that lines people up by seat row has to wait until everyone has arrived before it can let the first one through (stateful, a barrier). And the officer looking for "any passenger carrying a guitar" can stop the moment one appears (short-circuiting).
Words you'll meet
New words in this topic, in plain English. Come back here whenever one feels fuzzy.
- Stateless operation
- An operation that decides about each element on its own, without remembering any other element.
filterandmapare stateless. - Stateful operation
- An operation that must remember earlier elements, like
distinct(what it has seen) orsorted(everything, to sort it). - Barrier
- A point in the pipeline where all elements must arrive before any can continue.
sortedis a barrier. - Short-circuiting
- Able to produce its result without processing every element, like
findFirststopping at the first match. - flatMap
- An operation that turns each element into a stream of zero or more elements and joins all of those streams into one.
- Reduction
- Combining all elements into a single result, like a sum or a maximum.
- Identity
- A starting value that doesn't change the result when combined:
0for addition,1for multiplication. - Associative
- An operation where grouping doesn't matter:
(a + b) + cequalsa + (b + c). Subtraction isn't associative. - Encounter order
- The order in which a stream's source presents its elements, such as index order for a
List.
Step by step
01A map of the operations
Before memorising methods, learn the two questions to ask of any operation: does it need to remember other elements (stateful or stateless)? Can it stop early (short-circuiting)? The answers predict memory use, whether the pipeline works on infinite input, and how well it parallelises.
02filter, map and flatMap
filter(predicate) keeps elements for which the predicate is true. map(function) replaces each element with the function's result, so a Stream<Order> mapped with Order::customer becomes a Stream<String>.
flatMap is for one-to-many: each order has a list of items, and you want one stream of all items. map(o -> o.items()) would give a stream of lists; flatMap(o -> o.items().stream()) opens each inner stream and pours its elements into the outer one. An empty inner stream simply contributes nothing.
03mapMulti: flatMap without creating streams (Java 16)
flatMap creates a small stream object for every element, which is wasteful when each element produces only zero, one or two results. mapMulti((element, sink) -> ...) hands you a Consumer called the sink; you call sink.accept(x) once per output element, or not at all.
It's also handy for type filtering: .<Circle>mapMulti((shape, sink) -> { if (shape instanceof Circle c) sink.accept(c); }) filters and casts in one step. You often need the explicit type witness <Circle> because the compiler can't infer the output type from the lambda.
List<String> tagged = orders.stream()
.<String>mapMulti((o, sink) -> {
for (String item : o.items()) {
sink.accept(o.id() + ":" + item); // zero or more outputs per order
}
})
.toList();04Stateful operations: distinct, sorted, limit, skip
distinct() passes an element only if it hasn't seen an equal one before. It keeps a HashSet of seen elements, so memory grows with the number of distinct values, and it relies on correct equals and hashCode.
sorted() (natural order) and sorted(comparator) can't emit anything until the last element has arrived: the smallest might be the last one. So sorted collects everything into a buffer, sorts it, then pushes elements on. On an ordered stream the sort is stable. On an infinite stream it runs until memory runs out.
skip(n) drops the first n elements and limit(n) stops after n. Together they implement paging (skip(page * size).limit(size)). limit is short-circuiting: once n elements have passed, the source stops being read.
05takeWhile and dropWhile (Java 9)
takeWhile(p) passes elements while p is true and stops at the first failure, even if later elements would pass again. dropWhile(p) skips elements while p is true and passes everything after the first failure.
Compare with filter, which tests every element independently. For temperatures 18, 21, 25, 31, 24, takeWhile(t -> t < 30) gives 18, 21, 25, but filter(t -> t < 30) gives 18, 21, 25, 24. takeWhile is short-circuiting, so it also works on infinite sorted streams where filter would run forever.
06Matching and finding
anyMatch(p) stops at the first true; allMatch(p) stops at the first false; noneMatch(p) stops at the first true. On an empty stream they return false, true and true: that's standard logic ("all of nothing" is vacuously true), and a common source of bugs when an empty list means "no data" rather than "all good".
findFirst() returns the first element in encounter order as an Optional (Topic 10.7). findAny() may return any element; in a sequential stream it's usually the first, but in parallel it returns whichever is found first, which is faster when you don't care.
07reduce: three forms
reduce(0, Integer::sum): start from the identity and fold each element in. Always returns a value; for an empty stream, the identity.
reduce(Integer::sum): no identity, so the first element is the starting point, and the result is an Optional that's empty for an empty stream.
reduce(0, (sum, item) -> sum + item.price(), Integer::sum): the accumulator takes a result-so-far and an element of a different type, and the combiner merges two partial results. The combiner is only used in parallel, but it must still be correct. For most real reductions, mapToInt(...).sum() or a collector (Topic 10.6) is clearer.
int total = prices.stream().reduce(0, Integer::sum); // 0 if empty
Optional<Integer> t2 = prices.stream().reduce(Integer::sum); // Optional.empty if empty
int total3 = cart.stream().reduce(0,
(sum, item) -> sum + item.price(), // accumulator: Integer, Item
Integer::sum); // combiner: Integer, Integer
int best = prices.stream().reduce(Integer.MIN_VALUE, Math::max);08peek is for debugging, and side effects aren't guaranteed
peek(action) runs an action as each element passes, without changing it. It's handy to see what flows through a pipeline while debugging. Don't put real logic in it.
The library is allowed to skip work whose result can't change the answer. Since Java 9, list.stream().peek(...).count() returns the list's size without running peek at all, because nothing in the pipeline can change the count. Add a filter and the pipeline runs again. Code that counted events or updated metrics inside peek silently stopped working when teams upgraded from Java 8.
Try it yourself
- 1
Move the barrier
In the barrier example, move
.sorted()to after.map(n -> n * 10). Predict the print order, then run. Then add.limit(2)aftersorted(): how manysawlines appear now, and why not fewer? - 2
takeWhile on unsorted data
Change the temperatures to start with
35. Predict the result oftakeWhile(t -> t < 30)anddropWhile(t -> t < 30), then run. This is whytakeWhileis mostly used on sorted or naturally ordered data. - 3
Find the bug in a reduction
Replace
reduce(0, Integer::sum)withreduce(1, Integer::sum). Predicttotal1(it's off by one). Then writereduce(1, (a, b) -> a * b)over the prices: is1a correct identity for multiplication?
Code & diagrams
Expected output
map: [[tea, samosa], [], [coffee, tea, cake]]
flatMap: [tea, samosa, coffee, tea, cake]
distinct + sorted: [cake, coffee, samosa, tea]
mapMulti: [A1:tea, A1:samosa, A3:coffee, A3:tea, A3:cake]
lines to words: [hello, world, lazy, streams]Before sorted, elements flow one by one. sorted must see all four before it can send out 1, so every 'saw' comes before any 'got'.
Expected output
stateless only:
saw 5
got 50
saw 2
got 20
saw 8
got 80
saw 1
got 10
with sorted in the middle:
saw 5
saw 2
saw 8
saw 1
got 10
got 20
got 50
got 80Expected output
limit on infinite: [64, 81, 100]
takeWhile < 30: [18, 21, 25]
dropWhile < 30: [31, 24, 35, 19]
filter < 30: [18, 21, 25, 24, 19]
skip 2, limit 3: [25, 31, 24]
anyMatch > 30: true
allMatch > 10: true
noneMatch > 40: true
empty anyMatch: false
empty allMatch: true
anyMatch on an infinite stream finished: truetoArray() without an argument returns Object[]; passing String[]::new gives a properly typed array.
Expected output
1170 Optional[1170] 1170
empty reduce: Optional.empty
cheapest: pen, priciest: bag
items over 100: 2
first starting with b: book
[pen, book, bag] is a String[]Pipeline A can't change the size of a sized list, so since Java 9 count() returns 3 without running it. On Java 8 'peek A' lines would print.
Expected output
count A = 3
peek B 2
peek B 3
count B = 2Break it on purpose
Errors are the best teachers. Make each change, read the error, guess what went wrong, then reveal the answer.
Break #1
Return the wrong type from IntStream.map
Write IntStream.range(0, 3).map(i -> "item " + i).forEach(System.out::println);.
Break #2
Sort an infinite stream
Write Stream.iterate(1, n -> n + 1).sorted().limit(3).toList() hoping for [1, 2, 3].
Break #3
Call get() on an empty findFirst
Write names.stream().filter(n -> n.startsWith("z")).findFirst().get() when no name starts with z.
Myth vs fact
Myth
limit(3) makes any pipeline cheap.
Fact
Only if nothing before it is a full barrier. sorted().limit(3) still buffers and sorts the entire input; limit only saves work for operations that come before it and are stateless.
Myth
takeWhile is a faster filter.
Fact
They mean different things. takeWhile stops at the first element that fails; filter checks them all. On unsorted data they give different results.
Myth
allMatch on an empty list returns false.
Fact
It returns true (vacuous truth). If an empty input should count as failure, check for emptiness separately.
Myth
peek is a good place for logging or metrics.
Fact
peek is not guaranteed to run (count() may skip it), and in parallel it runs on many threads in any order. Use it only for temporary debugging.
When it breaks
Metrics counted inside peek stop increasing after a Java upgrade
What you see
A job that did records.stream().peek(r -> metrics.increment()).count() reports zero processed records on Java 9+, because count() on a sized source no longer runs the pipeline. Dashboards show the job doing nothing while it actually works.
Fix & prevent
Don't put required side effects in peek or in intermediate lambdas. Count from the terminal result (long n = ...count(); metrics.add(n);) or do the side effect in forEach. Add a test that asserts the metric, so behavioural changes in the library are caught.
A stream with sorted() over a huge export runs the service out of memory
What you see
An endpoint does repository.streamAll().sorted(byDate).limit(100) over millions of rows. Heap spikes, long GC pauses, and finally OutOfMemoryError, because sorted buffers every row before limit sees one.
Fix & prevent
Push sorting and limiting into the database (ORDER BY ... LIMIT 100). If it must happen in memory, keep only the top N with a bounded PriorityQueue. Review every sorted() on an unbounded source.
Pro corner
Extra depth for experienced readers. New to this? Skip it for now and come back later.
- ▸
Stateful operations are what make parallel streams expensive:
sortedsorts each chunk and merges;distincton an ordered parallel stream must keep the first occurrence, which needs coordination;limiton an ordered parallel stream must track encounter positions. Calling.unordered()relaxes this and can makedistinctandlimitmuch cheaper when order doesn't matter. - ▸
sorted()on a sequential stream collects into anArrayList(or an exactly sized array when the size is known) and sorts withArrays.sort/TimSort, which is stable, then replays. ForStream<T>without a comparator, elements must beComparableor you get aClassCastExceptionat sort time, not compile time. - ▸
Short-circuiting changes how the source is driven: with no short-circuit op, the source's
forEachRemainingpushes everything; with one, the pipeline usestryAdvancein a loop and checkscancellationRequested()between elements. That's also whyflatMapbefore Java 10 wasn't lazy withfindFirst: the inner stream was fully pushed (JDK-8075939, fixed in Java 10). - ▸
Gatherers (JEP 485, final in Java 24) generalise intermediate operations: a
Gatherer<T, A, R>has an initializer, an integrator that can push zero or more results and signal "stop", a combiner for parallel use, and a finisher.Gatherers.windowFixed(3),windowSliding(2),fold,scanandmapConcurrent(n, f)are built in. Like collectors for terminals, they turn custom stateful steps into reusable pieces.
Remember this
- 1
Stateless intermediate operations handle each element independently:
filter(keep or drop),map(one in, one out),flatMap(one in, zero or more out, flattened),mapMulti(Java 16, the same idea pushing to a consumer),peek(look without changing), and themapToInt/mapToObjconversions. They cost nothing extra and work on infinite streams. - 2
Stateful intermediate operations need to remember earlier elements.
distinct()keeps a set of everything it has seen (usingequals/hashCode, Topic 4.8).sorted()must buffer every element before emitting the first one, so it's a barrier: it breaks the element-by-element flow and never finishes on an infinite stream.limit(n)andskip(n)keep a counter, andtakeWhile/dropWhile(Java 9) remember whether the condition has failed yet. - 3
Short-circuiting operations can finish without looking at every element:
limitandtakeWhileamong intermediates;anyMatch,allMatch,noneMatch,findFirstandfindAnyamong terminals. They're what letStream.iterate(1, n -> n + 1)(infinite) produce an answer. On an empty stream,anyMatchisfalseandallMatchistrue("every one of zero elements matches"). - 4
reducecombines all elements into one value.reduce(identity, accumulator)needs an identity (a value that changes nothing:0for+,1for*,""for concatenation) and an associative accumulator ((a + b) + c == a + (b + c)). Without an identity,reduce(accumulator)returns anOptional, empty for an empty stream. The three-argument form adds a combiner for when the result type differs from the element type. - 5
Encounter order is the order a source defines: a
Listhas one, aHashSetdoesn't. Sequential streams keep it,sortedcreates it, andfindFirstandforEachOrderedrespect it even in parallel.forEachdoesn't promise order in parallel (Topic 10.8).peekexists for debugging; it isn't guaranteed to run (since Java 9,count()can skip the pipeline entirely when the size is known). - 6
Since Java 24 (JEP 485), stream gatherers let you write your own intermediate operations with
stream.gather(...).Gatherersships ready-made ones such as fixed and sliding windows,fold,scanandmapConcurrent. They fill the gap that used to force people back into loops for "group every 3 elements" or "running total".
Explain it without notes
What is the difference between stateless and stateful intermediate operations? Give examples and explain why it matters.
Which operations are short-circuiting, and why do they matter for infinite streams?
Explain map versus flatMap, and when you'd use mapMulti.
What are the rules for the identity and accumulator in reduce, and what happens if you break them?
Why might a side effect inside peek or map not run?
Practice
Given sentences "the cat sat", "the dog ran", print the distinct words in alphabetical order using flatMap.
From List.of(3, 8, 12, 5, 20, 7), print the first two numbers greater than 6, and then the sum of all numbers using reduce with an identity.
Print the first 5 multiples of 7 that are also odd, starting from an infinite IntStream.iterate(7, n -> n + 7).
Trade-offs
- ↔
flatMapis expressive but allocates a stream per element;mapMultiis faster for small fan-outs but more verbose. PreferflatMapfor clarity unless profiling says otherwise. - ↔
distinctandsortedare convenient but hold data in memory. For very large or unbounded inputs, a database query, an external sort, or a bounded data structure is safer than a stream barrier. - ↔
reduceis general but easy to get wrong (identity, associativity, combiner). Specialised terminals (sum,max,count) and collectors (joining,summingInt, Topic 10.6) are clearer and harder to misuse.
Done when you can
Done when you can classify any stream operation as stateless, stateful or short-circuiting.
Done when you can predict the print order of a pipeline that contains
sorted.Done when you can choose correctly between
map,flatMapandmapMulti.Done when you know the results of the match operations on an empty stream.
Done when you can write all three forms of
reducewith a correct identity.Done when you can explain why
peekand other side effects may not run.