Java Streams API Tutorial: Filter, Map, Reduce and Collect with Runnable Examples

Learn Java streams from the source-to-terminal mental model, then trace operations, reductions and collectors through runnable examples and practice tasks.

KnowledgeGate Team

Exam prep & CS education

Updated 24 Sep 20265 min read

You can read a stream chain of filter, map, sorted and collect, yet still be unable to predict what each stage produces or debug where the result goes wrong. Trace any pipeline stage by stage to see what each operation produces and which operation to use. Once you master tracing, continue your wider Java practice through the Coding & Skills courses.

Java Streams API: what a stream is and is not

A stream is a one-way pipeline that processes elements from a source. A List<Integer> stores values, while list.stream() describes how those values will be processed. Creating the stream neither copies the list nor turns it into an indexable container.

Use this three-part model:

  1. A source, such as a list

  2. Zero or more intermediate operations

  3. One terminal operation

filter, map, distinct and sorted are intermediate operations. collect, count, reduce, max and forEach are terminal operations.

For a first example, start with [3, 6, 7, 10]:

java
List<Integer> numbers = Arrays.asList(3, 6, 7, 10);
List<Integer> even = numbers.stream()
    .filter(n -> n % 2 == 0)
    .collect(Collectors.toList());

The lambda is a predicate. It keeps 6 and 10, so even is [6, 10]. The original list remains [3, 6, 7, 10].

Build and trace one complete stream pipeline

This runnable program traces filter and map, collects the transformed values, then reduces the original passing scores to one sum:

java
import java.util.*;
import java.util.stream.*;

public class StreamPipelineDemo {
    public static void main(String[] args) {
        List<Integer> scores = Arrays.asList(42, 75, 63, 88, 75, 51, 39);

        List<Integer> finalScores = scores.stream()
            .filter(score -> score >= 60)
            .map(score -> score + 5)
            .distinct()
            .sorted(Comparator.reverseOrder())
            .collect(Collectors.toList());

        int passingTotal = scores.stream()
            .filter(score -> score >= 60)
            .reduce(0, Integer::sum);

        System.out.println(finalScores); // [93, 80, 68]
        System.out.println(passingTotal); // 301
    }
}

Trace it from left to right:

Stage

Value

Source

[42, 75, 63, 88, 75, 51, 39]

After filter

[75, 63, 88, 75]

After map

[80, 68, 93, 80]

After distinct

[80, 68, 93]

After descending sorted

[93, 80, 68]

After collect

A new List<Integer> containing [93, 80, 68]

The duplicate 80 disappears only after mapping because the two source values of 75 both become 80. The list is the source, the four transformations are intermediate operations, and collect is the terminal operation that starts traversal and materialises the result.

Six-stage Java stream diagram showing source scores, filtering, a box labelled 'man score + 5', duplicate removal, descending sorting and collection into [93, 80, 68].

Intermediate operations: filter, map, flatMap, distinct and sorted

Intermediate operations are lazy. Defining a pipeline does not traverse the source. Traversal begins only when a terminal operation is called. filter and map work element by element, while distinct and sorted may need to remember elements while processing.

Suppose the source is [ ["Java", "SQL"], ["Java", "DSA"] ]. This pipeline flattens the nested lists, removes the repeated value, and sorts the result:

java
List<String> topics = nested.stream()
    .flatMap(List::stream)
    .distinct()
    .sorted()
    .collect(Collectors.toList());

Flattening first gives ["Java", "SQL", "Java", "DSA"]. The final result is ["DSA", "Java", "SQL"]. map(List::stream) would produce a stream of streams, but flatMap(List::stream) produces one stream of strings.

Order also changes meaning. On [1, 2, 3, 4], filtering even numbers and then squaring gives [4, 16]. Squaring first and then keeping values greater than 4 gives [9, 16]. Those pipelines answer different questions.

Terminal operations, reduction and primitive streams

Create a fresh stream for each result from the original scores:

java
long count = scores.stream().filter(s -> s >= 60).count();
int sum = scores.stream().filter(s -> s >= 60).reduce(0, Integer::sum);
Optional<Integer> max = scores.stream().filter(s -> s >= 60).max(Integer::compareTo);
OptionalDouble average = scores.stream().filter(s -> s >= 60)
    .mapToInt(Integer::intValue).average();

The four passing scores are 75, 63, 88 and 75. Therefore, count is 4, sum is 75 + 63 + 88 + 75 = 301, max is Optional[88], and average is 301 / 4 = 75.25, represented as OptionalDouble[75.25].

An empty input has no maximum or average, which is why these operations return optional results. max.orElse(0) supplies a deliberate fallback. For numeric work, mapToInt also gives direct access to sum() and average() without unnecessary boxing.

A stream object can be consumed only once. After stream.count(), calling stream.findFirst() on the same object throws IllegalStateException. Call scores.stream() again for a new traversal.

Collect results into lists, strings and groups

Collectors build useful result containers. Given skills = ["Java", "SQL", "DSA", "Java", "Spring", "SQL"]:

java
Map<String, Long> frequency = skills.stream().collect(
    Collectors.groupingBy(skill -> skill, Collectors.counting()));

String names = skills.stream()
    .distinct()
    .sorted()
    .collect(Collectors.joining(", "));

The logical frequency entries are Java=2, SQL=2, DSA=1 and Spring=1. A default Map does not guarantee any particular key order when printed. The joined string is exactly "DSA, Java, SQL, Spring".

Use toList or toSet for collections, joining for one string, and groupingBy for a map of groups. Use reduce when combining elements into one value, such as the score sum 301. It is not a replacement for every collector.

Common Java stream mistakes and what to do instead

Do not reuse a consumed stream. Create a fresh stream from its collection. Also avoid adding to or removing from the source list during traversal, because that interferes with the pipeline. Produce a new result instead.

Mutating an external list inside forEach is another tempting pattern. It is harder to reason about and becomes fragile with parallel execution. Let the pipeline return its result:

java
List<String> cleaned = words.stream()
    .filter(Objects::nonNull)
    .map(String::trim)
    .collect(Collectors.toList());

Filtering with Objects::nonNull before String::trim also prevents a null dereference.

Finally, parallelStream() is not an automatic speed switch. Small inputs, ordered work, shared mutable state and cheap per-element operations can erase any benefit. Prefer a clear sequential stream until measurement identifies a real bottleneck.

How exercises and interviews test stream reasoning

For words = ["gate", "java", "api", "stream", "java"], filter for length at least 4, convert to uppercase, remove duplicates, sort naturally, and collect. The stages are:

  1. After filtering: ["gate", "java", "stream", "java"]

  2. After mapping: ["GATE", "JAVA", "STREAM", "JAVA"]

  3. After distinct: ["GATE", "JAVA", "STREAM"]

  4. After sorting: ["GATE", "JAVA", "STREAM"]

The sum of the final string lengths is 4 + 4 + 6 = 14.

Now take [5, 12, 7, 12, 20, 3]. Retain values greater than 5, square them, remove duplicates and sort ascending:

  • After filter: [12, 7, 12, 20]

  • After square: [144, 49, 144, 400]

  • After distinct: [144, 49, 400]

  • Final result: [49, 144, 400]

Pipeline judgement includes cost as well as output. Once you can trace the result, use time complexity and asymptotic notation to reason about how the work changes with the input size, and sorting algorithms: complexity and comparison to understand why sorting requires broader reasoning than element-wise transformation. Since sorted() is stateful, include it only when the output order is required.

The short version and the next Java step

Identify the source, read intermediate operations from left to right, compute each stage, and finish with one terminal operation. Create a fresh stream if you need another result. Use the final list [93, 80, 68] and the filtered-score sum 301 as quick self-checks.

For structured language practice, continue with the Java course with concepts, MCQs and coding questions. When you are ready to apply Java alongside data structures and problem solving, move to DSA using Java.