Topic 14.4
The JIT Compiler
In one line
HotSpot starts by interpreting bytecode, counts which methods and loops are hot, and compiles those to machine code: quickly with C1, then aggressively with C2 using the profile it gathered. C2's big wins are inlining, escape analysis and speculative optimisations that are undone (deoptimised) if their assumptions break.
Think of it like this
A new waiter reads every order from the menu book, word by word (the interpreter). After taking the same order a few hundred times he writes a quick cheat card for it (C1). For the handful of orders customers ask for all day, the head chef designs a perfect, pre-prepared routine based on what customers actually order (C2). If the menu suddenly changes, the routine is thrown away and the waiter goes back to the book until he learns the new pattern (deoptimisation).
Words you'll meet
New words in this topic, in plain English. Come back here whenever one feels fuzzy.
- JIT compiler
- Just-in-time compiler: it turns bytecode into machine code while the program runs, for the parts that run often.
- Interpreter
- The part of the JVM that runs bytecode one instruction at a time, without compiling it. Slow, but starts instantly.
- C1 and C2
- HotSpot's two JIT compilers. C1 compiles quickly with simple optimisations; C2 compiles slowly but produces much faster code.
- Tiered compilation
- Running code first in the interpreter, then C1, then C2 as it gets hotter.
- Profile
- Statistics the JVM collects while code runs: how often branches are taken, which classes appear at each call.
- Inlining
- Copying a called method's body into the caller, so there is no call left and both can be optimised together.
- Escape analysis
- Checking whether an object can be seen outside the method that created it. If not, the JIT may skip creating it on the heap.
- Deoptimisation
- Throwing away compiled code whose assumptions turned out wrong and going back to the interpreter.
- OSR (on-stack replacement)
- Switching a method that is in the middle of a long loop from interpreted to compiled code without waiting for it to return.
- Warm-up
- The early period after start-up when code is still interpreted or lightly compiled, so the program runs slower.
Step by step
01The tiers and how code moves between them
Every method starts in tier 0. Counters in the interpreter (and in tier 3 code) trigger compilation requests that go into a queue served by background compiler threads; your thread keeps running the old version until the new one is installed. If a C2 assumption breaks, the method drops back to tier 0 and climbs again.
02Watch the JIT at work
-XX:+PrintCompilation prints one line per compilation: milliseconds since start, a compile id, flags (% = OSR, s = synchronized, ! = has exception handlers, n = native), the tier, the method and its bytecode size. made not entrant marks a version being replaced, either because a higher tier took over or because of deoptimisation. The output is sample output; yours differs in every detail.
03Inlining decisions
With diagnostic options you can see each inlining decision. Reasons such as inline (hot), too big, hot method too big and virtual call explain why a call was or wasn't inlined. This is sample output; the exact wording differs slightly between JDK versions.
04Speculation and uncommon traps
C2 compiles for the program it has seen, not every program that could run. A branch that was never taken is compiled as an uncommon trap instead of real code; a call site that only saw ArrayList gets a class check plus the inlined ArrayList.get. This is what makes Java fast, and also why a rare code path (an error branch, a new subclass loaded by a plugin) causes a short slow-down: deoptimisation, a bit of interpreted execution, recompilation.
-Xlog:deoptimization=debug (recent JDKs) or JFR's jdk.Deoptimization event (Topic 14.5) show when and why it happens. Code that deoptimises over and over is usually polymorphic in an unexpected way.
05Escape analysis in practice
Escape analysis works per compiled method *after inlining*. An object passed to a method that wasn't inlined is treated as escaping, which is one reason inlining matters so much. Objects stored in fields, static variables, arrays that escape, or passed to other threads all escape.
The effect is real: iterators in an enhanced for over an ArrayList, Optionals and small records in hot loops often cost nothing after C2. But it is not guaranteed and it's easily lost when a method grows too big to inline, so don't design around it; just don't fear small short-lived objects.
06Measuring correctly with JMH
Hand-written timing loops lie: the first iterations are interpreted, the JIT may delete a loop whose result is unused (dead-code elimination), fold constants, or hoist work out of the loop. JMH runs warm-up iterations, forks fresh JVMs, and provides Blackhole to consume results so they can't be eliminated.
import org.openjdk.jmh.annotations.*;
import java.util.concurrent.TimeUnit;
@BenchmarkMode(Mode.AverageTime)
@OutputTimeUnit(TimeUnit.NANOSECONDS)
@Warmup(iterations = 5) @Measurement(iterations = 5) @Fork(2)
@State(Scope.Thread)
public class SquareBench {
int x = 42; // a field, so the JIT can't fold it as a constant
@Benchmark
public int square() { return x * x; } // return the result so JMH consumes it
}07Start-up vs peak performance
The JIT gives excellent peak performance but costs CPU at start-up and needs warm-up. Options: -XX:TieredStopAtLevel=1 for short-lived tools (fast start, lower peak), CDS and AOT class loading (Topic 14.1), sending some warm-up traffic before marking an instance ready, or GraalVM native images when start-up matters more than peak throughput.
Try it yourself
- 1
See the tiers for real
Run the first example on your own JDK with
java -XX:+PrintCompilation Main.java | grep Main::. Find the lines forMain::squareand note which tiers appear. Then run with-Xintand with-XX:TieredStopAtLevel=1and see what changes in the compilation log. - 2
Make a call site megamorphic
Add a fourth shape,
record Circle(int r), to "What the profile sees at a call site", and calltotalArea100,000 times with only squares first, then once with a circle. Run with-XX:+UnlockDiagnosticVMOptions -XX:+PrintInlining -XX:+PrintCompilationand look formade not entrantafter the circle appears. - 3
Watch escape analysis
Make
n50,000,000 in "Objects that escape and objects that don't", call onlysumLocal, and run with-Xlog:gc. Then add-XX:-DoEscapeAnalysisand run again. Predict which run shows more young GCs (the second: everyPointbecomes a real allocation).
Code & diagrams
A teaching model: the JVM's real decisions also depend on loop counts and compiler load, so run with -XX:+PrintCompilation to see the true moments on your machine.
Expected output
call 1 -> tier 0: interpreted, counting calls
call 200 -> tier 3: C1 code that also collects a profile
call 5000 -> tier 4: C2 code optimised from the profile
sum of squares 1..6000 = 72018001000The program models the receiver-type profile C2 uses. In a real JVM, the moment a third type shows up at a hot monomorphic site, compiled code hits its type guard and deoptimises.
Expected output
13 | [Square] | monomorphic: one type check, then area() inlined
10 | [Square, Rect] | bimorphic: two type checks, both bodies inlined
6 | [Square, Rect, Triangle] | megamorphic: real virtual call, no inliningSame result, different cost once hot: compare the allocation rate of each method with JFR or -Xlog:gc on your own JDK, using a much larger n.
Expected output
no escape: 5150
escapes: 5150
points kept: 100, last Point[x=5050, y=100]-XX:+PrintCompilation # one line per compiled method
-XX:+UnlockDiagnosticVMOptions -XX:+PrintInlining # inlining decisions
-Xint # interpreter only (for comparison)
-XX:TieredStopAtLevel=1 # C1 only: faster start-up, lower peak
-XX:-TieredCompilation # C2 only after interpretation (old behaviour)
-XX:ReservedCodeCacheSize=256m # room for compiled code
-XX:-DoEscapeAnalysis # turn escape analysis off, to measure its effectBreak it on purpose
Errors are the best teachers. Make each change, read the error, guess what went wrong, then reveal the answer.
Break #1
Benchmarking with a hand-written loop
Time Math.sqrt in a loop with System.nanoTime() and never use the result.
long start = System.nanoTime();
for (int i = 0; i < 100_000_000; i++) {
Math.sqrt(i); // result thrown away
}
System.out.println((System.nanoTime() - start) / 1_000_000 + " ms");Break #2
Code cache full
Run a large application with -XX:ReservedCodeCacheSize=4m.
Myth vs fact
Myth
Java is interpreted, so it's slow.
Fact
Hot code is compiled to optimised machine code, often using run-time information a static compiler doesn't have. Steady-state Java is competitive with C++ for most server workloads.
Myth
Making methods final makes them faster.
Fact
C2 already inlines monomorphic virtual calls using the profile and class hierarchy analysis. final is for design, not speed.
Myth
Small methods are slow because of call overhead.
Fact
Small hot methods are inlined and the call disappears. Huge methods are the problem: they exceed inlining limits and compile less well.
Myth
new always allocates on the heap.
Fact
Semantically yes, but if the object doesn't escape, C2 may scalar-replace it and never allocate it.
Interview problem
The problem
Slow for the first minute after every deploy
A Java service responds in 15 ms at p99 normally, but for the first 60-90 seconds after each deploy, p99 is 400 ms and some requests time out. What's happening and what would you do?
You're given
- Kubernetes rolling deploys, 2 CPU per pod
- Load balancer sends full traffic as soon as the readiness probe passes
- No errors in the logs
The interviewer follows up
Would -XX:TieredStopAtLevel=1 help?
Would a GraalVM native image help?
Pro corner
Extra depth for experienced readers. New to this? Skip it for now and come back later.
- ▸
MaxInlineLevel(nesting depth) was raised from 9 to 15 in JDK 14, which helped deeply layered code such as streams and Scala/Kotlin collections. Inlining decisions are per call site and per profile, so the same method can be inlined in one caller and not another. - ▸
Interface and virtual calls that are megamorphic go through the itable/vtable. Code that walks a
List<Shape>with many implementations can be sped up by splitting by type, or with aswitchover a sealed hierarchy (Topic 11.2), which compiles to type checks the JIT can optimise. - ▸
Loops counted with
int("counted loops") get range-check elimination, unrolling and vectorisation;longinduction variables historically optimised worse (improved in recent JDKs). Counted loops historically had no safepoint polls inside, which could delay GC pauses ("time to safepoint"); loop strip mining (JDK 10, on by default with the low-pause collectors such as G1 and ZGC) largely fixed that. - ▸
JIT compilation uses its own threads (
C1 CompilerThread0,C2 CompilerThread0in thread dumps) and CPU. In a container with one CPU, compilation competes with your code during warm-up;-XX:CICompilerCountcontrols the thread count.
Remember this
- 1
javacproduces portable bytecode and does almost no optimisation. The JVM's Just-In-Time (JIT) compiler does the optimising at run time, when it can see the real CPU and the real behaviour of the program. HotSpot has two JIT compilers: C1 (the "client" compiler) compiles fast with simple optimisations, and C2 (the "server" compiler) compiles slowly but produces highly optimised code. Tiered compilation (the default since Java 8) uses both. - 2
The tiers: 0 interpreter, 1 C1 without profiling (for trivial methods), 2 C1 with light profiling, 3 C1 with full profiling, 4 C2. A typical hot method goes 0 -> 3 -> 4. Promotion is driven by counters of invocations and loop back-edges (jumps back to a loop's start), with defaults such as
Tier3InvocationThreshold=200andTier4InvocationThreshold=5000, adjusted for how busy the compiler threads are. A long-running loop inside a method that is called once is compiled while it runs, by on-stack replacement (OSR). - 3
Inlining is the most important optimisation, because it opens the door to all others: the callee's body is copied into the caller, removing the call and letting the compiler optimise both together. C2 inlines small methods always (bytecode under
MaxInlineSize, 35 bytes) and hot methods up toFreqInlineSize(325 bytes on x86-64), up to a nesting depth. Small methods such as getters cost nothing in hot code; giant methods don't get inlined and are harder to optimise. - 4
Java calls are virtual by default, so how can C2 inline
shape.area()? With the profile: if the call site has only ever seenSquare(monomorphic), C2 inlinesSquare.area()behind a cheap class check; with two types (bimorphic) it inlines both; with three or more (megamorphic) it usually makes a real virtual call. If the check ever fails, the compiled code hits an uncommon trap, the method is deoptimised (marked "not entrant", execution continues in the interpreter) and later recompiled with the new profile. Class Hierarchy Analysis (CHA) also lets C2 inline a method that no loaded subclass overrides, and deoptimise if one is loaded later. - 5
Escape analysis asks whether an object can be seen outside the method (or thread) that created it. If it can't escape, C2 can scalar-replace it (keep its fields in registers, so no heap allocation at all), and remove locks on it (lock elision). Other C2 optimisations include constant folding, dead-code elimination, loop unrolling, range-check elimination, auto-vectorisation with SIMD instructions, and intrinsics: hand-written machine code for methods such as
System.arraycopy,Mathfunctions,String.equalsandArrays.fill. - 6
Consequences for you: a Java program has a warm-up period during which it's slower; measuring with
System.nanoTimearound a single call gives wrong answers (use JMH, the Java Microbenchmark Harness); compiled code lives in the code cache (Topic 14.2); and-Xint(interpreter only) or-XX:TieredStopAtLevel=1(C1 only, faster start-up) change the trade-off. GraalVM offers an alternative JIT written in Java and native-image ahead-of-time compilation; the experimental Graal JIT that shipped inside OpenJDK 10-16 was removed in JDK 17.
Explain it without notes
Explain tiered compilation in HotSpot.
Why is inlining called the mother of all optimisations?
What is escape analysis and what does it enable?
What is deoptimisation, and when does it happen?
Practice
Write a program that counts how many times a small method is called inside a loop and prints the count at which it would cross the tier-3 and tier-4 thresholds (200 and 5000), for loop sizes 100, 1000 and 10000.
Rewrite a loop over a List<Shape> with three shape types so that each type is processed in its own loop (keeping call sites monomorphic), and show that the total area is the same.
Write a method that builds a StringBuilder, appends three values and returns the String, and explain in a comment whether the StringBuilder escapes.
Trade-offs
- ↔
JIT compilation gives peak performance tuned to real behaviour, at the cost of warm-up time and CPU at start-up; AOT (native images) flips that trade-off.
- ↔
Speculative optimisation is fast in the common case but causes deoptimisation hiccups when behaviour changes.
- ↔
C1-only (
TieredStopAtLevel=1) suits short-lived processes; full tiered compilation suits long-running servers.
Done when you can
Done when you can describe tiers 0 to 4 and what moves code between them.
Done when you can explain inlining, monomorphic/bimorphic/megamorphic call sites and deoptimisation.
Done when you can explain escape analysis and scalar replacement.
Done when you can read -XX:+PrintCompilation output.
Done when you can explain why benchmarks need JMH and warm-up.