Topic 14.5
JVM Tools: jcmd, jstack, jmap and JFR
In one line
The JDK ships tools that look inside a running JVM without restarting it: jcmd (the Swiss-army knife), jstack for thread dumps, jmap and heap dumps for memory, jstat for GC counters, and Java Flight Recorder (JFR) for low-overhead, always-on recording of what the JVM did. Knowing which one answers which question is most of production troubleshooting.
Think of it like this
A doctor examining a patient. A photo of the patient right now shows posture (a thread dump: what every thread is doing at this instant). An X-ray shows everything inside (a heap dump: every object and who holds it). A heart-rate monitor worn all day records a stream of measurements you can scroll back through after the patient felt dizzy (JFR). And the nurse who can take temperature, blood pressure or a blood test on request is **jcmd**. You choose the tool by the question you're asking.
Words you'll meet
New words in this topic, in plain English. Come back here whenever one feels fuzzy.
- PID
- Process id: the number the operating system gives a running program. Tools use it to pick which JVM to look at.
- Thread dump
- A snapshot of every thread in a JVM: its name, its state and the methods it is inside right now.
- Heap dump
- A file containing every object in the heap and the references between them, for finding what uses memory.
- Retained size
- How much memory would be freed if one object were collected, counting everything only it keeps alive.
- Dominator tree
- A view of a heap dump that shows which objects keep which others alive, biggest first.
- JFR (Java Flight Recorder)
- A recorder built into the JVM that logs events such as GC pauses, hot methods and lock waits with very little overhead.
- JMX
- Java Management Extensions: the standard way for a JVM to publish metrics and accept management commands, locally or over the network.
- Attach API
- The mechanism that lets tools like
jcmdconnect to a JVM that is already running.
Step by step
01Find the process and ask it questions
jcmd with no arguments lists the JVMs you can attach to. Then every question is jcmd <pid> <command>. Sample output below.
02Reading a thread dump
Each thread entry has a header line and a stack. In the header: the thread name (name your pool threads, Topic 13.6), daemon, the CPU time it has used, nid (the OS thread id in hex) and a short state. The next line is the Java state, then the stack, innermost frame first. Lines starting with - show locks: locked <0x...> (held), waiting to lock <0x...> (blocked on it), parking to wait for (a java.util.concurrent lock or condition).
Sample output; the exact header format changes a little between JDK versions.
03Deadlocks are reported for you
At the end of a thread dump the JVM checks for cycles of threads waiting on each other's monitors or java.util.concurrent locks and prints them. This is the same check ThreadMXBean.findDeadlockedThreads() performs, which the third runnable example uses. Sample output:
04Memory: histogram first, heap dump when needed
GC.class_histogram is quick and safe to run a few times a minute apart: a class whose instance count only grows is your leak suspect. The [B and [C names are byte[] and char[]. Sample output:
When you need *who holds them*, take a heap dump and open it in MAT: start from the dominator tree, pick the biggest retained size, and follow path to GC roots. The live option (and jcmd GC.heap_dump, which dumps live objects by default) runs a full GC first so the dump contains only reachable objects.
05GC at a glance with jstat
jstat -gcutil <pid> <interval> prints occupancy percentages (survivors S0/S1, eden E, old O, Metaspace M, compressed class space CCS), young/full/concurrent GC counts and the time spent in each. Old occupancy that keeps climbing after each cycle points to a growing live set. Sample output:
06Recording with JFR
Start a recording at launch or on a running JVM, then analyse it. settings=profile collects more detail than the default with a little more overhead. The jfr tool prints, summarises and (in JDK 21) shows predefined views such as hot-methods and gc with jfr view. Sample output:
07High CPU: from OS thread to Java stack
top -H -p <pid> shows CPU per OS thread. Convert the busiest thread id to hex (printf '%x\n' 11295 prints 2c1f) and search the thread dump for nid=0x2c1f. If it's a GC or C2 compiler thread, the problem is memory or warm-up; if it's an application thread, its stack shows the hot loop. JFR's method profiling (jdk.ExecutionSample) gives the same answer statistically over time.
Try it yourself
- 1
Take a real thread dump
Change the deadlock example so the threads are not daemons and
maindoesn't detect anything (justt1.join()). Run it on your own JDK; it hangs. In another terminal runjcmdto find the PID, thenjcmd <pid> Thread.print. Find the twotransfer-threads and theFound one Java-level deadlocksection at the end. - 2
Record and read a flight recording
Run any program that loops for 30 seconds with
java -XX:StartFlightRecording=duration=30s,filename=rec.jfr Main.java. Then runjfr summary rec.jfrto see which event types were recorded andjfr print --events jdk.ExecutionSample rec.jfr | head -40to see sampled stacks. If you have JDK Mission Control, open the file and look at the Method Profiling page. - 3
Watch a leak in the histogram
Run the leaky
SESSIONSmap from Topic 14.3 in an endless loop with aThread.sleep(1)per request. Takejcmd <pid> GC.class_histogram | head -8three times, a minute apart. Predict which lines grow ([Bandjava.util.HashMap$Node, andjava.lang.Integerfor the keys).
Code & diagrams
StackWalker (Java 9) walks frames lazily, so it's much cheaper than creating an exception to get a stack trace. Naming threads is what makes dumps readable.
Expected output
"main" RUNNABLE
at Main.queryDatabase
at Main.main
"http-worker-1" RUNNABLE
at Main.queryDatabase
at Main.loadOrder
at Main.handleRequestIn a dump, many threads BLOCKED on one lock means contention; many WAITING (parking) inside a pool or connection pool means they're starved of a resource. The daemon flag lets the JVM exit while sleeper and waiter are still parked.
Expected output
sleeper TIMED_WAITING
waiter WAITING
blocked BLOCKED
blocked TERMINATEDThe latch guarantees each thread holds its first lock before reaching for the second, so the deadlock happens every time. Lock ordering (Topic 13.9) is the cure.
Expected output
deadlocked threads: 2
transfer-1 is BLOCKED, waiting for a lock held by transfer-2
transfer-2 is BLOCKED, waiting for a lock held by transfer-1import jdk.jfr.*;
@Name("com.acme.OrderPlaced")
@Label("Order Placed")
@Category("Orders")
class OrderPlaced extends Event {
@Label("Order id") long orderId;
@Label("Items") int items;
}
public class Main {
public static void main(String[] args) {
for (long id = 1; id <= 3; id++) {
OrderPlaced e = new OrderPlaced();
e.begin();
e.orderId = id;
e.items = (int) id * 2;
// ... place the order ...
e.commit(); // recorded only if a recording with this event enabled is running
}
}
}
// java -XX:StartFlightRecording=filename=orders.jfr Main.java
// jfr print --events com.acme.OrderPlaced orders.jfr
// Events show up in JMC next to GC pauses and CPU samples, with duration and thread.Break it on purpose
Errors are the best teachers. Make each change, read the error, guess what went wrong, then reveal the answer.
Break #1
Attaching as the wrong user
Run jcmd <pid> Thread.print as your own user against a JVM started by a service account (or from outside its container).
Break #2
Heap dump that fills the disk
Enable -XX:+HeapDumpOnOutOfMemoryError with an 8 GB heap and no HeapDumpPath, in a container with a small writable layer.
Myth vs fact
Myth
Profiling production is too risky.
Fact
JFR is designed for production: the default settings typically cost around 1% overhead. Many teams run it continuously.
Myth
One thread dump tells you what's slow.
Fact
A single dump is one instant. Take several a few seconds apart, or use JFR sampling, to separate stuck threads from busy ones.
Myth
jmap is the only way to get a heap dump.
Fact
jcmd <pid> GC.heap_dump is the recommended way; -XX:+HeapDumpOnOutOfMemoryError captures one at the moment of failure; JMX's HotSpotDiagnosticMXBean can do it from code.
When it breaks
All 200 Tomcat threads are WAITING in HikariPool.getConnection and the service stops responding.
What you see
Thread dumps show the pool exhausted; one slow query or a connection leak (a connection not closed on an error path) holds connections for minutes.
Fix & prevent
Close connections with try-with-resources, set query and connection timeouts, enable the pool's leak detection, and size the pool against the database's limits.
An incident happened at 03:12 and by the time anyone looks, the JVM was restarted.
What you see
No thread dump, no heap dump, no profile: the root cause is guessed.
Fix & prevent
Run JFR continuously (-XX:StartFlightRecording=maxage=30m,disk=true,dumponexit=true,filename=/data/app.jfr), collect GC logs to a volume, and dump heap on OOM. Then the evidence survives the restart.
Pro corner
Extra depth for experienced readers. New to this? Skip it for now and come back later.
- ▸
Thread dumps only show threads at safepoints, so a thread in a tight compiled loop is shown at the nearest safepoint poll, not the exact instruction (safepoint bias). JFR's execution sampler and async-profiler sample threads without waiting for safepoints, which makes their CPU profiles more accurate.
- ▸
JFR event streaming (JDK 14,
jdk.jfr.consumer.RecordingStream) lets an application consume its own events live, for example to export GC pause times or lock contention as metrics without parsing files. - ▸
For virtual threads (Topic 13.10), a classic thread dump doesn't list the possibly millions of virtual threads;
jcmd <pid> Thread.dump_to_file -format=json <file>(JDK 21) writes all of them, grouped by their containers. - ▸
jcmd <pid> VM.native_memory baselineand laterVM.native_memory summary.diffgive a native-memory diff over time, the fastest way to find which non-heap area is growing (Topic 14.2).
Remember this
- 1
**
jcmdtalks to a running JVM through the Attach API and runs diagnostic commands**.jcmdalone lists Java processes;jcmd <pid> helplists commands. The ones you'll use most:Thread.print(thread dump),GC.heap_info,GC.class_histogram(object counts per class),GC.heap_dump <file>,VM.flags,VM.system_properties,VM.native_memory(with NMT on, Topic 14.2),Compiler.codecache, andJFR.start/JFR.dump/JFR.stop. You must run it as the same OS user as the target JVM, and in a container usually from inside the container. - 2
A thread dump (
jstack <pid>,jcmd <pid> Thread.print, or sendingkill -3/SIGQUIT, which prints it to the JVM's standard output) shows every thread's name, state and stack. It answers "why is it stuck?" and "why is CPU high?". Look for many threadsBLOCKEDon the same lock, threadsWAITINGon a pool or connection, the same stack repeated, and theFound one Java-level deadlocksection. Take three dumps a few seconds apart: a thread stuck in the same frame in all of them is the problem; one seen once is just passing through. - 3
A heap dump (
jcmd <pid> GC.heap_dump /tmp/heap.hprof,jmap -dump:live,format=b,file=heap.hprof <pid>, or automatically with-XX:+HeapDumpOnOutOfMemoryError) is an.hproffile with every object. Analyse it with Eclipse MAT or VisualVM: the dominator tree and retained size show which object keeps the most memory alive, and path to GC roots shows why it can't be collected (Topic 14.3). A dump pauses the JVM and can be as large as the heap, and it contains your data (passwords, personal data), so handle it like a production database backup. For a quick look,GC.class_histogramlists counts and bytes per class in seconds. - 4
Java Flight Recorder (open-sourced into OpenJDK 11; earlier a commercial Oracle feature) records events inside the JVM: GC pauses, allocations by stack trace, method samples (where CPU goes), lock contention, I/O, exceptions, class loading, compilation, safepoints, and your own custom events, with overhead typically around 1% in the default configuration. Start it with
-XX:StartFlightRecordingat launch orjcmd <pid> JFR.startlater; open the.jfrfile in JDK Mission Control (JMC) or use thejfrcommand-line tool. Many teams run it continuously with a ring buffer and dump the last few minutes when something goes wrong. - 5
Other tools: **
jpslists JVMs;jstat -gcutil <pid> 1000prints GC occupancy and counts every second;jinfoshows flags and properties;jconsole(in the JDK) and VisualVM (a separate download since JDK 9) are graphical monitors that connect locally or through JMX**;jhsdbinspects hung processes and core files. The same data is available in code throughjava.lang.management(ThreadMXBean,MemoryMXBean,GarbageCollectorMXBean), which is how monitoring agents and Micrometer/Prometheus exporters get JVM metrics. - 6
A good order of investigation: metrics first (CPU, heap after GC, GC pause time, thread count) to see *what* is wrong; then JFR to see *where* over a time window; then a thread dump for hangs and contention or a heap dump for memory. On Linux, combine
top -H -p <pid>(CPU per thread) with a thread dump: convert the hot thread's id to hex and find the matchingnid=0x...in the dump to see exactly which Java code burns the CPU.
Explain it without notes
A service is hung: requests time out but CPU is near zero. How do you investigate?
CPU is at 100% on a Java service. How do you find the code responsible?
What's the difference between a thread dump, a heap dump and a JFR recording?
Why should heap dumps be treated as sensitive data?
Practice
Write a program that names three threads worker-1 to worker-3, starts them, joins them, and prints the names of all threads whose names start with worker- while they are alive (use a latch so they're all alive at once).
Use ThreadMXBean to print whether thread contention monitoring is supported and the number of threads that have been started by the JVM since start-up is at least 1.
Write a program that checks the JVM has at least one GarbageCollectorMXBean and that every collection count is at least zero, printing only collectors found: true (the real names and counts vary between runs and collectors).
Trade-offs
- ↔
Heap dumps give complete answers but pause the JVM, are huge and contain sensitive data; histograms and JFR old-object samples are lighter first steps.
- ↔
Continuous JFR costs a little CPU and disk but means the data from the incident already exists when you need it.
- ↔
Shipping the full JDK in production images makes diagnosis easy but enlarges the image and attack surface; a slim JRE needs JFR/JMX enabled up front or a sidecar tool.
Done when you can
Done when you can list JVMs and run diagnostic commands with jcmd.
Done when you can read a thread dump: states, locks held and waited for, deadlock section.
Done when you can take a heap dump safely and find a leak with dominator tree and path to GC roots.
Done when you can start, dump and read a JFR recording.
Done when you can map a hot OS thread to its Java stack.