All articles
2026-05-02•24 min read•JVM · Performance

JVM internals, performance tuning, and diagnostics in Java 21

How the JVM actually runs Java 21 — bytecode, class loading, JIT, GC — plus an evidence-first method for tuning and diagnosing production incidents.

pC
Prashant Chaturvedi
Engineer

A Spring Boot service goes slow at 2 a.m., the heap graph climbs like it has somewhere to be, and somebody suggests -Xmx4g because it worked once on a different service. This article is the antidote to that conversation. It explains how the JVM actually executes your code — bytecode, class loading, memory areas, JIT compilation, garbage collection — and then turns that model into a diagnostic method: given a symptom, which subsystem is suspect, which evidence proves it, and which change is justified.

The running example is a government scheme eligibility engine: a Spring Boot service that loads citizen facts from registries, evaluates policy rules, persists auditable decisions, and publishes events. It is a useful specimen because it has real CPU work (policy evaluation), real allocation pressure (request payloads, decision objects), real blocking (registry calls), and real retention risk (evidence caches) — the four ingredients of almost every JVM incident.

You should be comfortable shipping a Spring Boot service; no prior JVM internals knowledge is assumed. Examples were run on OpenJDK 21 (HotSpot) on Linux; every command shown was executed and verified against that version. Where behavior is implementation-specific rather than specification-guaranteed, the article says so.

JVM, JDK, JRE, and HotSpot — the names people conflate

Four terms get used interchangeably in incident channels, which matters because they define what is guaranteed and what is not:

  • JVM (Java Virtual Machine) — the abstract machine defined by the Java Virtual Machine Specification. It specifies the class-file format, bytecode semantics, runtime data areas, and the load-link-initialize lifecycle.
  • HotSpot — the JVM implementation in OpenJDK and Oracle JDK. Tiered JIT compilation, metaspace, object layout, and the collector algorithms are all HotSpot details, not spec guarantees.
  • JDK (Java Development Kit) — the toolchain: javac, java, jcmd, jstack, jfr, javap.
  • JRE (Java Runtime Environment) — a legacy distribution name for “just the runtime.” Production images today are usually a JDK or a trimmed runtime built with jlink.

The distinction that pays rent all article long:

ConceptSpec-level guaranteeHotSpot implementation detail
HeapShared runtime area holding objects and arraysRegions or generations, depending on collector
Method areaShared storage for class-level structuresClass metadata in native-memory-backed metaspace
ExecutionBytecode instructions executeInterpreter plus tiered JIT compilation
Class loadingClasses are loaded, linked, initializedBootstrap, platform, application, and custom loaders
Garbage collectionUnreachable objects may be reclaimedG1, ZGC, Parallel, Serial, each with different algorithms

Everything in the right-hand column can change between JDK vendors and versions. Everything in the left-hand column cannot, which is why bytecode from 2008 still runs.

From source to machine code

Nothing executes your .java file. javac compiles it to bytecode in a .class file; at runtime the JVM loads the class, interprets the bytecode, profiles what is actually happening, and compiles the hot parts to native machine code.

Java source

javac compiler

.class bytecode

Class loading

Verify, link, initialize

Bytecode interpreter

Runtime profiling

JIT compilation

Optimized native code

CPU execution

Java source

javac compiler

.class bytecode

Class loading

Verify, link, initialize

Bytecode interpreter

Runtime profiling

JIT compilation

Optimized native code

CPU execution

Take the smallest piece of the scheme engine:

EligibilityRules.java
public final class EligibilityRules {
public boolean isIncomeEligible(long annualIncome, long threshold) {
return annualIncome <= threshold;
}
}

Compile and disassemble it:

terminal
javac --release 21 EligibilityRules.java
javap -c -v EligibilityRules.class

Here is the real output for the one method that matters:

javap -c EligibilityRules.class (excerpt)
public boolean isIncomeEligible(long, long);
Code:
0: lload_1
1: lload_3
2: lcmp
3: ifgt 10
6: iconst_1
7: goto 11
10: iconst_0
11: ireturn

Read it once and the mystery evaporates: load the first long argument, load the second, compare them, branch if greater, push a boolean, return. The verbose (-v) output adds the metadata layer — major version: 65 (Java 21’s class-file version), the constant pool, method descriptors like (JJ)Z, and a StackMapTable attribute.

What a class file contains

A class file carries everything the JVM needs to verify, link, and execute the class — not just instructions.

Class file

Magic number and version

Constant pool

Access flags

This class and superclass

Implemented interfaces

Field definitions

Method definitions and bytecode

Attributes: annotations, debug data, records, modules

Class file

Magic number and version

Constant pool

Access flags

This class and superclass

Implemented interfaces

Field definitions

Method definitions and bytecode

Attributes: annotations, debug data, records, modules

The constant pool deserves the attention. Bytecode does not contain literal field and method addresses; it contains indexes into a table of symbolic references — class names, method names, descriptors, strings. Resolution, the process of turning a symbolic reference into a live runtime pointer, happens later, during linking. That indirection is why classes can be compiled against an interface they have never met.

Try it on a record from the engine:

EligibilityRequest.java
public record EligibilityRequest(
String citizenId,
String schemeCode,
long annualHouseholdIncome
) {}

javap -c -v EligibilityRequest.class shows a normal class with generated accessors, equals, hashCode, toString, and a Record attribute naming the components. The syntax is concise; the execution model is not special — records ride the same class-loading and verification machinery as everything else.

The practical point of bytecode inspection is not memorizing opcodes. It is that when a library claims to do X, javap shows what was actually compiled — useful the day a Lombok annotation, an AOP proxy, or a Kotlin default argument produces bytecode that does not match the source you were reading.

Class loading, linking, and initialization

The JVM brings classes into execution lazily, through three stages:

Loading

Linking

Verification

Preparation

Resolution

Initialization

Ready

Unloading

Loading

Linking

Verification

Preparation

Resolution

Initialization

Ready

Unloading

  • Loading finds the bytes — from a JAR, a directory, a module image, generated bytecode, or a custom loader — and creates the JVM’s internal representation.
  • Linking makes the class runnable: verification checks structural and type-safety constraints in the bytecode; preparation allocates static fields and zeroes them; resolution converts constant-pool symbols into direct references, eagerly or on first use.
  • Initialization runs the static initializers — static field assignments and static {} blocks — compiled into a method the JVM calls <clinit>. It happens on first active use: creating an instance, calling a static method, touching a non-constant static field.

The loader hierarchy and parent delegation

LoaderLoadsWhat getClassLoader() returns
BootstrapCore java.* classesnull — by convention, not by absence
PlatformJDK platform modulesjdk.internal.loader.ClassLoaders$PlatformClassLoader
ApplicationYour classes and dependency JARsjdk.internal.loader.ClassLoaders$AppClassLoader
CustomPlugins, app-server modules, generated classesWhatever you wrote

Most loaders delegate to their parent first and only define a class themselves when the parent cannot. One consequence that bites in production: a class’s identity is its fully qualified name plus its defining loader. Two loaders can load com.foo.Plugin independently, producing two types that cannot be cast to each other — the root cause of ClassCastException: com.foo.Plugin cannot be cast to com.foo.Plugin and of class-loader leaks in plugin systems.

Run this against a real classpath and the hierarchy stops being abstract:

ClassLoaderDemo.java
public final class ClassLoaderDemo {
public static void main(String[] args) {
printLoader("Application class", ClassLoaderDemo.class);
printLoader("JDK class", String.class);
}
private static void printLoader(String label, Class<?> type) {
System.out.printf("%-24s %-45s loader=%s%n",
label, type.getName(), type.getClassLoader());
}
}
observed output
Application class ClassLoaderDemo loader=jdk.internal.loader.ClassLoaders$AppClassLoader@18b4aac2
JDK class java.lang.String loader=null

That null is the bootstrap loader — the class has a loader, you just cannot hold a reference to it.

Initialization order, verified

SchemeRules.java
public final class SchemeRules {
static final String POLICY_VERSION = loadPolicyVersion();
static {
System.out.println("SchemeRules initialized");
}
private static String loadPolicyVersion() {
System.out.println("Loading policy version");
return "2026-01";
}
public static void main(String[] args) {
System.out.println("POLICY_VERSION=" + POLICY_VERSION);
}
}
observed output
Loading policy version
SchemeRules initialized
POLICY_VERSION=2026-01

Static initializers run in source order — the field assignment before the block — inside a single <clinit> the JVM executes once, under a lock that makes initialization thread-safe.

Production note — keep network I/O, config lookups, and database calls out of static initializers. If <clinit> throws, the JVM marks the class erroneous and every subsequent use in that loader fails with NoClassDefFoundError: Could not initialize class … — an error that survives until restart and never carries the original stack trace again.

Where the JVM keeps data at runtime

The spec defines shared and per-thread data areas; the diagram maps them to the parts of a running service.

Class files and JARs

Class loader subsystem

Class metadata / metaspace

Java heap

Application thread

Program counter

JVM stack: frames

Native method stack

Execution engine

Garbage collector

Interpreter

JIT compiler

Code cache

JNI and native libraries

Class files and JARs

Class loader subsystem

Class metadata / metaspace

Java heap

Application thread

Program counter

JVM stack: frames

Native method stack

Execution engine

Garbage collector

Interpreter

JIT compiler

Code cache

JNI and native libraries

  • Heap — shared across threads; holds objects and arrays; the GC’s territory. Exhaustion produces OutOfMemoryError: Java heap space, which is only one of several OOM flavors.
  • JVM stack — one per thread, one frame per in-flight method call, holding local variables and the operand stack. Unbounded recursion gives you StackOverflowError; java.lang.OutOfMemoryError: unable to create native thread is what happens when you have too many threads — each stack costs native memory even when the heap is calm.
  • Program counter — per-thread pointer into the current bytecode instruction.
  • Native method stack — for JNI calls into native code.
  • Method area / metaspace — class metadata: runtime constant pools, field and method structures, statics. In HotSpot this lives in metaspace, which is native memory, not heap, and unbounded by default (MaxMetaspaceSize is effectively unlimited until you set it). Steady metaspace growth at runtime usually means generated proxies, dynamic languages, or a class-loader leak preventing unloading.
  • Code cache — where JIT-compiled native code lives. Pressure here is rare but real in workloads with heavy dynamic code generation.

Production note — learn the OOM taxonomy once: Java heap space, Metaspace, unable to create native thread, Requested array size exceeds VM limit, Direct buffer memory, and Compressed class space. Each points at a different area and a different fix; “increase -Xmx” answers only the first.

Object allocation and lifetime

Allocation in HotSpot is cheap but not free. Each thread gets a thread-local allocation buffer (TLAB) — a private slab of eden where bumping a pointer creates an object with no locking. Cross-thread coordination only happens when the TLAB refills. This is why Java allocates small objects faster than most people expect, and why “allocation is free” is still wrong: every byte allocated is a byte the GC eventually walks, copies, or reclaims.

Two implementation details deserve mental-model status rather than trivia status:

  • References are opaque. Compressed oops, object headers, and alignment are HotSpot internals. Do not compute memory footprints from first principles; measure them (jcmd … GC.class_histogram, JFR, or JOL in a test).
  • Escape analysis can delete allocations. When the JIT proves an object never escapes its method or thread, it can scalar-replace it — the object never exists at runtime. Source-level allocation counts and runtime allocation counts are different numbers, which is exactly why allocation profiling needs a profiler, not code review.

And the corollary that causes most “memory leaks” in Java:

EvidenceCache.java (illustrative — do not ship)
public final class EvidenceCache {
private static final Map<String, byte[]> CACHE = new ConcurrentHashMap<>();
public void cacheEvidence(String evaluationId, byte[] payload) {
CACHE.put(evaluationId, payload);
}
}

This is not a GC failure. The static map is a GC root, and every byte[] it holds is reachable by design. Garbage collection reclaims objects that are unreachable; it cannot rescue objects your code keeps on purpose. The fix is an eviction policy or a different owner, not a different collector.

Interpret, profile, compile: how JIT works

The JVM starts interpreting bytecode immediately — no compile step stands between you and a running service. While interpreting, it collects a profile: which methods run hot, which call sites are monomorphic, which branches are taken. Hot methods get compiled to native code, and the profile feeds the optimizer assumptions about your program’s actual behavior.

Bytecode

Interpret

Collect runtime profile

C1: fast compilation

C2: deeper optimization

Optimized machine code

Deoptimize if assumptions break

Bytecode

Interpret

Collect runtime profile

C1: fast compilation

C2: deeper optimization

Optimized machine code

Deoptimize if assumptions break

Tiered compilation is the short version: interpreter → quick C1 compilation with profiling → deeper C2 compilation for the truly hot code. The loop back to the interpreter — deoptimization — is the part that makes the system trustworthy: if a new class is loaded that breaks an “only one implementation exists” assumption, the JVM throws away the optimized code and resumes interpreted execution at the exact right point. Correctness is never traded for speed; the bet is just recalculated.

Warm-up is real. A method’s first thousand calls run interpreted or C1-compiled; its millionth runs fully optimized native code. Timing a method once from main measures startup, class loading, and interpreter speed — not steady-state performance.

What the JIT optimizes — and why naive benchmarks lie

OptimizationWhat it doesWhy you care
InliningPastes a small callee’s body at the call siteKills call overhead and feeds every other optimization
Dead-code eliminationDeletes work whose result is unobservableMakes careless benchmarks measure nothing
Escape analysisProves an object stays inside a scopeEnables scalar replacement and lock elision
Scalar replacementReplaces an object with loose fieldsRemoves the allocation entirely
Loop optimizationUnrolls, hoists, vectorizes loopsWhere the big throughput wins live
DevirtualizationCompiles a virtual call to a direct targetHot polymorphic paths get static-call speed
DeoptimizationBails out when an assumption breaksThe safety net that makes the rest possible

The classic trap:

MisleadingBenchmark.java
public final class MisleadingBenchmark {
public static void main(String[] args) {
long start = System.nanoTime();
for (int i = 0; i < 1_000_000; i++) {
String value = "scheme-" + i;
}
System.out.println(System.nanoTime() - start + " ns");
}
}

value is never read, so dead-code elimination can erase the loop body; the JIT may also compile mid-loop (on-stack replacement) and fold the concatenation. The number printed tells you about the benchmark, not the service. For isolated microbenchmarks use JMH, which exists precisely to defeat these optimizations; for service behavior use load tests, metrics, and JFR — nobody’s p99 ever improved because a for loop in main got faster.

How garbage collection decides what dies

GC reclaims objects that are unreachable from GC roots: live thread stacks, static fields, JNI references, and JVM-internal references.

GC roots

Live thread-stack references

Static fields

JNI references

JVM internal references

Request context

Static cache

Eligibility decision

Evidence payload

GC roots

Live thread-stack references

Static fields

JNI references

JVM internal references

Request context

Static cache

Eligibility decision

Evidence payload

That framing converts “memory leak” from a mystery into a question: which root holds a path to this object, and why? The EvidenceCache example is a static-field root. A ThreadLocal on a pooled thread is a thread-stack root. A listener nobody deregisters is a reference chain from some long-lived publisher.

Three more ideas complete the model:

  • Stop-the-world vs concurrent. Some GC phases pause all application threads; others run alongside them. Every collector pauses sometimes; they differ in how often, for how long, and what they charge you in throughput for shorter pauses.
  • The generational hypothesis. Most objects die young; a minority live long. Collectors that exploit this (G1, generational ZGC, Parallel, Serial all do) collect young objects cheaply and often, old objects thoroughly and rarely. It is a hypothesis about workloads, not a law — mostly true for request-response services.
  • Reference strengths. SoftReference, WeakReference, and PhantomReference weaken reachability for caches and cleanup coordination. They are niche tools — an unbounded ConcurrentHashMap with a TTL beats a soft-reference cache in every way that matters operationally — and they never substitute for an explicit eviction policy.

G1: the default collector

G1 is HotSpot’s default since JDK 9 and the right starting point for a Spring Boot service unless measurement says otherwise. It splits the heap into fixed-size regions (the JVM picked 4 MB regions on my machine; anywhere from 1 to 32 MB depending on heap size) rather than two monolithic generations, and it targets a pause-time goal while keeping throughput reasonable.

G1 heap

Region

Region

Region

Region

Region

Young regions

Old regions

Humongous regions

G1 heap

Region

Region

Region

Region

Region

Young regions

Old regions

Humongous regions

The mechanics that matter for diagnosis:

  • Young collections evacuate live objects out of young regions — stop-the-world, frequent, fast.
  • Concurrent marking walks the live-object graph while your threads run.
  • Mixed collections reclaim chosen old regions after marking — the “garbage-first” part: G1 collects the regions with the most reclaimable garbage per unit of work.
  • Remembered sets record cross-region references so a collection does not scan the whole heap.
  • Humongous objects — anything bigger than half a region — allocate straight into dedicated regions on a slower path and can trigger marking cycles early. Big byte[] payloads and oversized batch fetches are the usual suspects.

A baseline command for a lab environment:

terminal — learning baseline, not a production recommendation
mkdir -p logs # the JVM will not create this; startup fails without it
java \
-Xms1g \
-Xmx1g \
-XX:+UseG1GC \
-XX:MaxGCPauseMillis=200 \
-Xlog:gc*,safepoint:file=logs/gc.log:time,level,tags \
-jar scheme-engine.jar

Three honest notes about those flags. MaxGCPauseMillis=200 is already G1’s default — setting it explicitly just documents intent. The log directory must exist before launch; -Xlog opens the file but will not create missing parent directories. And -Xlog:gc* is the flag that should have been on before the incident — GC logging is cheap enough to leave on permanently with rotation. Reading that log is its own skill — Reading the G1 GC log, for real this time covers the three numbers to look at first.

Tuning order for G1: fix retention and gratuitous allocation first (application problems dominate), then size the heap for the measured live set plus headroom, then set a pause target only if you have a latency budget that demands it, changing one flag at a time against a repeatable load.

ZGC for low-pause collection

ZGC does nearly all of its work — marking, relocation, reference processing — concurrently, using colored pointers and load barriers so your threads never stop to let objects move. Its design goal is pause times measured in fractions of a millisecond, largely independent of heap size. Since JDK 21 (JEP 439) it is generational by default, which fixed the old ZGC’s appetite for allocation rate.

Evaluate ZGC when:

  • p99 or p999 latency has a hard budget and GC pauses are demonstrably eating it;
  • the heap or allocation rate makes G1’s pauses material even after application fixes and sane sizing;
  • you have CPU and memory headroom — concurrent collection is not free, it is just concurrent.

What it is not: a throughput upgrade. When you compare collectors, compare p50/p95/p99 latency, throughput under identical load, CPU utilization, allocation rate, and resident memory — not average GC pause, which is the one number both collectors make look good while hiding different sins.

Shenandoah is the other low-pause collector, available in some OpenJDK distributions; Serial and Parallel remain the right answers for tiny and batch workloads respectively. For anything deeper — collector selection, the four flags that matter, when to stop tuning — the JVM tuning series walks through it chapter by chapter.

Heap size is not process size: the JVM in containers

A container gets OOM-killed when process memory exceeds the limit, and the Java heap is only one line item in that budget:

Container memory limit

Java heap

Metaspace

Thread stacks

Direct / NIO buffers

Code cache

GC native structures

JNI and native libraries

Allocator overhead and other native memory

Container memory limit

Java heap

Metaspace

Thread stacks

Direct / NIO buffers

Code cache

GC native structures

JNI and native libraries

Allocator overhead and other native memory

The classic failure: the JVM respects the container limit (container awareness has been in place since the JDK 8u191 era), but -XX:MaxRAMPercentage defaults to 25% — so a pod with a 4 GiB limit gets a 1 GiB heap by default, which looks absurdly small until you remember the JVM is reserving the other 75% for everything else in that diagram, plus a margin. Setting MaxRAMPercentage to 75% or an explicit -Xmx is usually right for a dedicated service container — but never 100%, because metaspace, thread stacks, direct buffers, code cache, and GC internals all live outside -Xmx.

Sizing procedure:

  1. Start from the container limit.
  2. Reserve measured headroom for native categories (NMT, covered later, is how you measure).
  3. Set -Xmx — or MaxRAMPercentage — to the remainder.
  4. Load-test the exact image with the exact flags you will deploy.
  5. Watch resident set size against the limit, not just heap against -Xmx.

Production note — thread count is a memory line item. A few hundred platform threads at ~1 MB of stack each is hundreds of MB you will not find in the heap. Virtual threads — final since JDK 21, JEP 444 — move thread stacks into heap-managed memory and change that math dramatically, but they do not fix a downstream registry that answers in 30 seconds; blocking is still blocking.

A method for performance tuning

Tuning is an experiment, not a flag collection. The loop:

Yes

No

Observed symptom

Capture baseline

Form one hypothesis

Collect JFR, GC logs, metrics, dumps

Make one constrained change

Repeat equivalent workload

Compare latency, throughput, CPU, memory

Improved without regression?

Document and retain

Revert and revise hypothesis

Yes

No

Observed symptom

Capture baseline

Form one hypothesis

Collect JFR, GC logs, metrics, dumps

Make one constrained change

Repeat equivalent workload

Compare latency, throughput, CPU, memory

Improved without regression?

Document and retain

Revert and revise hypothesis

Start from the symptom, not the subsystem you happen to know:

SymptomLikely suspectsFirst evidence
High p99 latencyGC pauses, blocking I/O, lock contention, queuing, CPU saturationJFR, request metrics, GC log
High CPUHot loop, serialization, retry storm, GC overhead, lock spinningJFR CPU samples, thread dumps
Heap growthCache retention, queues, session data, listeners, class-loader leakClass histogram, heap dump
Container OOMKilledHeap plus native memory exceeding the limitContainer metrics, NMT, heap stats
Slow startupClass loading, component scanning, initialization, remote callsStartup metrics, JFR
Request timeoutsPool exhaustion, downstream latency, connection-pool contentionThread dumps, executor metrics, traces

The discipline that separates this from superstition: fix application behavior before JVM flags; define the workload and the objective before touching anything; baseline first; one variable per experiment; keep the evidence and the exact command line with each result; and check that the bottleneck did not just move to the next subsystem.

Java Flight Recorder: the black box

JFR records low-overhead JVM events — method sampling, allocation, GC, locks, threads, socket and file I/O, class loading — into a file you open in Java Mission Control or feed to jfr print. It is the single best first instrument for “the service is slow and I don’t know why.”

Capture an incident window without restarting:

terminal
jcmd <pid> JFR.start \
name=scheme-engine-incident \
settings=profile \
duration=5m \
filename=recordings/scheme-engine-incident.jfr

(Verified on JDK 21 — jcmd -l lists local JVMs to find <pid>, and JFR.check confirms the recording is running.) For always-on coverage, start the JVM with -XX:StartFlightRecording=filename=...,settings=profile instead.

The views that answer questions:

  • Method profiling — where CPU time actually goes.
  • Allocation — which call sites generate the pressure; the fastest path to “why is GC busy.”
  • Garbage collection — pause durations, frequency, heap occupancy around collections.
  • Threads and locks — parked, blocked, and waiting states; contended monitors.
  • Socket and file I/O — the slow read() hiding behind “the endpoint is slow.”
  • Class loading — unexpected dynamic generation or reload churn.

Security note — JFR recordings, heap dumps, thread dumps, and GC logs can contain request payloads, citizen identifiers, SQL fragments, and environment values. In a system that touches citizen data, treat all four as sensitive operational artifacts: restrict access, set retention, and do not attach them to tickets that fly to vendors.

Reading thread dumps

A thread dump is a snapshot of every thread’s state and stack. It answers the questions latency graphs cannot: what are request threads waiting on?

terminal
jcmd <pid> Thread.print -l > dumps/threads-$(date +%s).txt

Take three or four, thirty seconds apart — one dump is a photograph, a sequence is a time series, and a thread BLOCKED on the same monitor in four consecutive dumps is a finding, not a coincidence.

StateWhat it means in practice
RUNNABLEExecuting or ready to execute — could be real work, could be spinning, could be a native read()
BLOCKEDQueued on an intrinsic (synchronized) monitor — check the waiting to lock address across dumps
WAITINGParked indefinitely — queue take, Object.wait, LockSupport.park
TIMED_WAITINGWaiting with a deadline — sleep, poll(timeout), parkNanos
TERMINATEDFinished

The canonical scheme-engine incident:

RegistryLookup.java (illustrative)
ExecutorService registryPool = Executors.newFixedThreadPool(20);
CompletableFuture<RegistryResponse> lookup(String citizenId) {
return CompletableFuture.supplyAsync(
() -> slowRegistryClient.fetch(citizenId), registryPool);
}

The registry slows to 30-second responses, all 20 pool threads park inside fetch, the request queue behind supplyAsync grows without bound, and p99 goes vertical while CPU sits near zero — a signature that is unmistakable in a thread dump and invisible in a CPU graph. Doubling the pool buys minutes; the fix is timeouts on the registry call, a bounded queue with backpressure, a bulkhead so one slow dependency cannot starve the service, and retries with a budget.

Production note — Thread.print does not show virtual threads. For services on JDK 21 virtual threads, use jcmd <pid> Thread.dump_to_file threads.txt (add -format=json for tooling), which dumps platform and virtual threads together.

Reading heap dumps

Reach for a heap dump on a real signal: an OutOfMemoryError, sustained heap growth across traffic cycles, or a leak hypothesis you need the object graph to confirm. Dumps are large and pause the VM while written — deliberate, not reflexive.

terminal
jcmd -l # find the JVM
jcmd <pid> VM.flags # record the exact flag set
jcmd <pid> GC.class_histogram # cheap first look: instance counts by class
jcmd <pid> GC.heap_dump dumps/scheme-engine.hprof

For automatic capture on the fatal path:

JVM flags
-XX:+HeapDumpOnOutOfMemoryError
-XX:HeapDumpPath=/var/log/scheme-engine/dumps

The analysis workflow, whether you open the dump in Eclipse MAT, VisualVM, or JMC:

OOM or sustained heap growth

Capture heap dump

Class histogram

Dominator tree

Largest retained sets

Paths to GC roots

Owning cache, queue, listener, static

Fix ownership or lifecycle

Repeat workload, compare

OOM or sustained heap growth

Capture heap dump

Class histogram

Dominator tree

Largest retained sets

Paths to GC roots

Owning cache, queue, listener, static

Fix ownership or lifecycle

Repeat workload, compare

The vocabulary that does the work: shallow size is what an object occupies itself; retained size is what becomes collectable if that object disappears; the dominator tree ranks objects by retained size; paths to GC roots explains why the object is still alive. The goal is never “find the biggest class” — it is “find the reference chain that should not exist.” In the scheme engine’s case the histogram would show byte[] dominating, the dominator tree would point at EvidenceCache.CACHE, and the GC-root path would say static field — three tools converging on the bug from the retention section.

Usual suspects, in rough order of how often they are actually guilty: unbounded maps and caches, queues outpacing consumers, request-context objects outliving requests, listeners never deregistered, ThreadLocal values riding pooled threads, oversized log or retry buffers, class-loader leaks in plugin/reload environments, and ORM sessions held past their transaction.

Native memory and NMT

When RSS climbs while heap graphs stay flat, the growth is native — and now you have the map to look for it: metaspace, thread stacks, code cache, direct buffers, GC internals, JNI allocations, allocator overhead.

Native Memory Tracking is opt-in at JVM start:

JVM flag
-XX:NativeMemoryTracking=summary

detail instead of summary tracks call-site-level allocation at a higher cost — enable it for a diagnosis window, not permanently. Once running:

terminal
jcmd <pid> VM.native_memory summary # categories with reserved/committed
jcmd <pid> VM.native_memory baseline # snapshot after warm-up
jcmd <pid> VM.native_memory summary.diff # growth since the baseline

The workflow that works: baseline after warm-up, run the suspect workload, diff, and the category that grew names the subsystem. Class growing means metaspace — chase loaders. Thread growing means a leak in ExecutorService creation. Internal or Other growing under a Netty service usually means direct buffers. NMT does not see all native memory — JNI libraries and some mmap’d regions are invisible to it — so a fully-explained NMT report with unexplained RSS growth means the leak is below the JVM, in native code.

The incident playbook

Each of these is the tuning loop aimed at one symptom.

High GC pause time. Confirm the user-visible symptom first. Capture request metrics, GC log, a short JFR recording. Verify pauses correlate with latency spikes — they often do not. Measure allocation rate and post-GC occupancy. Fix application allocation and retention before heap sizing; resize before collector flags; switch collectors last. Re-run the same load.

Suspected memory leak. Confirm growth across comparable traffic cycles, not a one-time climb to steady state. Histogram first if a dump is too expensive, dump at high occupancy, dominator tree, path to GC roots, fix the owner, verify with a second capture under the same load.

CPU saturation. Confirm saturation at host/container level, not just “service slow.” JFR plus several thread dumps. Identify whether the CPU is application compute, GC work, serialization, compression, retry loops, or lock spinning — five different fixes wearing one costume.

Thread-pool starvation. Executor active counts, queue depth, connection-pool metrics, downstream latency; multiple thread dumps to find the common stack; then timeouts, bulkheads, bounded queues, backpressure — not bigger pools.

Container OOMKilled. First separate the two killers: a JVM OutOfMemoryError in the logs versus the container runtime killing the process — different causes, different fixes. For the container kill, compare RSS, heap, thread count, direct buffers, and NMT against the limit, reserve headroom beyond -Xmx, and verify under realistic concurrency.

A lab plan that builds the model

Everything in this article is verifiable on a small Spring Boot service with one endpoint — POST /api/v1/schemes/{schemeCode}/eligibility-evaluations — plus a fake registry client you can make slow on demand.

  1. Bytecode: javap -c -v an eligibility-rule class; find the constant pool and the StackMapTable.
  2. Class loading: run ClassLoaderDemo; then read SchemeRules.POLICY_VERSION and predict the print order before running it.
  3. Stack: recurse without a base case; observe StackOverflowError; rerun with -Xss256k and watch the depth change.
  4. Retention: deploy EvidenceCache unbounded; watch old-generation growth across collections.
  5. GC log: drive allocation pressure, read the log, compute the allocation rate from eden fill intervals.
  6. JFR: record under load; find the top allocation site; compare it to what the code review predicted.
  7. Thread dump: make the registry sleep 30 s; capture three dumps; count threads parked in fetch.
  8. Heap dump: histogram, dominator tree, GC-root path; name the retaining owner.
  9. NMT: baseline, load, diff; reconcile the growth category with what you deployed.
  10. Container: run under a memory limit, pick -Xmx with native headroom, load it, watch RSS.

Bytecode

Class loading

Memory areas

JIT and warm-up

GC logging

JFR analysis

Thread dumps

Heap dumps

Native memory

Container sizing

Bytecode

Class loading

Memory areas

JIT and warm-up

GC logging

JFR analysis

Thread dumps

Heap dumps

Native memory

Container sizing

Each lab produces one artifact — a javap output, a gc.log, a .jfr, a .hprof — that you keep next to the flag set that produced it. That collection of evidence is the real deliverable; the mental model is what lets you produce it again during an incident.

Final checklist

Before changing a JVM option in production:

  • Is there a measured symptom with a stated latency, throughput, error-rate, or memory impact?
  • Is the workload representative and repeatable?
  • Do you have the evidence for this symptom class — metrics, GC log, JFR, dumps?
  • Is the problem heap, native memory, CPU, contention, I/O, or a downstream dependency?
  • Has application behavior been ruled out before collector or flag changes?
  • Is there container headroom beyond -Xmx for metaspace, threads, and direct buffers?
  • Are dumps and recordings handled as sensitive artifacts?
  • Was exactly one meaningful change made since the last baseline?
  • Was the result verified under equivalent load?
  • Are the flags, evidence, decision, and trade-off written down?

Where to go next

The JVM is a managed execution environment, not a performance lottery. The model here — bytecode in, profiled compilation, reachability-driven collection, native memory outside the heap — is enough to classify nearly every production symptom you’ll meet, and the evidence tools turn classification into diagnosis in minutes instead of days.

For the scheme engine specifically, that means keeping the policy code boring and testable, and reaching for the diagnostics when operations get interesting: retained evidence payloads, a registry that answers late, a container that dies at 3 a.m. with a flat heap graph. To go deeper on any single stop in this tour: Reading the G1 GC log for the evidence format that matters most, and the JVM tuning series for heap sizing, collector selection, and the discipline of knowing when to stop.

JVMPerformance

Related Articles

Type to search the site.

↑↓ navigate⏎ openPowered by Pagefind