Production Java problems rarely announce themselves with a simple error message.
Sometimes CPU suddenly reaches 100%. Sometimes memory keeps increasing for days. Sometimes the average API latency looks perfectly healthy while a small percentage of users experience several seconds of delay.
Sometimes there are no errors in the logs at all.
This is where production troubleshooting becomes different from normal application development.
A senior Java engineer needs to understand not only Java code, but also the JVM, garbage collection, threads, databases, networking, connection pools, application behavior, deployment changes, and observability.
This article covers 20 real-world Java production scenarios and explains how an experienced engineer would investigate them.
I would not immediately restart the application.
My first objective would be to determine which process and which threads are consuming the CPU.
First, I would confirm whether the CPU increase is:
Then I would identify the CPU-consuming threads.
Application CPU = 100%
|
v
Which process?
|
v
Which thread?
|
v
Which Java method?
|
v
Why is it consuming CPU?
On Linux, tools such as top or pidstat can help identify the process and busy threads. A thread dump can then help map the Java thread to its stack trace.
Profiling tools or Java Flight Recorder can provide deeper information about hot methods.
For example, a thread dump may show:
Thread-42
at com.example.OrderService.calculate(...)
at com.example.OrderController.process(...)
Profiling might then reveal that calculate() is responsible for a large percentage of CPU usage.
I would also compare the current release with the previous release.
Senior-level answer: My first step is not to increase CPU or restart the server. I would identify the CPU-consuming process and threads, correlate them with application stack traces and profiling data, and then determine whether the cause is code, traffic, GC, or infrastructure.
Increasing memory usage does not automatically mean there is a memory leak.
Java may legitimately use available heap memory and later reclaim it through garbage collection.
I would first determine whether memory is actually being retained after garbage collection.
Heap Usage
/\
/ \ /\
/ \ / \
__/ \___/ \____
Healthy pattern:
Memory increases → GC → memory decreases
Potential leak:
Memory increases → GC → only small reduction
Memory increases → GC → only small reduction
Memory continuously grows
I would examine:
A heap dump can show which classes are consuming memory and, more importantly, why those objects remain reachable.
For example:
static Map<String, Object> cache = new HashMap<>();
Every request
|
v
Add entry
|
v
No eviction
|
v
Memory continuously increases
I would also check whether the application is simply experiencing higher traffic or larger workloads.
Senior-level answer: I would not call it a memory leak merely because heap usage is increasing. I would compare pre- and post-GC memory usage and use heap analysis to determine which objects are being retained and why.
This is a classic tail-latency problem.
Suppose:
Average = 180 ms
p50 = 150 ms
p95 = 300 ms
p99 = 5 seconds
The average looks healthy, but one percent of requests are extremely slow.
If the API receives one million requests per day, one percent means:
1,000,000 × 1%
= 10,000 slow requests
So p99 problems can represent a significant user impact.
I would use distributed tracing and latency histograms rather than relying only on averages.
I would also compare slow requests with normal requests:
Normal request:
API → DB → Response
200 ms
Slow request:
API → wait for connection
→ DB query
→ downstream timeout
→ retry
→ Response
5 sec
Senior-level answer: Average latency can hide serious production problems. I would investigate the tail using p95, p99, traces, dependency latency, queueing, connection pools, GC, and lock contention.
A large number of WAITING threads is not automatically a problem.
Some applications legitimately have many threads waiting for work.
The important question is:
What are those threads waiting for?
I would inspect the thread dump and categorize the states:
WAITING
TIMED_WAITING
BLOCKED
RUNNABLE
Then I would examine the stack traces.
For example:
200 threads
150 → waiting for database
30 → waiting for HTTP
15 → waiting for locks
5 → processing
The problem is probably not simply "too many threads".
The application has a dependency problem that is causing threads to remain occupied.
I would also correlate thread states with:
Senior-level answer: I would identify what the WAITING threads are waiting on rather than blindly increasing the thread pool size.
Frequent Full GC activity is a serious JVM performance signal.
I would first determine:
The key question is whether the JVM is unable to reclaim enough memory.
Before GC:
Heap = 90%
After GC:
Heap = 88%
Then:
Heap = 92%
Heap = 95%
Heap = 98%
Full GC
Heap = 87%
This pattern could indicate significant long-lived object retention or excessive allocation pressure.
I would use GC logs, JVM metrics, Java Flight Recorder, heap dumps, and allocation profiling.
I would avoid immediately increasing the heap.
If the application has a memory leak, a larger heap may simply delay the failure.
Senior-level answer: First determine whether Full GC is caused by allocation pressure, insufficient heap, or retained objects. Then use GC and heap evidence to identify the root cause rather than treating heap size as the universal solution.
I would not assume that the database query itself is the only factor.
The database console might report:
Query execution = 40 ms
while the application reports:
Database operation = 800 ms
I would investigate the entire application-to-database path.
For example:
Total application time = 800 ms
Connection wait = 300 ms
SQL execution = 50 ms
Result transfer = 150 ms
Hibernate mapping = 200 ms
Additional queries = 100 ms
The SQL query itself is fast.
The overall database interaction is not.
Senior-level answer: I would compare database execution time with end-to-end application time and break the operation into connection acquisition, SQL execution, result transfer, ORM mapping, and additional queries.
This pattern strongly suggests a problem that accumulates over time.
I would investigate:
A restart resets many runtime resources:
Restart
|
+-- Clears heap
+-- Clears thread state
+-- Clears in-memory caches
+-- Resets connection pools
+-- Clears queues
+-- Resets application state
If performance is good immediately afterward and gradually deteriorates, I would compare resource metrics over time.
For example:
Day 1:
Heap after GC = 35%
Day 3:
Heap after GC = 50%
Day 5:
Heap after GC = 70%
Day 7:
Heap after GC = 85%
This strongly suggests retained memory or another accumulating resource.
I would also investigate thread count and database connection behavior.
Senior-level answer: The restart itself is an important diagnostic clue. I would look for resources that accumulate over time rather than focusing only on the immediate symptom.
Low CPU with stuck requests strongly suggests the application is waiting for something.
Possible exhausted resources include:
For example:
CPU = 15%
HikariCP:
Active = 50
Idle = 0
Pending = 300
Result:
Requests wait for database connections
Another possibility is thread pool exhaustion:
200 worker threads
200 → waiting for slow HTTP service
New requests
|
v
Queue
|
v
High latency
I would use thread dumps, pool metrics, distributed traces, and operating-system resource metrics.
Senior-level answer: Low CPU tells me that computation may not be the bottleneck. I would look for exhausted concurrency, connection, locking, or I/O resources.
Thousands of threads are usually a strong signal that something in the concurrency design has gone wrong.
I would investigate:
A common mistake is creating an executor repeatedly:
for (...) {
ExecutorService executor =
Executors.newFixedThreadPool(10);
}
If those executors are not properly managed, the application can create an excessive number of threads.
Another problem is creating one thread for every incoming request.
I would inspect thread names in a thread dump.
pool-101-thread-1
pool-102-thread-1
pool-103-thread-1
...
This can reveal that hundreds of executors have been created.
I would also check whether the application is experiencing thread leaks.
Senior-level answer: I would identify who is creating the threads, whether they belong to expected executors, whether those executors are bounded and reused, and whether blocking operations are causing thread accumulation.
This is an interesting production scenario because the obvious JVM resource metrics look healthy.
I would compare the old and new versions and inspect the complete request path.
Possible causes include:
For example:
Before deployment:
API = 200 ms
After deployment:
API = 1,000 ms
CPU = normal
Memory = normal
Trace:
Database = 600 ms
Downstream = 250 ms
Application = 150 ms
The JVM looks healthy because the bottleneck is external to CPU and memory.
I would also compare database query counts before and after the deployment.
Senior-level answer: Healthy CPU and memory do not prove that the application is healthy from a latency perspective. I would use distributed traces and dependency metrics to identify what changed in the request path.
This situation strongly suggests an instance-specific problem.
I would compare the unhealthy instance against healthy instances.
Instance A → 150 ms
Instance B → 170 ms
Instance C → 5,000 ms
Instance D → 160 ms
I would compare:
It could be that one instance:
I would compare request traces routed specifically through that instance.
If removing the instance from service immediately restores overall performance, I would preserve the evidence and then investigate it separately.
Senior-level answer: I would isolate the unhealthy instance, compare its runtime and infrastructure metrics against healthy peers, and use instance-specific traces and logs to determine what makes it different.
The absence of errors in application logs does not mean that there are no failures.
The failure may occur:
I would first determine exactly what "failure" means.
Failure type?
HTTP 5xx?
HTTP 4xx?
Timeout?
Connection reset?
Slow response?
Incorrect response?
Client error?
I would use request IDs or trace IDs to follow individual failing requests.
I would also examine:
Intermittent failures often indicate:
Senior-level answer: I would not rely exclusively on application error logs. I would trace failed requests across the complete request path and determine whether the failure is application-level, infrastructure-level, dependency-level, or client-level.
This is a classic example of a timing-sensitive problem.
One possibility is a race condition.
Adding logging or debugging changes execution timing:
Without logging:
Thread A ────────────────┐
Thread B ────────────┐ |
| |
v v
Race
With logging:
Thread A ────────wait───────
Thread B ────────────wait───
Race disappears
Logging changes thread scheduling and timing, so the race may become much harder to reproduce.
Other possibilities include:
I would avoid concluding immediately that the problem is caused by logging.
Instead, I would capture non-intrusive evidence using metrics, thread dumps, tracing, profiling, and carefully selected logging.
I would specifically investigate shared mutable state and synchronization.
Senior-level answer: If observability changes the behavior, I would suspect timing-sensitive problems such as race conditions or synchronization issues and use low-impact diagnostic techniques to preserve the original behavior.
If network infrastructure looks healthy, I would investigate timeout sources above the network layer.
Possible causes include:
For example:
Network latency = 20 ms
But:
Thread waits = 1,500 ms
HikariCP waits = 800 ms
Database query = 500 ms
Total > timeout
The network is healthy, but the request still times out.
I would compare:
Connection timeout
Read timeout
Application timeout
Gateway timeout
Load balancer timeout
Database timeout
Timeouts can occur at multiple layers, and the shortest timeout may determine the actual behavior.
Senior-level answer: A network timeout does not necessarily mean a network problem. I would trace the request and measure queueing, connection acquisition, JVM pauses, database latency, and downstream processing time.
I would try to reproduce the production workload under controlled conditions.
I would first identify the relationship between traffic and failure.
100 req/sec → healthy
500 req/sec → healthy
1,000 req/sec → latency increases
2,000 req/sec → failures
3,000 req/sec → service crashes
This suggests a capacity or concurrency-related problem.
I would perform controlled load testing while monitoring:
Potential causes include:
I would gradually increase load rather than immediately generating maximum traffic.
This helps identify the point at which the system begins to degrade.
Senior-level answer: I would reproduce the traffic pattern under controlled load, monitor resource and dependency saturation, identify the first signal that changes before the crash, and use that evidence to locate the bottleneck.
Millions of objects do not automatically mean a memory leak.
The key question is:
Should these objects still be reachable?
I would investigate the reference chain to the objects.
Large number of objects
|
v
Who references them?
|
v
Why are they still reachable?
|
v
Should they have been released?
For example:
Application
|
v
Static Cache
|
v
HashMap
|
v
Millions of Customer objects
If the cache has no eviction policy, the objects may remain reachable indefinitely.
I would also compare heap dumps over time.
Heap Dump 1 → 100,000 objects
Heap Dump 2 → 500,000 objects
Heap Dump 3 → 2,000,000 objects
If the count continuously increases under a stable workload, that is strong evidence of a leak or unbounded retention.
I would use heap-analysis tools to inspect:
Senior-level answer: I would not identify a leak based only on object count. I would determine why those objects remain reachable and compare retention patterns across heap dumps.
If multiple threads are waiting for the same lock, that lock may be a serialization point.
Thread A ──┐
Thread B ──┤
Thread C ──┼──→ Lock
Thread D ──┤
Thread E ──┘
|
v
One thread
executes
I would identify:
For example:
synchronized
void process() {
updateMemory();
callExternalService();
writeToDatabase();
performCalculation();
}
If the synchronized method calls an external service, the lock may be held for hundreds of milliseconds or several seconds.
That can create severe contention.
I would consider reducing the critical section:
Synchronize only the shared-state operation.
Do not hold the lock while:
- Calling external services
- Performing slow I/O
- Running expensive calculations
Depending on the design, I might also replace coarse-grained synchronization with more appropriate concurrency primitives or redesign the shared state.
Senior-level answer: I would identify the lock owner, measure lock hold time, inspect the protected code, and determine whether the critical section can be reduced or redesigned.
This usually indicates significant allocation or memory pressure.
I would compare:
Application CPU
Allocation rate
GC CPU
GC frequency
GC pause time
Heap occupancy
Object promotion
A sudden increase in allocation rate may be caused by a new feature or workload.
For example:
Before deployment:
Allocation = 500 MB/sec
After deployment:
Allocation = 4 GB/sec
The application may now be generating large numbers of temporary objects.
Possible causes include:
I would use Java Flight Recorder, allocation profiling, GC logs, and heap analysis.
I would avoid immediately increasing heap size because the underlying problem may be excessive allocation rather than insufficient memory.
Senior-level answer: I would identify whether the problem is high allocation, poor object retention, or insufficient heap and then trace the allocation back to the responsible code path.
The feature flag provides a very useful comparison point.
Feature OFF
|
v
Latency = 200 ms
Feature ON
|
v
Latency = 1.5 sec
I would compare the same workload with the feature enabled and disabled.
I would examine:
Suppose the feature introduces an additional database lookup:
Before:
API → Database
200 ms
After:
API → Database
|
+→ Feature query
|
+→ Another database query
|
1,500 ms
The feature itself may be functionally correct but introduce an unexpected performance cost.
I would also check whether the feature creates additional calls only for certain users or request types.
Once the bottleneck is confirmed, the feature can be optimized or temporarily disabled.
Senior-level answer: Use the feature flag as an experiment boundary. Compare the same workload with the feature on and off, inspect traces and resource metrics, identify the additional cost, and then optimize or roll back the feature.
During an incident, I would avoid immediately making random configuration changes.
My first five checks would be:
I would check health endpoints, error rate, request rate, and availability from the user's perspective.
Requests
Errors
Latency
Availability
I would check recent:
I would check:
I would check:
I would identify a failing request and trace it through the system.
Client
|
v
Load Balancer
|
v
Java Service
|
+---- Database
|
+---- Redis
|
+---- Kafka
|
+---- Downstream Service
The objective during the first few minutes is to answer:
What is broken, when did it start, what changed, and where is the request failing?
If the service is severely impacting users, I would prioritize mitigation.
Depending on the evidence, mitigation might include:
Root-cause analysis can continue after the immediate user impact has been reduced.
Senior-level answer: My first priority is to establish impact and stabilize the system. Then I use metrics, logs, traces, JVM data, and recent changes to identify the root cause.
These questions may appear to cover different problems, but they test a few fundamental skills expected from senior Java engineers.
The common theme is:
Production Problem
|
v
Measure
|
v
Collect Evidence
|
v
Identify Bottleneck
|
v
Confirm Root Cause
|
v
Apply Targeted Fix
|
v
Verify
|
v
Prevent Recurrence
A restart can restore service temporarily, but it may destroy valuable evidence about the underlying problem.
Adding CPU, memory, threads, or connections may hide the symptom without fixing the cause.
Some problems occur in the database, network, load balancer, Kubernetes platform, or downstream services.
p95 and p99 are often much more useful for understanding real user experience.
The application may be blocked on I/O, locks, connection pools, or external dependencies.
Java's garbage collector and heap behavior must be understood before concluding that memory is leaking.
If you change five configurations simultaneously, it becomes difficult to determine which change actually fixed the problem.
When faced with an unfamiliar production problem, a useful framework is:
1. Confirm the symptom
↓
2. Measure the impact
↓
3. Establish a baseline
↓
4. Check recent changes
↓
5. Examine metrics
↓
6. Examine logs
↓
7. Trace affected requests
↓
8. Profile if necessary
↓
9. Identify the bottleneck
↓
10. Apply the smallest effective fix
↓
11. Verify recovery
↓
12. Prevent recurrence
This approach is useful because it prevents troubleshooting from becoming guesswork.
Yes. These scenarios focus on JVM internals, concurrency, performance, observability, databases, networking, and production incident troubleshooting. They are particularly useful for senior Java, Spring Boot, backend, and microservices interviews.
The most important skill is evidence-driven diagnosis. Instead of immediately changing code or configuration, identify the symptom, measure it, trace the request, find the bottleneck, and confirm the root cause.
A restart may be an appropriate mitigation when users are severely affected, but it should not replace investigation. Restarting can remove valuable diagnostic evidence and may only temporarily hide memory leaks, thread leaks, connection leaks, or other accumulating problems.
Average latency can hide a small percentage of extremely slow requests. In high-volume systems, even one percent of slow requests can affect thousands of users.
Thread dumps are extremely useful for identifying blocked threads, deadlocks, lock contention, thread pool exhaustion, and unexpected waiting behavior.
First establish the impact and stabilize the system if necessary. Then determine when the problem started, what changed, which components are affected, and where requests are failing.
Java production troubleshooting is not about memorizing a list of JVM commands.
It is about understanding how the application behaves under real-world conditions.
A senior engineer should be able to move from:
"The application is slow."
to:
"The p99 latency increased from 400 ms to 5 seconds."
Then:
"The JVM CPU is normal, but 70% of requests are waiting for database connections."
Then:
"The connection pool is exhausted because a new query introduced by the latest deployment holds connections for several seconds."
Then:
"After rolling back the deployment, latency returned to normal."
That is the difference between guessing and
troubleshooting.
Measure first. Follow the evidence. Find the bottleneck. Fix the root cause.
That mindset is one of the most valuable skills a senior Java engineer can bring to a production environment.
0 Comments