Java Production Problems – 20 Real-World Scenarios with Deep Answers


Production Java problems rarely announce themselves with a simple error message.

Sometimes CPU suddenly reaches 100%. Sometimes memory keeps increasing for days. Sometimes the average API latency looks perfectly healthy while a small percentage of users experience several seconds of delay.

Sometimes there are no errors in the logs at all.

This is where production troubleshooting becomes different from normal application development.

A senior Java engineer needs to understand not only Java code, but also the JVM, garbage collection, threads, databases, networking, connection pools, application behavior, deployment changes, and observability.

This article covers 20 real-world Java production scenarios and explains how an experienced engineer would investigate them.

1. Your Java application suddenly uses 100% CPU. What is the first thing you check?

I would not immediately restart the application.

My first objective would be to determine which process and which threads are consuming the CPU.

First, I would confirm whether the CPU increase is:

  • One application instance or all instances
  • One CPU core or all cores
  • Sudden or gradual
  • Associated with increased traffic
  • Associated with a recent deployment

Then I would identify the CPU-consuming threads.

Application CPU = 100%

        |
        v

Which process?
        |
        v

Which thread?
        |
        v

Which Java method?
        |
        v

Why is it consuming CPU?

On Linux, tools such as top or pidstat can help identify the process and busy threads. A thread dump can then help map the Java thread to its stack trace.

Profiling tools or Java Flight Recorder can provide deeper information about hot methods.

Possible causes

  • Infinite or extremely expensive loops
  • Unexpected traffic increase
  • Expensive serialization/deserialization
  • Excessive object creation
  • Regular-expression processing
  • High garbage collection activity
  • CPU-intensive business logic
  • Accidental polling loops
  • Logging or formatting overhead
  • New code introduced in a deployment

For example, a thread dump may show:

Thread-42
    at com.example.OrderService.calculate(...)
    at com.example.OrderController.process(...)

Profiling might then reveal that calculate() is responsible for a large percentage of CPU usage.

I would also compare the current release with the previous release.

Senior-level answer: My first step is not to increase CPU or restart the server. I would identify the CPU-consuming process and threads, correlate them with application stack traces and profiling data, and then determine whether the cause is code, traffic, GC, or infrastructure.

2. Memory usage keeps increasing but there is no OutOfMemoryError yet. What would you investigate?

Increasing memory usage does not automatically mean there is a memory leak.

Java may legitimately use available heap memory and later reclaim it through garbage collection.

I would first determine whether memory is actually being retained after garbage collection.

Heap Usage

     /\
    /  \       /\
   /    \     /  \
__/      \___/    \____

Healthy pattern:
Memory increases → GC → memory decreases


Potential leak:

Memory increases → GC → only small reduction
Memory increases → GC → only small reduction
Memory continuously grows

I would examine:

  • Heap usage
  • Old-generation usage
  • GC frequency
  • GC effectiveness
  • Allocation rate
  • Object counts
  • Heap dumps
  • Container memory limits

A heap dump can show which classes are consuming memory and, more importantly, why those objects remain reachable.

Common causes

  • Unbounded collections
  • Incorrect caching
  • Static collections
  • Listeners not removed
  • ThreadLocal misuse
  • Large objects retained by long-lived objects
  • Unbounded queues
  • Improper session storage
  • Large buffers

For example:

static Map<String, Object> cache = new HashMap<>();

Every request
      |
      v
Add entry
      |
      v
No eviction
      |
      v
Memory continuously increases

I would also check whether the application is simply experiencing higher traffic or larger workloads.

Senior-level answer: I would not call it a memory leak merely because heap usage is increasing. I would compare pre- and post-GC memory usage and use heap analysis to determine which objects are being retained and why.

3. An API has normal average latency but terrible p99 latency. What could be happening?

This is a classic tail-latency problem.

Suppose:

Average = 180 ms
p50     = 150 ms
p95     = 300 ms
p99     = 5 seconds

The average looks healthy, but one percent of requests are extremely slow.

If the API receives one million requests per day, one percent means:

1,000,000 × 1%
= 10,000 slow requests

So p99 problems can represent a significant user impact.

Possible causes

  • Occasional slow database queries
  • Connection pool waits
  • Thread pool exhaustion
  • GC pauses
  • Slow downstream services
  • Network latency
  • Cache misses
  • Lock contention
  • Large requests
  • Cold starts or scaling events

I would use distributed tracing and latency histograms rather than relying only on averages.

I would also compare slow requests with normal requests:

Normal request:
API → DB → Response
200 ms

Slow request:
API → wait for connection
    → DB query
    → downstream timeout
    → retry
    → Response
5 sec

Senior-level answer: Average latency can hide serious production problems. I would investigate the tail using p95, p99, traces, dependency latency, queueing, connection pools, GC, and lock contention.

4. A thread dump shows hundreds of WAITING threads. What would you look for?

A large number of WAITING threads is not automatically a problem.

Some applications legitimately have many threads waiting for work.

The important question is:

What are those threads waiting for?

I would inspect the thread dump and categorize the states:

WAITING
TIMED_WAITING
BLOCKED
RUNNABLE

Then I would examine the stack traces.

Possible causes

  • Waiting for database connections
  • Waiting for HTTP responses
  • Waiting for locks
  • Waiting on queues
  • Thread pool starvation
  • Slow external services
  • Incorrect synchronization
  • Long timeouts

For example:

200 threads

150 → waiting for database
30  → waiting for HTTP
15  → waiting for locks
5   → processing

The problem is probably not simply "too many threads".

The application has a dependency problem that is causing threads to remain occupied.

I would also correlate thread states with:

  • HikariCP metrics
  • HTTP client metrics
  • Database latency
  • Downstream service latency
  • Thread pool queue length

Senior-level answer: I would identify what the WAITING threads are waiting on rather than blindly increasing the thread pool size.

5. Your application has frequent Full GC pauses. How would you investigate?

Frequent Full GC activity is a serious JVM performance signal.

I would first determine:

  • How frequently Full GC occurs
  • How long the pauses are
  • Heap occupancy before GC
  • Heap occupancy after GC
  • Allocation rate
  • Old-generation usage
  • Object promotion behavior

The key question is whether the JVM is unable to reclaim enough memory.

Before GC:
Heap = 90%

After GC:
Heap = 88%

Then:

Heap = 92%
Heap = 95%
Heap = 98%
Full GC
Heap = 87%

This pattern could indicate significant long-lived object retention or excessive allocation pressure.

Possible causes

  • Memory leak
  • Large caches
  • Excessive allocation
  • Large collections
  • Incorrect heap sizing
  • Large temporary objects
  • Traffic increase
  • Application behavior change

I would use GC logs, JVM metrics, Java Flight Recorder, heap dumps, and allocation profiling.

I would avoid immediately increasing the heap.

If the application has a memory leak, a larger heap may simply delay the failure.

Senior-level answer: First determine whether Full GC is caused by allocation pressure, insufficient heap, or retained objects. Then use GC and heap evidence to identify the root cause rather than treating heap size as the universal solution.

6. A database query is fast in isolation but slow when called from the application. What would you check?

I would not assume that the database query itself is the only factor.

The database console might report:

Query execution = 40 ms

while the application reports:

Database operation = 800 ms

I would investigate the entire application-to-database path.

Things to check

  • Connection acquisition time
  • Connection pool exhaustion
  • Network latency
  • DNS behavior
  • Transaction overhead
  • ORM-generated SQL
  • Parameter differences
  • Result-set size
  • Object mapping time
  • Lazy loading
  • N+1 queries

For example:

Total application time = 800 ms

Connection wait       = 300 ms
SQL execution         = 50 ms
Result transfer       = 150 ms
Hibernate mapping     = 200 ms
Additional queries    = 100 ms

The SQL query itself is fast.

The overall database interaction is not.

Senior-level answer: I would compare database execution time with end-to-end application time and break the operation into connection acquisition, SQL execution, result transfer, ORM mapping, and additional queries.

7. Your application works perfectly after restart but becomes slower after running for several days. What could cause it?

This pattern strongly suggests a problem that accumulates over time.

I would investigate:

  • Memory leaks
  • Growing caches
  • Thread leaks
  • Connection leaks
  • Increasing queue sizes
  • File descriptor leaks
  • Stale connections
  • Increasing database data
  • Resource fragmentation
  • Scheduled jobs accumulating state

A restart resets many runtime resources:

Restart
   |
   +-- Clears heap
   +-- Clears thread state
   +-- Clears in-memory caches
   +-- Resets connection pools
   +-- Clears queues
   +-- Resets application state

If performance is good immediately afterward and gradually deteriorates, I would compare resource metrics over time.

For example:

Day 1:
Heap after GC = 35%

Day 3:
Heap after GC = 50%

Day 5:
Heap after GC = 70%

Day 7:
Heap after GC = 85%

This strongly suggests retained memory or another accumulating resource.

I would also investigate thread count and database connection behavior.

Senior-level answer: The restart itself is an important diagnostic clue. I would look for resources that accumulate over time rather than focusing only on the immediate symptom.

8. CPU usage is low, but requests are stuck. What resources could be exhausted?

Low CPU with stuck requests strongly suggests the application is waiting for something.

Possible exhausted resources include:

  • Database connections
  • HTTP connections
  • Application threads
  • Thread pool queues
  • File descriptors
  • Connection pools
  • Locks
  • Semaphores
  • External service capacity
  • Database capacity

For example:

CPU = 15%

HikariCP:
Active = 50
Idle = 0
Pending = 300

Result:
Requests wait for database connections

Another possibility is thread pool exhaustion:

200 worker threads

200 → waiting for slow HTTP service

New requests
     |
     v
Queue
     |
     v
High latency

I would use thread dumps, pool metrics, distributed traces, and operating-system resource metrics.

Senior-level answer: Low CPU tells me that computation may not be the bottleneck. I would look for exhausted concurrency, connection, locking, or I/O resources.

9. A Java application suddenly starts creating thousands of threads. What would you investigate?

Thousands of threads are usually a strong signal that something in the concurrency design has gone wrong.

I would investigate:

  • Thread creation rate
  • Thread pool configuration
  • Executors being created repeatedly
  • Whether thread pools are shut down
  • Request-driven thread creation
  • Scheduled tasks
  • Blocking operations
  • Third-party libraries
  • Virtual versus platform threads

A common mistake is creating an executor repeatedly:

for (...) {
    ExecutorService executor =
        Executors.newFixedThreadPool(10);
}

If those executors are not properly managed, the application can create an excessive number of threads.

Another problem is creating one thread for every incoming request.

I would inspect thread names in a thread dump.

pool-101-thread-1
pool-102-thread-1
pool-103-thread-1
...

This can reveal that hundreds of executors have been created.

I would also check whether the application is experiencing thread leaks.

Senior-level answer: I would identify who is creating the threads, whether they belong to expected executors, whether those executors are bounded and reused, and whether blocking operations are causing thread accumulation.

10. A deployment increases response time without increasing CPU or memory usage. Where would you look?

This is an interesting production scenario because the obvious JVM resource metrics look healthy.

I would compare the old and new versions and inspect the complete request path.

Possible causes include:

  • New database queries
  • N+1 queries
  • New downstream calls
  • Connection pool waiting
  • Lock contention
  • Changed timeout configuration
  • Changed serialization
  • Cache behavior
  • Network calls
  • Feature flags

For example:

Before deployment:

API = 200 ms

After deployment:

API = 1,000 ms

CPU = normal
Memory = normal

Trace:

Database = 600 ms
Downstream = 250 ms
Application = 150 ms

The JVM looks healthy because the bottleneck is external to CPU and memory.

I would also compare database query counts before and after the deployment.

Senior-level answer: Healthy CPU and memory do not prove that the application is healthy from a latency perspective. I would use distributed traces and dependency metrics to identify what changed in the request path.

11. One API instance is slow while other instances are healthy. How would you isolate the problem?

This situation strongly suggests an instance-specific problem.

I would compare the unhealthy instance against healthy instances.

Instance A → 150 ms
Instance B → 170 ms
Instance C → 5,000 ms
Instance D → 160 ms

I would compare:

  • Application version
  • Configuration
  • Environment variables
  • CPU
  • Memory
  • Thread count
  • Connection pool
  • GC activity
  • Network connectivity
  • Node placement
  • Pod events

It could be that one instance:

  • Has a stale connection pool
  • Is experiencing GC pressure
  • Runs on an unhealthy node
  • Has incorrect configuration
  • Has a network issue
  • Has a different application version

I would compare request traces routed specifically through that instance.

If removing the instance from service immediately restores overall performance, I would preserve the evidence and then investigate it separately.

Senior-level answer: I would isolate the unhealthy instance, compare its runtime and infrastructure metrics against healthy peers, and use instance-specific traces and logs to determine what makes it different.

12. Logs show no errors, but users report intermittent failures. How would you investigate?

The absence of errors in application logs does not mean that there are no failures.

The failure may occur:

  • Before reaching the application
  • At the API gateway
  • At the load balancer
  • During a timeout
  • Inside a downstream service
  • During network communication
  • Because of client-side behavior

I would first determine exactly what "failure" means.

Failure type?

HTTP 5xx?
HTTP 4xx?
Timeout?
Connection reset?
Slow response?
Incorrect response?
Client error?

I would use request IDs or trace IDs to follow individual failing requests.

I would also examine:

  • Access logs
  • Gateway logs
  • Load balancer metrics
  • Application metrics
  • Distributed traces
  • Client telemetry
  • Network metrics

Intermittent failures often indicate:

  • Only one unhealthy Pod
  • Connection pool exhaustion
  • Race conditions
  • Load balancing issues
  • Timeouts
  • Resource contention
  • Specific input patterns

Senior-level answer: I would not rely exclusively on application error logs. I would trace failed requests across the complete request path and determine whether the failure is application-level, infrastructure-level, dependency-level, or client-level.

13. A production issue disappears whenever debugging or logging is enabled. What could explain this behavior?

This is a classic example of a timing-sensitive problem.

One possibility is a race condition.

Adding logging or debugging changes execution timing:

Without logging:

Thread A ────────────────┐
Thread B ────────────┐   |
                     |   |
                     v   v
                  Race


With logging:

Thread A ────────wait───────
Thread B ────────────wait───

Race disappears

Logging changes thread scheduling and timing, so the race may become much harder to reproduce.

Other possibilities include:

  • Timeout-sensitive behavior
  • Connection timing
  • Network timing
  • Buffering
  • Deadlocks or lock contention
  • Heisenbug-like behavior

I would avoid concluding immediately that the problem is caused by logging.

Instead, I would capture non-intrusive evidence using metrics, thread dumps, tracing, profiling, and carefully selected logging.

I would specifically investigate shared mutable state and synchronization.

Senior-level answer: If observability changes the behavior, I would suspect timing-sensitive problems such as race conditions or synchronization issues and use low-impact diagnostic techniques to preserve the original behavior.

14. Your application is experiencing random timeout errors, but network monitoring looks normal. What would you investigate?

If network infrastructure looks healthy, I would investigate timeout sources above the network layer.

Possible causes include:

  • Thread pool exhaustion
  • Connection pool exhaustion
  • Slow database queries
  • GC pauses
  • Lock contention
  • Downstream service latency
  • Connection acquisition delays
  • Application-level queues
  • Incorrect timeout configuration

For example:

Network latency = 20 ms

But:

Thread waits = 1,500 ms
HikariCP waits = 800 ms
Database query = 500 ms

Total > timeout

The network is healthy, but the request still times out.

I would compare:

Connection timeout
Read timeout
Application timeout
Gateway timeout
Load balancer timeout
Database timeout

Timeouts can occur at multiple layers, and the shortest timeout may determine the actual behavior.

Senior-level answer: A network timeout does not necessarily mean a network problem. I would trace the request and measure queueing, connection acquisition, JVM pauses, database latency, and downstream processing time.

15. A Java service crashes only under heavy traffic. How would you reproduce and diagnose the problem?

I would try to reproduce the production workload under controlled conditions.

I would first identify the relationship between traffic and failure.

100 req/sec → healthy

500 req/sec → healthy

1,000 req/sec → latency increases

2,000 req/sec → failures

3,000 req/sec → service crashes

This suggests a capacity or concurrency-related problem.

I would perform controlled load testing while monitoring:

  • CPU
  • Heap
  • GC
  • Thread count
  • Thread pool queues
  • Database connections
  • HTTP connections
  • Database latency
  • Request latency
  • Error rate

Potential causes include:

  • Memory exhaustion
  • Thread exhaustion
  • Connection pool exhaustion
  • Database saturation
  • CPU saturation
  • Queue growth
  • Race conditions triggered by concurrency
  • Rate limits

I would gradually increase load rather than immediately generating maximum traffic.

This helps identify the point at which the system begins to degrade.

Senior-level answer: I would reproduce the traffic pattern under controlled load, monitor resource and dependency saturation, identify the first signal that changes before the crash, and use that evidence to locate the bottleneck.

16. A memory dump shows millions of objects of one particular class. How would you determine whether it is a memory leak?

Millions of objects do not automatically mean a memory leak.

The key question is:

Should these objects still be reachable?

I would investigate the reference chain to the objects.

Large number of objects
        |
        v
Who references them?
        |
        v
Why are they still reachable?
        |
        v
Should they have been released?

For example:

Application
    |
    v
Static Cache
    |
    v
HashMap
    |
    v
Millions of Customer objects

If the cache has no eviction policy, the objects may remain reachable indefinitely.

I would also compare heap dumps over time.

Heap Dump 1 → 100,000 objects
Heap Dump 2 → 500,000 objects
Heap Dump 3 → 2,000,000 objects

If the count continuously increases under a stable workload, that is strong evidence of a leak or unbounded retention.

I would use heap-analysis tools to inspect:

  • Object counts
  • Shallow size
  • Retained size
  • Reference chains
  • Dominators

Senior-level answer: I would not identify a leak based only on object count. I would determine why those objects remain reachable and compare retention patterns across heap dumps.

17. A thread dump shows multiple threads waiting for the same lock. How would you identify the bottleneck?

If multiple threads are waiting for the same lock, that lock may be a serialization point.

Thread A ──┐
Thread B ──┤
Thread C ──┼──→ Lock
Thread D ──┤
Thread E ──┘
             |
             v
         One thread
         executes

I would identify:

  • Which thread owns the lock
  • Which threads are waiting
  • What code is protected by the lock
  • How long the lock is held
  • Whether the critical section is larger than necessary

For example:

synchronized
void process() {

    updateMemory();

    callExternalService();

    writeToDatabase();

    performCalculation();
}

If the synchronized method calls an external service, the lock may be held for hundreds of milliseconds or several seconds.

That can create severe contention.

I would consider reducing the critical section:

Synchronize only the shared-state operation.

Do not hold the lock while:
- Calling external services
- Performing slow I/O
- Running expensive calculations

Depending on the design, I might also replace coarse-grained synchronization with more appropriate concurrency primitives or redesign the shared state.

Senior-level answer: I would identify the lock owner, measure lock hold time, inspect the protected code, and determine whether the critical section can be reduced or redesigned.

18. An application suddenly starts spending most of its time inside GC. What would you check?

This usually indicates significant allocation or memory pressure.

I would compare:

Application CPU
Allocation rate
GC CPU
GC frequency
GC pause time
Heap occupancy
Object promotion

A sudden increase in allocation rate may be caused by a new feature or workload.

For example:

Before deployment:

Allocation = 500 MB/sec

After deployment:

Allocation = 4 GB/sec

The application may now be generating large numbers of temporary objects.

Possible causes include:

  • Large JSON serialization
  • String creation
  • Large collections
  • Repeated object conversion
  • Excessive logging
  • Large query result processing
  • New caching behavior
  • Incorrect batching

I would use Java Flight Recorder, allocation profiling, GC logs, and heap analysis.

I would avoid immediately increasing heap size because the underlying problem may be excessive allocation rather than insufficient memory.

Senior-level answer: I would identify whether the problem is high allocation, poor object retention, or insufficient heap and then trace the allocation back to the responsible code path.

19. A service becomes slow only after a particular feature is enabled. How would you isolate the regression?

The feature flag provides a very useful comparison point.

Feature OFF
    |
    v
Latency = 200 ms

Feature ON
    |
    v
Latency = 1.5 sec

I would compare the same workload with the feature enabled and disabled.

I would examine:

  • Request latency
  • Database queries
  • CPU
  • Memory
  • GC
  • Thread activity
  • Downstream calls
  • Cache behavior
  • Payload size

Suppose the feature introduces an additional database lookup:

Before:

API → Database
      200 ms


After:

API → Database
      |
      +→ Feature query
      |
      +→ Another database query
      |
      1,500 ms

The feature itself may be functionally correct but introduce an unexpected performance cost.

I would also check whether the feature creates additional calls only for certain users or request types.

Once the bottleneck is confirmed, the feature can be optimized or temporarily disabled.

Senior-level answer: Use the feature flag as an experiment boundary. Compare the same workload with the feature on and off, inspect traces and resource metrics, identify the additional cost, and then optimize or roll back the feature.

20. You receive a production alert at 2 AM: "Java service is unhealthy." What would be your first 5 checks?

During an incident, I would avoid immediately making random configuration changes.

My first five checks would be:

1. Is the service actually unavailable?

I would check health endpoints, error rate, request rate, and availability from the user's perspective.

Requests
Errors
Latency
Availability

2. Did something change?

I would check recent:

  • Deployments
  • Configuration changes
  • Infrastructure changes
  • Database changes
  • Feature flag changes

3. What are the JVM resources doing?

I would check:

  • CPU
  • Memory
  • GC
  • Thread count
  • Heap usage

4. Are dependencies healthy?

I would check:

  • Database
  • Redis
  • Kafka
  • Downstream services
  • External APIs
  • Connection pools

5. What do traces and logs show?

I would identify a failing request and trace it through the system.

Client
  |
  v
Load Balancer
  |
  v
Java Service
  |
  +---- Database
  |
  +---- Redis
  |
  +---- Kafka
  |
  +---- Downstream Service

The objective during the first few minutes is to answer:

What is broken, when did it start, what changed, and where is the request failing?

Incident response mindset

If the service is severely impacting users, I would prioritize mitigation.

Depending on the evidence, mitigation might include:

  • Rollback
  • Removing an unhealthy instance
  • Disabling a feature
  • Scaling a bottleneck
  • Opening a circuit breaker
  • Failing over to a healthy dependency

Root-cause analysis can continue after the immediate user impact has been reduced.

Senior-level answer: My first priority is to establish impact and stabilize the system. Then I use metrics, logs, traces, JVM data, and recent changes to identify the root cause.

What These Production Scenarios Are Really Testing

These questions may appear to cover different problems, but they test a few fundamental skills expected from senior Java engineers.

JVM troubleshooting

  • CPU profiling
  • Heap analysis
  • Garbage collection
  • Thread dumps
  • Memory leaks

Concurrency

  • Thread contention
  • Locks
  • Thread pools
  • Deadlocks
  • Race conditions

Distributed systems

  • Network latency
  • Connection pools
  • Database bottlenecks
  • Downstream dependencies
  • Timeouts

Observability

  • Metrics
  • Logs
  • Distributed tracing
  • Profiling
  • Health checks

The common theme is:

Production Problem
       |
       v
Measure
       |
       v
Collect Evidence
       |
       v
Identify Bottleneck
       |
       v
Confirm Root Cause
       |
       v
Apply Targeted Fix
       |
       v
Verify
       |
       v
Prevent Recurrence

Common Mistakes When Troubleshooting Java Production Problems

1. Restarting immediately

A restart can restore service temporarily, but it may destroy valuable evidence about the underlying problem.

2. Increasing resources blindly

Adding CPU, memory, threads, or connections may hide the symptom without fixing the cause.

3. Looking only at application logs

Some problems occur in the database, network, load balancer, Kubernetes platform, or downstream services.

4. Looking only at average latency

p95 and p99 are often much more useful for understanding real user experience.

5. Assuming low CPU means the application is healthy

The application may be blocked on I/O, locks, connection pools, or external dependencies.

6. Treating every memory increase as a leak

Java's garbage collector and heap behavior must be understood before concluding that memory is leaking.

7. Changing multiple things at once

If you change five configurations simultaneously, it becomes difficult to determine which change actually fixed the problem.

Senior Java Production Troubleshooting Framework

When faced with an unfamiliar production problem, a useful framework is:

1. Confirm the symptom
        ↓
2. Measure the impact
        ↓
3. Establish a baseline
        ↓
4. Check recent changes
        ↓
5. Examine metrics
        ↓
6. Examine logs
        ↓
7. Trace affected requests
        ↓
8. Profile if necessary
        ↓
9. Identify the bottleneck
        ↓
10. Apply the smallest effective fix
        ↓
11. Verify recovery
        ↓
12. Prevent recurrence

This approach is useful because it prevents troubleshooting from becoming guesswork.

Key Takeaways

  • 100% CPU requires identifying the CPU-consuming threads and methods.
  • Increasing memory usage should be analyzed across garbage-collection cycles.
  • Average latency can hide serious p99 problems.
  • WAITING threads must be analyzed based on what they are waiting for.
  • Frequent Full GC requires investigation of allocation, retention, and heap behavior.
  • Database execution time and application database-operation time are not necessarily the same.
  • Performance degradation after several days can indicate an accumulating resource problem.
  • Low CPU with stuck requests often indicates waiting or resource exhaustion.
  • Thousands of threads require investigation of thread creation and executor management.
  • A deployment can increase latency without increasing CPU or memory.
  • An unhealthy application instance should be compared against healthy peers.
  • Intermittent failures require correlation across logs, metrics, traces, and infrastructure.
  • Debugging or logging can change timing and hide race conditions.
  • Timeouts can occur because of application-level waiting even when the network is healthy.
  • Load testing should identify the point at which the system begins to degrade.
  • A large object count does not automatically mean a memory leak.
  • Lock contention can create a bottleneck even when CPU is low.
  • High GC activity should be correlated with allocation and object-retention behavior.
  • Feature flags can be used to isolate production regressions.
  • During incidents, stabilize the system first and then perform deeper root-cause analysis.

Frequently Asked Questions

Are these questions suitable for senior Java interviews?

Yes. These scenarios focus on JVM internals, concurrency, performance, observability, databases, networking, and production incident troubleshooting. They are particularly useful for senior Java, Spring Boot, backend, and microservices interviews.

What is the most important skill in Java production troubleshooting?

The most important skill is evidence-driven diagnosis. Instead of immediately changing code or configuration, identify the symptom, measure it, trace the request, find the bottleneck, and confirm the root cause.

Should I restart a Java application when it becomes unhealthy?

A restart may be an appropriate mitigation when users are severely affected, but it should not replace investigation. Restarting can remove valuable diagnostic evidence and may only temporarily hide memory leaks, thread leaks, connection leaks, or other accumulating problems.

Why are p99 latency and tail latency important?

Average latency can hide a small percentage of extremely slow requests. In high-volume systems, even one percent of slow requests can affect thousands of users.

How important are thread dumps for Java troubleshooting?

Thread dumps are extremely useful for identifying blocked threads, deadlocks, lock contention, thread pool exhaustion, and unexpected waiting behavior.

What should a senior engineer do first during a production incident?

First establish the impact and stabilize the system if necessary. Then determine when the problem started, what changed, which components are affected, and where requests are failing.

Conclusion

Java production troubleshooting is not about memorizing a list of JVM commands.

It is about understanding how the application behaves under real-world conditions.

A senior engineer should be able to move from:

"The application is slow."

to:

"The p99 latency increased from 400 ms to 5 seconds."

Then:

"The JVM CPU is normal, but 70% of requests are waiting for database connections."

Then:

"The connection pool is exhausted because a new query introduced by the latest deployment holds connections for several seconds."

Then:

"After rolling back the deployment, latency returned to normal."

That is the difference between guessing and
troubleshooting.

Measure first. Follow the evidence. Find the bottleneck. Fix the root cause.

That mindset is one of the most valuable skills a senior Java engineer can bring to a production environment.

Java Production Problems – 20 Real-World Scenario-Based Interview Questions

Post a Comment

0 Comments

Close Menu