Your Spring Boot application works perfectly in development. Unit tests pass, APIs return the expected responses, and everything looks stable.
But what happens when hundreds of requests arrive simultaneously? What happens when a database transaction fails, an asynchronous task throws an exception, a cache contains stale data, or two application instances process the same business operation?
Some of the most dangerous production bugs are not obvious coding errors. They are problems caused by concurrency, transaction boundaries, asynchronous execution, serialization, configuration, and distributed-system behavior.
In this article, we will explore 20 hidden Spring Boot bugs that can break production, explain why they happen, and examine practical solutions that experienced Java developers should understand.
Scenario: A Spring Boot service stores the current user's information in an instance variable. The application works correctly during development, but production requests occasionally return another user's data.
Spring beans are singleton-scoped by default. A singleton bean is generally shared across requests and application threads within its Spring application context.
Consider this implementation:
@Service
public class UserService {
private String currentUser;
public String getUserDetails(String userId) {
currentUser = userId;
// Simulate business processing
return "Details for " + currentUser;
}
}Imagine two requests execute concurrently:
Request A:
currentUser = "Alice"
Request B:
currentUser = "Bob"
Request A:
returns "Details for Bob"The shared mutable field creates a race condition. One request can overwrite the value another request expects to use.
Store request-specific values in local variables or pass them explicitly through method parameters.
@Service
public class UserService {
public String getUserDetails(String userId) {
return "Details for " + userId;
}
}Local variables belong to individual method invocations, although objects referenced by those variables can still be shared.
For authenticated user information, use Spring Security's authentication context or another appropriate request-scoped mechanism rather than a mutable field in a singleton service.
Production lesson: Singleton beans are appropriate for stateless services. Do not store request-specific mutable state in their instance fields.
Scenario: Your service runs on three instances behind a load balancer. A user updates a setting, but subsequent requests sometimes return the old value.
One common cause is storing application state in local memory.
@Service
public class SettingsService {
private final Map<String, String> settings =
new ConcurrentHashMap<>();
public void update(String key, String value) {
settings.put(key, value);
}
public String get(String key) {
return settings.get(key);
}
}Even though ConcurrentHashMap supports concurrent access within one JVM, each application instance has its own map.
Instance A: theme = DARK
Instance B: theme = LIGHT
Instance C: theme = LIGHTA thread-safe collection does not make data consistent across multiple application instances.
First, determine whether the data should be shared or whether local state is intentional.
Also compare instance-specific configuration, application versions, cache contents, and deployment timestamps.
Production lesson: A concurrent collection solves local thread-safety problems, not distributed consistency problems.
Scenario: An API starts several asynchronous operations. One operation fails, but the response still indicates that everything succeeded.
CompletableFuture represents the outcome of an asynchronous computation. Exceptions are often captured in the future rather than thrown immediately on the thread that started the task.
Consider:
@GetMapping("/report")
public String generateReport() {
CompletableFuture.runAsync(() -> {
throw new IllegalStateException("Report failed");
});
return "Report generated successfully";
}The API returns before the asynchronous operation completes. The caller does not wait for the future, inspect its result, or handle its exception.
If the response depends on the asynchronous result, compose the futures and observe their outcomes.
CompletableFuture<String> reportFuture =
CompletableFuture.supplyAsync(
() -> generateReportData(),
reportExecutor
);
return reportFuture
.thenApply(data -> ResponseEntity.ok(data))
.exceptionally(ex -> {
log.error("Report generation failed", ex);
return ResponseEntity.internalServerError()
.body("Unable to generate report");
});The example illustrates asynchronous response composition; a controller would need a suitable declared return type, such as CompletableFuture<ResponseEntity<String>>, and an appropriately configured executor.
If the operation is independent of the response, handle its failure in the asynchronous workflow and expose it through logs, metrics, tracing, or a durable job-status mechanism.
Also remember that CompletableFuture.allOf() does not automatically provide successful results for all tasks. The combined future can complete exceptionally if a component future fails.
Production lesson: Starting asynchronous work is not the same as confirming that the work succeeded.
Scenario: An API works for most requests but occasionally throws IllegalStateException with a message indicating that the stream has already been operated upon or closed.
Java streams are designed for a single traversal. Once a terminal operation consumes a stream, the same stream cannot be reused.
Stream<String> names = users.stream()
.map(User::getName);
long count = names.count();
List<String> result = names.toList();The second terminal operation attempts to reuse an already-consumed stream.
Create a fresh stream for each independent traversal.
long count = users.stream()
.map(User::getName)
.count();
List<String> result = users.stream()
.map(User::getName)
.toList();If you need to reuse the calculated result, materialize it into a collection once.
List<String> names = users.stream()
.map(User::getName)
.toList();
long count = names.size();For large datasets, consider the memory implications of materializing an entire collection.
Production lesson: Reuse data when appropriate, but do not reuse a consumed Stream instance.
Scenario: A service changes an entity's field, but the database record remains unchanged.
JPA typically tracks changes to managed entities within a persistence context. When a managed entity is modified inside a transaction, the persistence provider can detect the change and synchronize it with the database during flushing.
However, an entity retrieved in one persistence context may become detached after that context ends.
Customer customer =
customerRepository.findById(id).orElseThrow();
// In a separate operation, after the entity is detached:
customer.setName("New Name");
// No managed persistence context tracks this change.Changing a detached object's field does not automatically persist the modification.
Load the entity within the intended transaction and modify the managed instance.
@Transactional
public void updateCustomer(Long id, String name) {
Customer customer =
customerRepository.findById(id).orElseThrow();
customer.setName(name);
}With a typical JPA transaction, dirty checking synchronizes the change at flush or commit. An explicit save() call is not generally required for an already managed entity, although repository APIs and application design may use one for clarity.
For detached entities, consider reloading the entity and applying changes or using merge() appropriately. Be careful: merging copies state into a managed instance and can overwrite fields if the detached object contains stale data.
Production lesson: Understand the difference between managed, detached, and transient entities before diagnosing missing database updates.
Scenario: An API returns an entity, but serialization fails with a recursion-related error or produces an excessively large JSON response.
Bidirectional relationships can create cyclic object graphs.
For example:
Order
|
v
Customer
|
v
Orders
|
v
Customer
|
v
Orders ...If the JSON serializer follows both sides of the relationship, it may repeatedly traverse the same objects.
The most maintainable solution for most REST APIs is to define explicit response DTOs.
public record OrderResponse(
Long id,
String orderNumber,
String customerName
) {}Then map only the data the API should expose.
Jackson annotations such as @JsonManagedReference and @JsonBackReference, or @JsonIdentityInfo, can help in specific serialization models. However, they do not replace deliberate API design.
DTOs also reduce accidental exposure of internal entity fields and make the response contract independent of the database model.
Production lesson: Do not assume that a JPA entity graph is a suitable REST response model.
Scenario: A customer changes an address successfully, but the API continues returning the previous address.
The database and cache are two separate representations of data. Updating the database does not automatically invalidate every cache that may contain the old value.
Database: address = "New Address"
Cache: address = "Old Address"This is a cache consistency problem.
For a common cache-aside design, a typical approach is:
For example, a Spring service can use transaction-aware event handling to invalidate a cache after the database transaction commits.
@Transactional
public void updateCustomer(Long id, String address) {
Customer customer =
customerRepository.findById(id).orElseThrow();
customer.setAddress(address);
applicationEventPublisher.publishEvent(
new CustomerUpdatedEvent(id)
);
}A listener can invalidate the cache after commit:
@TransactionalEventListener(
phase = TransactionPhase.AFTER_COMMIT
)
public void onCustomerUpdated(CustomerUpdatedEvent event) {
customerCache.evict(event.customerId());
}This example is appropriate when the cache is local to the application and the listener can reliably access it. For distributed caches, multiple instances, or critical invalidation requirements, additional coordination or a durable event/outbox design may be needed.
Also consider TTL, cache stampedes, failed invalidation, and how stale a response is allowed to be.
Production lesson: Define cache consistency and invalidation behavior before adding a cache to a business-critical workflow.
Scenario: An order is saved, and an event is published to send a confirmation or update another system. The application crashes before the operation completes.
Spring application events are commonly delivered within the application process. An ordinary event listener does not automatically create a durable message that survives a process crash.
Consider this workflow:
Save Order
|
v
Publish Event
|
v
Listener starts
|
X
Application crashes
|
X
Operation may never completeMoving work to an asynchronous listener does not, by itself, make it durable.
Choose a reliability strategy based on the business requirement.
With an outbox design, the business record and an event record are written in the same database transaction. A separate publisher delivers pending events to a broker and records delivery progress.
Database Transaction
|
+-- Save Order
|
+-- Save Outbox Event
|
v
Commit together
|
v
Outbox Publisher
|
v
Message BrokerThis avoids the dual-write problem in which the database commit succeeds but event publication fails. The publisher and consumers still need recovery and duplicate-handling logic.
Production lesson: An in-memory event is not a durable delivery guarantee.
Scenario: A notification method is annotated with @Async. The main business operation succeeds, but notification processing fails without any visible error in the calling method.
An asynchronous method runs independently of the calling thread. The caller may return before the task finishes.
When an asynchronous method returns void, there is no returned future through which the caller can inspect the outcome.
@Async
public void sendNotification() {
throw new IllegalStateException("Notification failed");
}The exception cannot be propagated back through a method call that has already returned normally.
For work whose outcome matters to the caller, return a future and compose or inspect the result.
@Async("notificationExecutor")
public CompletableFuture<Void> sendNotification() {
notificationClient.send();
return CompletableFuture.completedFuture(null);
}The caller can observe the future, attach failure handling, or combine it with other asynchronous tasks. The caller must actually inspect or compose the future for that handling to be useful.
For fire-and-forget work, configure an AsyncUncaughtExceptionHandler for applicable void-returning @Async methods. Also use structured logging, metrics, tracing, and an explicit retry strategy when required.
Configure and monitor the executor's thread count and queue capacity. A reliable asynchronous design also needs a clear policy for rejected tasks and application shutdown.
Production lesson: Asynchronous work needs its own error-handling and observability strategy.
Scenario: A request contains an invalid amount or missing field, but the controller still processes it.
Declaring validation annotations does not automatically guarantee that every incoming request is validated.
For example:
public class PaymentRequest {
@NotNull
@Positive
private BigDecimal amount;
// getters and setters
}The controller must trigger validation for the request body, and the appropriate validation dependency and configuration must be available.
@PostMapping("/payments")
public ResponseEntity<String> createPayment(
@Valid @RequestBody PaymentRequest request) {
return ResponseEntity.ok("Processed");
}Without the appropriate validation trigger, invalid values may reach business logic.
@Valid or an appropriate @Validated setup?For nested request objects, cascading validation may require @Valid on the nested field. For monetary amounts, also validate business constraints such as currency, scale, and allowed ranges.
Production lesson: Validation annotations only help when validation is actually invoked and the constraints represent the required business rules.
Scenario: A payment fails, but the API returns HTTP 200 with a message saying the operation failed.
The controller may be catching every exception and returning a normal response:
@PostMapping("/payments")
public String pay() {
try {
paymentService.process();
return "Success";
} catch (Exception e) {
return "Payment failed";
}
}Because the method returns normally, Spring MVC can produce a successful HTTP response unless another response status is specified.
Define a consistent error contract and map failures to appropriate HTTP status codes.
For example, a validation error may use HTTP 400, an authentication failure may use HTTP 401, an authorization failure may use HTTP 403, and an unexpected server failure may use HTTP 500.
Use centralized exception handling where appropriate:
@RestControllerAdvice
public class ApiExceptionHandler {
@ExceptionHandler(InvalidPaymentException.class)
public ResponseEntity<ApiError> handleInvalidPayment(
InvalidPaymentException ex) {
return ResponseEntity.badRequest()
.body(new ApiError(
"INVALID_PAYMENT",
"The payment request is invalid"
));
}
}The example assumes an ApiError response type exists.
Not every business failure should become an HTTP error. The response should follow the API contract and distinguish accepted asynchronous work, business rejections, and technical failures.
Production lesson: HTTP status codes and response bodies should accurately communicate the outcome of the request.
Scenario: A new field is added to a JPA entity, and it automatically appears in a REST response. The field contains internal or sensitive information.
When entities are serialized directly, changes to the persistence model can unexpectedly change the API response.
For example:
@Entity
public class Customer {
private Long id;
private String name;
private String email;
private String internalNotes;
}If the entity is returned directly and the serializer exposes the field, the API may reveal data that clients should never receive.
Use response DTOs that explicitly define the public API contract.
public record CustomerResponse(
Long id,
String name,
String email
) {}Map only approved fields into the response object.
Additional safeguards include authorization checks, response-contract tests, sensitive-field classification, and security reviews when entity models change.
Serialization annotations can help in specific cases, but they are not a complete substitute for an explicit API response model.
Production lesson: Data that exists inside the application should not automatically become public API data.
Scenario: An administrator changes a timeout or feature setting in a configuration system, but the running application continues behaving as before.
Spring Boot configuration is generally bound during application startup or bean initialization. Changing an external configuration source does not automatically update every already-created object.
For example:
@Component
@ConfigurationProperties(prefix = "payment")
public class PaymentProperties {
private Duration timeout;
// getters and setters
}If the configuration source changes after startup, the value already bound into the application may remain unchanged unless the application uses an appropriate refresh mechanism.
Choose a configuration strategy based on the required behavior.
Some Spring Cloud environments provide refresh capabilities, but availability and behavior depend on the components and versions used. Refreshing a property source does not mean every singleton bean will automatically reconfigure itself correctly.
For critical timeouts, connection pools, and security settings, test whether changing the value at runtime is safe before enabling dynamic updates in production.
Production lesson: Configuration refresh is an application lifecycle and consistency problem, not just a property-file change.
Scenario: A Spring Boot test passes when executed alone but fails when the complete suite runs.
This is often caused by test isolation problems.
For example, a test may insert a customer with a fixed unique email address. When the test runs again, the insert fails because the record remains in the database.
Use deterministic test data, isolated containers or databases where appropriate, and explicit synchronization for asynchronous tests.
Production lesson: A flaky test is often evidence of hidden coupling, shared state, timing assumptions, or an incomplete test lifecycle.
Scenario: A database schema migration is successful, but some application instances fail during a rolling deployment.
During a rolling deployment, old and new application versions may run simultaneously. If a database migration immediately removes or renames a column required by the old version, the old instances can fail.
Old Application ----+
|
v
Database
^
|
New Application ----+Both versions may need to operate against the same schema during the transition.
Use a backward-compatible expand-and-contract migration strategy.
Database migrations should be tested against the previous application version and the new version. Destructive changes should have a recovery plan.
Production lesson: A successful migration is not necessarily a safe migration. Consider every application version that may run during deployment and rollback.
Scenario: A scheduled report or cleanup task does not run at the expected time, especially during periods of high application load.
Spring scheduling behavior depends on how scheduling is configured and which scheduler executes the task.
Potential causes include:
For example, a slow scheduled task can occupy a scheduler thread for several minutes. Other tasks may then start later than expected if they share that constrained scheduler.
Configure a suitable scheduler and isolate long-running work where appropriate. For critical jobs, consider a dedicated worker or job-processing system with explicit execution tracking and recovery.
Also remember that @Scheduled normally runs independently on every application instance. If a job must execute once across the deployment, use distributed coordination or a single designated job runner.
Production lesson: Scheduling is not a guarantee that a business job will execute successfully at an exact wall-clock time. Monitor actual execution and completion.
Scenario: A payment or order is successfully processed by the server, but the client does not receive the response before its timeout. The client retries, and the business operation runs twice.
A client timeout does not prove that the server failed to complete the operation.
Client sends request
|
v
Server processes payment
|
v
Payment succeeds
|
X
Response delayed or lost
|
v
Client retries
|
v
Payment processed againThis is a fundamental distributed-systems problem: the client may not know whether the first operation succeeded.
Use idempotency for critical business operations.
The client sends a stable idempotency key for the logical operation:
POST /payments
Idempotency-Key: order-123-payment-1The server stores the key, request identity, processing state, and outcome according to the business requirements. Repeated requests with the same key should not create a second payment.
The implementation must handle concurrent duplicate requests safely, define key expiry, and reject reuse of the same key for an incompatible request.
Database constraints and transaction design can help enforce uniqueness. For external payment providers, use their idempotency mechanisms where available and reconcile uncertain outcomes when necessary.
Production lesson: Critical APIs should assume requests can be delivered more than once.
Scenario: Debug logs contain bearer tokens, passwords, payment details, or customer information.
Logs are frequently collected by centralized logging platforms, retained for extended periods, and made accessible to multiple operational teams.
Logging sensitive information can turn a routine debugging practice into a security incident.
For example, log a request identifier, customer reference, operation name, and outcome rather than a complete payment request containing sensitive data.
log.info(
"Payment processing completed: requestId={}, status={}",
requestId,
status
);Use correlation IDs or trace IDs to connect logs without copying sensitive payloads into every message.
If a real credential is exposed, removing it from future logs is not sufficient. Follow the incident process, rotate the credential where necessary, and assess access to historical logs.
Production lesson: Logging should provide diagnostic value without exposing secrets or unnecessary personal data.
Scenario: The application starts, but requests fail because the database or another required dependency is unavailable. Kubernetes continues sending traffic to the instance.
Successful process startup does not necessarily mean the application is ready to serve traffic.
A liveness check and a readiness check have different purposes.
Spring Boot Actuator can expose health endpoints, and Kubernetes can use configured probes to make routing and restart decisions.
Configure readiness checks to reflect the dependencies required for the application to serve its intended traffic.
For example, a service that cannot handle any useful requests without its database may need database-aware readiness. A service that can serve some requests without an optional analytics service may not need to become entirely unready when analytics is unavailable.
A Spring Boot configuration may enable readiness health information through Actuator and Kubernetes probe support, depending on the Spring Boot version and setup.
For critical dependencies, also consider connection timeouts, startup behavior, retry policies, and how the application recovers when the dependency returns.
Important: Do not make every dependency failure a liveness failure. If all pods restart whenever a shared dependency becomes unavailable, the result can be a restart loop rather than recovery.
Production lesson: Readiness should control traffic admission; liveness should support recovery from a genuinely unhealthy process.
Scenario: A production issue occurs intermittently, but local testing never reproduces it.
Production differs from local development in many ways:
A race condition may only occur when two requests overlap. A query may become slow only when a large tenant's data is processed. A timeout may only happen when several downstream calls are slow at the same time.
Use observability to build an evidence-based timeline.
Structured logs: Record relevant fields such as request ID, trace ID, operation, outcome, application version, and safe business identifiers.
Metrics: Track request rates, error rates, p95/p99 latency, connection pool utilization, queue depth, JVM memory, GC pauses, and dependency latency.
Distributed tracing: Determine where request time is spent across controllers, databases, caches, and downstream services.
Production comparisons: Compare healthy and unhealthy instances, recent releases, configuration, traffic, and resource metrics.
Reproduction: Use sanitized production-like data, realistic concurrency, controlled fault injection, and repeatable load tests where possible.
For JVM-specific problems, use thread dumps, JFR recordings, GC logs, or heap dumps when appropriate and operationally safe.
Production lesson: If a bug cannot be reproduced locally, improve the evidence available from production instead of guessing at the cause.
Real incidents often involve several small mistakes that reinforce each other.
Consider an order-processing service:
Client submits an order
|
v
Request times out
|
v
Client retries
|
v
No idempotency protection
|
v
Duplicate order created
|
v
Async event triggers processing
|
v
Listener fails during a crash
|
v
Order and downstream systems disagree
|
v
Manual recovery requiredThe initial timeout was not necessarily the biggest problem. Missing idempotency, unreliable event handling, and inadequate operational visibility made the incident worse.
This is why production engineering must consider the complete workflow, not just individual methods.
Common examples include shared mutable state in singleton beans, incorrect transaction boundaries, missing asynchronous error handling, stale caches, duplicate operations after retries, and incompatible database migrations.
No. Singleton scope means that a bean instance is shared within its application context. It does not automatically make mutable instance fields thread-safe. Stateless services are generally easier to use safely across concurrent requests.
Asynchronous methods execute independently of the caller. Exceptions may be stored in a future or handled by an asynchronous exception handler rather than propagated through the original call stack. Monitoring and explicit failure handling are essential.
Entities can expose internal fields, create recursive serialization graphs, trigger unexpected lazy-loading queries, and tightly couple the API contract to the database model. DTOs provide more control over the response.
Use an idempotency mechanism with a stable key for each logical payment operation, enforce safe handling of concurrent duplicate requests, and coordinate with the payment provider's idempotency or reconciliation mechanisms when available.
Liveness helps determine whether a process should continue running. Readiness determines whether an instance should receive traffic. Treating every dependency failure as a liveness failure can create unnecessary restarts.
Correlate structured logs, metrics, distributed traces, application versions, configuration, JVM diagnostics, database behavior, and infrastructure metrics. Reproduce the failure with sanitized production-like data and realistic concurrency where possible.
Knowing Spring Boot annotations is important, but understanding the behavior behind those annotations is what helps prevent production failures.
A singleton bean can introduce a race condition. An asynchronous task can fail after the API returns. A database transaction can succeed while the cache remains stale. A retry can create duplicate payments. A successful deployment can break older application instances if the database schema changes incompatibly.
These are not merely theoretical interview questions. They represent real engineering problems that appear when applications encounter concurrency, failures, scale, and distributed-system boundaries.
Senior developers do not just ask whether the application works under normal conditions. They ask what happens when requests overlap, dependencies fail, messages are delivered twice, processes restart, and the system must recover.
That is the difference between writing working code and building reliable production software.
Which of these 20 hidden bugs have you encountered in production?
0 Comments