Spring Boot makes it easy to build production-ready applications.
But Spring Boot does not automatically make an application production-ready.
Many serious production incidents happen because developers use Spring Boot features correctly from a coding perspective, but without understanding their behavior under high traffic, failures, concurrency, distributed systems, database load, or Kubernetes resource constraints.
A connection pool that works perfectly in development can become a bottleneck in production. A simple JPA repository call can generate hundreds of SQL queries. A retry mechanism can turn a temporary downstream failure into a cascading outage. A scheduled job can execute multiple times because several application instances are running.
This article covers 25 common Spring Boot mistakes that can cause real production problems, along with why they happen, how to identify them, and what a better production approach looks like.
One of the most common mistakes is deploying a Spring Boot application with the default database connection pool configuration without understanding the production workload.
Spring Boot commonly uses HikariCP for JDBC connection pooling. The application does not create a new database connection for every request. Instead, connections are maintained in a pool and borrowed by application threads when needed.
The problem appears when the production traffic and database workload are significantly larger than the assumptions made during development.
For example, imagine a service with a small connection pool receiving hundreds of concurrent requests. If most requests require database access and connections remain occupied for a long time, new requests may wait for connections.
Request
|
v
Application Thread
|
v
Wait for DB Connection
|
v
HikariCP
|
v
DatabaseIncreasing the pool size blindly is not necessarily the solution. The database itself has a maximum capacity. Giving every application instance a huge connection pool can overwhelm the database.
Production lesson: Tune the connection pool based on application concurrency, query duration, database capacity, and the number of application instances.
Returning JPA entities directly from REST controllers looks convenient:
@GetMapping("/users/{id}")
public User getUser(@PathVariable Long id) {
return userRepository.findById(id).orElseThrow();
}But this tightly couples your API contract to your persistence model.
Your database entity is designed to represent persistence. Your API response is designed to represent what clients should receive. These are different responsibilities.
A better approach is to use DTOs.
public record UserResponse(
Long id,
String name,
String email
) {}Then explicitly map the entity to the response model.
This gives the API an independent contract and prevents persistence implementation details from leaking into the API.
Production lesson: Entities belong to the persistence layer. DTOs define the API contract.
One of the most expensive Spring Data JPA mistakes is allowing one database query to silently become hundreds or thousands.
Consider loading a list of orders and then accessing each order's customer:
List<Order> orders = orderRepository.findAll();
for (Order order : orders) {
System.out.println(order.getCustomer().getName());
}The first query might load all orders:
SELECT * FROM orders;Then accessing the customer for each order can trigger additional queries:
SELECT * FROM customer WHERE id = 1;
SELECT * FROM customer WHERE id = 2;
SELECT * FROM customer WHERE id = 3;
...
One request can therefore produce hundreds of SQL queries.
Depending on the use case, solutions include JOIN FETCH, entity graphs, DTO projections, carefully designed queries, or batch fetching.
Production lesson: A JPA repository method that looks simple in Java can generate a surprisingly expensive database workload.
@Transactional is powerful, but many developers treat it as a magic annotation.
Consider:
@Transactional
public void createOrder() {
saveOrder();
reduceInventory();
createPaymentRecord();
}The important question is not simply whether the annotation exists. You need to understand what code executes inside the transaction, what database resources participate in it, and what happens when an exception occurs.
Transactions can affect:
A transaction that remains open while performing slow external HTTP calls can hold a database connection and potentially database locks for much longer than necessary.
BEGIN TRANSACTION
Update database
Call external API <-- slow
Another database operation
COMMITThis can increase connection pool pressure and lock contention.
Production lesson: Keep transaction boundaries intentional and as small as practical while preserving business consistency.
This is a classic Spring interview question and a common production mistake.
public void methodA() {
methodB();
}
@Transactional
public void methodB() {
// database operation
}Developers may assume that calling methodB() automatically activates the transaction.
But Spring's declarative transaction management commonly relies on proxies. A call made internally through this.methodB() does not pass through the Spring proxy.
External caller
|
v
Spring Proxy
|
v
@Transactional methodBut self-invocation looks more like:
methodA()
|
+---- this.methodB()
|
X
proxy bypassedPossible solutions include moving the transactional method to another Spring bean or restructuring the service boundary.
Production lesson: Understanding Spring proxy behavior is essential when using annotations such as @Transactional, @Async, and caching annotations.
Retries can help recover from temporary failures.
But unlimited retries can turn a small outage into a much larger one.
Imagine Service A calls Service B.
Service A
|
v
Service B unavailable
|
v
Retry
|
v
Retry
|
v
Retry
|
v
More traffic to an already unhealthy serviceThis creates a retry storm.
Retries should normally have:
Not every error should be retried. Retrying a validation error or authentication failure usually makes no sense.
Production lesson: A retry is additional traffic. Always ask whether the dependency can handle that additional traffic during failure.
One of the most dangerous mistakes in a microservice is making an HTTP call without properly configured connection and read timeouts.
Consider:
Spring Boot API
|
v
External API
|
X
No responseIf there is no effective timeout, application threads can remain occupied waiting for the remote service.
Eventually:
External service slow
|
v
Requests wait
|
v
Threads become occupied
|
v
Thread pool exhaustion
|
v
New requests cannot execute
|
v
Service becomes unhealthyConfigure appropriate connection and response/read timeouts for every remote dependency.
Also consider connection pool limits, retry policies, circuit breakers, and fallback behavior.
Production lesson: Every remote call should have a clearly defined failure boundary.
A common reaction to slow requests is:
"Let's increase the thread pool."
This can make the problem worse.
Threads consume memory and CPU. More threads do not automatically mean more throughput.
For an I/O-heavy workload, increasing concurrency may help up to a point. Beyond that point, you may create more contention and resource pressure.
Too much concurrency
|
+-- CPU contention
+-- Memory pressure
+-- Database connection pressure
+-- Context switching
+-- Downstream overloadThread pool sizing should consider the workload, CPU count, blocking behavior, downstream capacity, database pool size, and request characteristics.
Production lesson: Concurrency is a resource. More threads are not free.
parallelStream() can make code look faster:
items.parallelStream()
.map(this::process)
.toList();But it uses shared JVM resources, including the common ForkJoinPool in typical usage.
Problems can occur when:
For example, if every element performs an HTTP request, parallelStream may generate a large amount of concurrent network traffic.
For production workloads, explicitly designed executors or asynchronous architectures are often more predictable.
Production lesson: Parallelism should be an intentional architecture decision, not a shortcut for performance.
This is another common cause of production memory problems.
Code such as:
List<Customer> customers = repository.findAll();may work perfectly with 10,000 records in development.
What happens when the table contains 10 million records?
The application may attempt to load a massive amount of data into memory.
Potential consequences include:
Use pagination, streaming, batch processing, or database-side operations depending on the use case.
Page<Customer> page =
repository.findAll(PageRequest.of(0, 500));Production lesson: Never assume production datasets will remain the size of development datasets.
@Async can move work away from the calling thread, but simply adding the annotation does not create a complete asynchronous architecture.
@Async
public void sendNotification() {
// work
}You need to understand which executor processes the work, how many threads it can create, what happens when the queue fills, and how failures are handled.
Without proper configuration and monitoring, asynchronous workloads can create:
Production executors should have explicit capacity and monitoring.
Production lesson: Asynchronous execution moves work somewhere else; it does not eliminate resource consumption.
Never treat configuration files inside source control as secure secret storage.
A configuration such as:
spring.datasource.password=mySecretPassword
payment.api-key=abc123can accidentally expose credentials through Git history, build artifacts, logs, or developer machines.
Production applications should use appropriate secret-management mechanisms and inject secrets securely at runtime.
Also remember that simply hiding a value from the main configuration file does not make it secure if the secret is subsequently printed into logs.
Production lesson: Configuration and secrets are different concerns. Treat credentials as secrets throughout their lifecycle.
Development and production have different characteristics.
Examples include:
Using the same configuration everywhere often leads to either over-provisioning development environments or under-configuring production.
Use environment-specific configuration while keeping the configuration structure consistent.
Production lesson: The application should behave predictably across environments, but the resource and operational configuration should reflect the environment.
Logs are important, but logs alone are not enough for modern production systems.
Imagine an API becomes slow.
Your logs may tell you that requests are arriving, but they may not tell you exactly where the 4 seconds are being spent.
You need multiple observability signals:
Logs
+
Metrics
+
Distributed Tracing
+
Health Checks
+
JVM Monitoring
+
Infrastructure MetricsUseful metrics include:
Production lesson: Logs explain events. Metrics show trends. Traces show where time is spent.
Returning internal exception details directly to clients can expose implementation information.
For example:
{
"error": "SQL error: password authentication failed for user admin"
}This is inappropriate for an external API.
Clients generally need a controlled error response rather than internal stack traces, SQL statements, filesystem paths, database details, or internal service information.
Use centralized exception handling with controlled response models.
{
"code": "INTERNAL_ERROR",
"message": "Something went wrong",
"traceId": "abc-123"
}The detailed technical information can remain in secure server-side logs and observability systems.
Production lesson: Error responses are part of your security boundary.
This pattern is dangerous:
try {
processOrder();
} catch (Exception e) {
log.error("Something went wrong", e);
}What happens after the exception?
If the application simply logs the exception and continues, the caller may receive a misleading success response.
Broad exception handling can also hide programming bugs and make operational diagnosis difficult.
Prefer handling exceptions according to business meaning.
catch (PaymentTimeoutException e) {
// specific recovery
}
catch (InsufficientBalanceException e) {
// business response
}Unexpected exceptions should generally be allowed to reach a centralized error-handling layer rather than being silently swallowed.
Production lesson: Exception handling should recover from known failures, not hide unknown failures.
During deployments, application instances are often terminated and replaced.
If shutdown is not handled correctly, requests currently being processed can be interrupted.
Load Balancer
|
v
Spring Boot Instance
|
X
Instance terminated
|
X
Request interruptedA graceful shutdown strategy allows the instance to stop receiving new work while completing existing requests within an appropriate termination window.
This becomes particularly important for:
Production lesson: Deployment is not just starting a new version. It is also safely shutting down the old version.
Consider:
@Scheduled(cron = "0 0 * * * *")
public void generateReport() {
// job
}If you have five application instances, the scheduled task may execute on every instance.
Instance 1 ----+
Instance 2 ----+
Instance 3 ----+----> Same scheduled job
Instance 4 ----+
Instance 5 ----+If the job sends emails, generates reports, processes payments, or modifies records, duplicate execution can cause serious problems.
Solutions include distributed locking, dedicated job workers, external schedulers, queue-based processing, or another architecture that guarantees appropriate execution semantics.
Production lesson: In a distributed system, "run once" requires coordination.
Caching can dramatically improve performance.
But stale data can create correctness problems.
Database
|
v
Cache contains old value
Database updated
|
X
Cache still contains old valueThe difficult question is not usually "How do I add a cache?"
The difficult question is:
"When is this cached value no longer valid?"
Before implementing caching, define:
Also consider cache stampedes when many requests attempt to rebuild the same missing cache entry simultaneously.
Production lesson: Caching is not simply a performance feature. It changes data-consistency behavior.
This becomes particularly dangerous during rolling deployments.
Imagine the old application expects:
customer_namewhile the new application expects:
full_nameIf both versions temporarily run during deployment, one version may fail against the database schema required by the other.
Old Application
|
+----------------+
|
Database
|
+----------------+
|
New ApplicationProduction deployments often require backward-compatible schema changes.
A common approach is:
This is commonly described as an expand-and-contract migration strategy.
Production lesson: Database changes must consider application versions that coexist during deployment.
Synchronous REST calls are simple and often appropriate.
But not every operation needs to block the user's request while waiting for another service.
Consider an order workflow:
Order API
|
+--> Payment
|
+--> Inventory
|
+--> Notification
|
+--> AnalyticsIf notification and analytics do not need to complete before the response is returned, synchronously waiting for them increases latency and coupling.
Asynchronous messaging can be more appropriate for independent, eventually consistent operations.
Order Service
|
v
Message Broker
| |
v v
Notification AnalyticsThis does not mean "use Kafka everywhere." Synchronous communication is still appropriate when an immediate response is required.
Production lesson: Choose synchronous or asynchronous communication based on business requirements, consistency needs, latency, and failure behavior.
This is especially dangerous for payments, orders, inventory, and financial operations.
Imagine a client sends:
POST /paymentsThe server successfully processes the payment, but the response is lost because of a network problem.
The client retries the request.
Request 1
|
v
Payment succeeds
|
X
Response lost
Request 2
|
v
Payment processed again
Without idempotency, the customer may be charged twice.
A common approach is to accept an idempotency key:
Idempotency-Key: payment-12345The server stores the result associated with that key and ensures repeated requests do not execute the business operation multiple times.
Production lesson: Network retries mean the same logical request can arrive more than once. Critical operations must be designed accordingly.
Java applications need memory not only for the Java heap.
A simplified view is:
Container Memory
|
+-- Java Heap
+-- Metaspace
+-- Thread Stacks
+-- Direct Buffers
+-- Native Memory
+-- JVM OverheadIf Kubernetes memory limits are set too aggressively, the container can be killed even when the configured Java heap appears to have available memory.
CPU limits can also affect application latency. CPU throttling may cause an application to become slow even when the code itself has not changed.
Monitor:
Production lesson: JVM tuning and Kubernetes resource configuration must be considered together.
A process being alive does not necessarily mean the application is healthy.
For example:
Spring Boot process
|
v
Running
Database
|
X
UnavailableThe process may still respond to a basic health check even though it cannot serve important requests.
However, blindly making every dependency mandatory for liveness can also create problems. A temporary dependency outage should not necessarily cause every application instance to restart.
This is why readiness and liveness need to have different purposes.
Dependency checks should be designed carefully based on the role of the dependency and the desired recovery behavior.
Production lesson: A health check should help the platform make the right operational decision, not simply return "UP."
This is perhaps the most important mistake on the list.
An API becomes slow.
The developer immediately starts optimizing Java code.
But the actual problem could be:
Consider an API whose response time increased from 300 milliseconds to 4 seconds.
Before modifying application code, look at the distributed trace:
Total request: 4,000 ms
Spring Boot processing: 150 ms
Database: 2,700 ms
Downstream API: 800 ms
Connection wait: 250 ms
Other: 100 msOptimizing 150 milliseconds of Java processing will barely change the user experience.
The correct question is:
"Where is the request actually spending its time?"
Measure first. Then optimize.
Production lesson: Performance engineering should be evidence-driven, not assumption-driven.
The most interesting production incidents usually involve several mistakes at the same time.
For example:
Traffic increases
|
v
N+1 queries
|
v
Database queries become slower
|
v
Connections remain occupied longer
|
v
HikariCP pool becomes exhausted
|
v
Application threads wait
|
v
API latency increases
|
v
Clients retry
|
v
More requests arrive
|
v
Database gets more traffic
|
v
System-wide degradationNotice that the original problem was not necessarily a single Java coding mistake.
It was a chain reaction.
When a Spring Boot service suddenly becomes slow or unhealthy, avoid immediately changing configuration.
Use a structured investigation.
Find where the request spends its time.
Request
|
+-- Controller
|
+-- Database
|
+-- Redis
|
+-- Payment Service
|
+-- External APILook at latency, errors, timeouts, retries, and saturation.
A production investigation becomes difficult when multiple configuration changes are made simultaneously.
Make a controlled change, measure the result, and continue from evidence.
| Production Symptom | Possible Mistakes to Investigate |
|---|---|
| High API latency | N+1, slow DB, downstream latency, thread pool, connection pool |
| Low CPU but slow requests | Database waits, connection pool, locks, network I/O |
| High CPU | Excessive computation, GC, thread contention, traffic spike |
| Memory continuously increasing | Large datasets, cache growth, object retention, memory leak |
| Duplicate operations | Retries, scheduled jobs, missing idempotency |
| Service unavailable during deployment | No graceful shutdown, incompatible schema, readiness problems |
| Database overloaded | N+1, oversized connection pools, excessive retries, inefficient queries |
| Stale data | Incorrect cache invalidation or TTL strategy |
| Pod restarts | Memory limits, JVM configuration, health checks, application crashes |
| Intermittent downstream failures | No timeout, retry storm, connection exhaustion, dependency saturation |
Each of these statements can sometimes be true, but none should be treated as a diagnosis.
A senior engineer does not simply ask whether the code works.
They ask:
This is the difference between writing application code and engineering a production system.
There is no single mistake responsible for every incident, but database-related issues such as inefficient queries, N+1 queries, connection pool exhaustion, and poor transaction boundaries are extremely common production bottlenecks.
Not technically, but DTOs are strongly useful when the API contract should be independent of the persistence model. They also help control the response shape and prevent accidental exposure of internal fields and relationships.
Low CPU does not mean the application is idle. Threads may be waiting for database connections, database queries, locks, network responses, external APIs, or other I/O operations.
Not automatically. First determine whether requests are waiting for connections and whether the database has enough capacity. Increasing the pool without understanding the database workload can make the problem worse.
Retries generate additional traffic. When a downstream service is already unhealthy, unlimited retries can overload it further and create a cascading failure.
Without timeouts, application threads can remain blocked while waiting for an unhealthy dependency. Enough blocked requests can eventually exhaust the application's available concurrency.
Spring's declarative transaction support commonly relies on proxies. An internal method call within the same object can bypass the proxy, meaning the transactional interceptor may not be invoked.
Caching introduces another copy of data. If invalidation, TTL, consistency, and failure behavior are not planned, applications can serve stale or incorrect data.
Networks are unreliable and clients may retry requests. Without idempotency, the same logical payment request can potentially execute more than once.
Start with evidence: compare p50/p95/p99 latency, request volume, errors, recent deployments, distributed traces, database latency, connection pool usage, downstream latency, JVM metrics, and Kubernetes resource metrics.
Spring Boot is rarely the real problem in a production incident.
The bigger problem is usually an incorrect assumption about how the application behaves under real-world conditions.
A connection pool that was sufficient during development may fail under production concurrency.
A JPA query that looked harmless may generate thousands of SQL statements.
A retry that helped during testing may become a retry storm during an outage.
A scheduled task that worked perfectly on one machine may execute five times when five application instances are deployed.
A Java application that appears healthy may actually be waiting on a database, downstream service, lock, network call, or Kubernetes resource.
The most important production mindset is simple:
Don't ask only whether the code works.
Ask how the system behaves when everything doesn't go as planned.
That is where production engineering begins.
0 Comments