Thread Pools
On this page 16
Part II — Concurrency · Interview reference
Thread pools bound concurrency, reuse threads, and decouple task submission from execution. Seniors must size pools from first principles, choose queue/rejection policies deliberately, and isolate workloads.
Why pools exist
| Without pool | With pool |
|---|---|
| Unbounded thread spawn | Bounded workers |
| High allocation / stack usage | Reuse |
| No backpressure | Queue + rejection = load shed |
| Hard to meter | Metrics: active, queued, rejected |
Anatomy (Executor framework mental model)
submit(task) → work queue → worker threads → completed / Future / callback
↑
rejection policy when saturated
Core parameters
| Parameter | Meaning |
|---|---|
corePoolSize | Threads kept ready (policy-dependent) |
maxPoolSize | Upper bound of workers |
keepAliveTime | Excess thread retirement |
workQueue | Buffer between submitters and workers |
RejectedExecutionHandler | What happens at capacity |
ThreadFactory | Naming, priority, UncaughtExceptionHandler |
Queue choices (critical tradeoffs)
| Queue | Behavior | Use when |
|---|---|---|
| SynchronousQueue | No capacity; handoff to thread or reject/create | Prefer scaling threads first (cached-style) |
| Bounded LinkedBlockingQueue / ArrayBlockingQueue | Buffer then reject | Controlled latency + backpressure |
| Unbounded LinkedBlockingQueue | Never rejects on queue | Dangerous: memory growth; maxPoolSize ignored in common ThreadPoolExecutor config when core is busy and queue is unbounded |
Interview trap: Unbounded queue + ThreadPoolExecutor → threads stay at core size; tasks pile in memory → OOM / huge latency. Know this interaction.
Rejection policies (Java reference)
| Policy | Effect |
|---|---|
| AbortPolicy | Throw (default) — caller must handle |
| CallerRunsPolicy | Run on submitter thread — natural backpressure |
| DiscardPolicy | Silent drop — usually wrong for business tasks |
| DiscardOldestPolicy | Drop oldest — rare; define semantics carefully |
Java under the hood — ThreadPoolExecutor
execute(task):
if workers < corePoolSize → add worker
else if workQueue.offer(task) → queued
else if workers < maxPoolSize → add worker
else → RejectedExecutionHandler
| Detail | Why it matters |
|---|---|
ctl atomic | Packs run-state (RUNNING/SHUTDOWN/…) + worker count in one int |
Executors.newFixedThreadPool(n) | core=max=n, unbounded LinkedBlockingQueue → never creates >core workers; OOM/latency under load |
newCachedThreadPool | core=0, max=huge, SynchronousQueue — scales threads first; can explode |
ForkJoinPool | Work-stealing deques; used by parallelStream / CompletableFuture default |
ScheduledThreadPoolExecutor | Delay/periodic tasks on a heap of trigger times |
| Virtual-thread executor | newVirtualThreadPerTaskExecutor() — still bound concurrency with semaphores / connection pools |
Prefer constructing ThreadPoolExecutor explicitly (bounded queue + named ThreadFactory + chosen rejection policy) over the Executors convenience factories.
Production case study (high volume)
Context: Checkout API used Executors.newFixedThreadPool(16) (unbounded queue) for payment + tax + fraud fan-out; Black Friday task latency hit minutes.
Why seniors care: Unbounded queue = latent OOM and unbounded queueing delay; maxPoolSize never engaged; overload becomes memory and p99, not clean 503s.
Failure / symptom: Heap climb; task age histograms balloon; clients time out and retry → more queue; rejection count stays zero (trap!).
Resolution: Bounded ArrayBlockingQueue + CallerRuns or Abort→503; queue-age SLO alerts; align pool with Hikari/Redis pools; load-test to saturation.
Seen at / similar to: Countless JVM retail outages; Amazon Builders’ Library load shedding; Netflix concurrency limits / adaptive concurrency.
Sizing — first principles
Let:
- C = number of CPUs
- U = target utilization (0..1)
- W/C = wait time / compute time ratio for tasks
CPU-bound: roughly threads ≈ C (or C+1)
IO-bound: threads ≈ C * (1 + W/C) (Little’s law intuition)
Then validate with load tests: saturation CPU, queue latency, error rate, DB pool limits.
Hard rule: Pool size must be consistent with downstream limits (DB connections, HTTP client pool, rate limits). Oversized app pools stampede the database.
Isolation & bulkheads
| Pattern | Intent |
|---|---|
| Separate pools per dependency | One slow dependency doesn’t consume all threads |
| Separate pools CPU vs blocking IO | Protect latency-sensitive work |
| Cap queue per tenant / API | Noisy neighbor control |
This is the thread analogue of microservices bulkheads.
Production case study (high volume)
Context: SaaS multi-tenant API: one shared pool for Stripe calls, Salesforce sync, and PDF generation; Salesforce degraded and consumed all workers. Why seniors care: No bulkhead → one dependency’s latency is everyone else’s outage; noisy neighbor tenants amplify. Failure / symptom: Checkout timeouts while PDF queue looks “busy”; thread dump: all workers in Salesforce SDK IO. Resolution: Separate pools/semaphores per dependency; tenant concurrency caps; shed non-critical work first; dashboard pool utilization per bulkhead. Seen at / similar to: Netflix Hystrix/Resilience4j bulkheads; AWS cell-based architecture themes; Shopify multi-tenant isolation patterns.
Common pool antipatterns
- Shared unbounded executor for everything
- Deadlock: block on
Futurefrom same pool (starvation & deadlock) - Ignoring
InterruptedException/ poor cancellation - ThreadLocal leaks across reused workers
- Creating pool per request
- Mismatched DB pool: 200 app threads → 10 DB connections → wait storms
Metrics that matter
- Active threads / pool size
- Queue depth & age (time in queue)
- Task execution time histogram
- Rejection count
- Downstream pool wait time
SLOs should include queueing delay, not only handler time.
Virtual threads note
With virtual threads, many servers move toward huge numbers of cheap blocking tasks and smaller dedicated pools for CPU-bound or pinning-prone sections. You still need bounding (semaphores, rate limits, connection pools) — unbounded concurrency remains unsafe.
Production case study (high volume)
Context: After Loom, a logistics tracking API removed platform pool limits; Postgres max connections exhausted in minutes at peak truck-ping volume.
Why seniors care: Virtual threads shift the bottleneck to connection pools and downstream; seniors size admission control, not just executors.
Failure / symptom: HikariPool timeout storms; DB too many connections; app CPU fine.
Resolution: Semaphore ≈ DB pool size; timeout budgets; per-tenant limits; treat rejection as first-class 503 with client backoff.
Seen at / similar to: Early JDK 21 adopters’ postmortems; classic “node-pg pool exhausted” / Rails Puma+DB stories in other stacks.
What interviewers probe
- How do you size a pool for a given service?
- Unbounded queue trap in
ThreadPoolExecutor. - CallerRunsPolicy as backpressure — pros/cons (latency on request thread, risk of reentrancy).
- Design pools for: incoming HTTP, DB, outbound HTTP, CPU-heavy encryption.
- What happens under overload? — shed load explicitly vs melt down.
- Diagnose thread pool exhaustion from symptoms (high queue, timeouts, 503s).
- Bulkheads under dependency failure at millions of req/day.
Senior-level expectation: Pool design as part of capacity planning; isolation; explicit overload behavior.
Pitfalls
- Equating “async” with “no need to bound concurrency”
- Silent discard policies on money/path-critical tasks
- Huge pools to “fix timeouts” without fixing slow dependencies
- Missing thread names → unreadable dumps
- Not propagating MDC/context into workers
- Unbounded queues that never reject (false sense of health)
Cross-references
- Threads
- Thread safety
- Thread starvation and deadlock
- Load balancing vs API gateway — overload & health
- TCP/IP stack — timeouts interact with queueing