Thread Pools

On this page 16

Part II — Concurrency · Interview reference

Thread pools bound concurrency, reuse threads, and decouple task submission from execution. Seniors must size pools from first principles, choose queue/rejection policies deliberately, and isolate workloads.


Why pools exist

Without poolWith pool
Unbounded thread spawnBounded workers
High allocation / stack usageReuse
No backpressureQueue + rejection = load shed
Hard to meterMetrics: active, queued, rejected

Anatomy (Executor framework mental model)

submit(task) → work queue → worker threads → completed / Future / callback
                   ↑
             rejection policy when saturated

Core parameters

ParameterMeaning
corePoolSizeThreads kept ready (policy-dependent)
maxPoolSizeUpper bound of workers
keepAliveTimeExcess thread retirement
workQueueBuffer between submitters and workers
RejectedExecutionHandlerWhat happens at capacity
ThreadFactoryNaming, priority, UncaughtExceptionHandler

Queue choices (critical tradeoffs)

QueueBehaviorUse when
SynchronousQueueNo capacity; handoff to thread or reject/createPrefer scaling threads first (cached-style)
Bounded LinkedBlockingQueue / ArrayBlockingQueueBuffer then rejectControlled latency + backpressure
Unbounded LinkedBlockingQueueNever rejects on queueDangerous: memory growth; maxPoolSize ignored in common ThreadPoolExecutor config when core is busy and queue is unbounded

Interview trap: Unbounded queue + ThreadPoolExecutor → threads stay at core size; tasks pile in memory → OOM / huge latency. Know this interaction.


Rejection policies (Java reference)

PolicyEffect
AbortPolicyThrow (default) — caller must handle
CallerRunsPolicyRun on submitter thread — natural backpressure
DiscardPolicySilent drop — usually wrong for business tasks
DiscardOldestPolicyDrop oldest — rare; define semantics carefully

Java under the hood — ThreadPoolExecutor

execute(task):
  if workers < corePoolSize → add worker
  else if workQueue.offer(task) → queued
  else if workers < maxPoolSize → add worker
  else → RejectedExecutionHandler
DetailWhy it matters
ctl atomicPacks run-state (RUNNING/SHUTDOWN/…) + worker count in one int
Executors.newFixedThreadPool(n)core=max=n, unbounded LinkedBlockingQueue → never creates >core workers; OOM/latency under load
newCachedThreadPoolcore=0, max=huge, SynchronousQueue — scales threads first; can explode
ForkJoinPoolWork-stealing deques; used by parallelStream / CompletableFuture default
ScheduledThreadPoolExecutorDelay/periodic tasks on a heap of trigger times
Virtual-thread executornewVirtualThreadPerTaskExecutor() — still bound concurrency with semaphores / connection pools

Prefer constructing ThreadPoolExecutor explicitly (bounded queue + named ThreadFactory + chosen rejection policy) over the Executors convenience factories.

Production case study (high volume)

Context: Checkout API used Executors.newFixedThreadPool(16) (unbounded queue) for payment + tax + fraud fan-out; Black Friday task latency hit minutes. Why seniors care: Unbounded queue = latent OOM and unbounded queueing delay; maxPoolSize never engaged; overload becomes memory and p99, not clean 503s. Failure / symptom: Heap climb; task age histograms balloon; clients time out and retry → more queue; rejection count stays zero (trap!). Resolution: Bounded ArrayBlockingQueue + CallerRuns or Abort→503; queue-age SLO alerts; align pool with Hikari/Redis pools; load-test to saturation. Seen at / similar to: Countless JVM retail outages; Amazon Builders’ Library load shedding; Netflix concurrency limits / adaptive concurrency.


Sizing — first principles

Let:

  • C = number of CPUs
  • U = target utilization (0..1)
  • W/C = wait time / compute time ratio for tasks

CPU-bound: roughly threads ≈ C (or C+1)

IO-bound: threads ≈ C * (1 + W/C) (Little’s law intuition)

Then validate with load tests: saturation CPU, queue latency, error rate, DB pool limits.

Hard rule: Pool size must be consistent with downstream limits (DB connections, HTTP client pool, rate limits). Oversized app pools stampede the database.


Isolation & bulkheads

PatternIntent
Separate pools per dependencyOne slow dependency doesn’t consume all threads
Separate pools CPU vs blocking IOProtect latency-sensitive work
Cap queue per tenant / APINoisy neighbor control

This is the thread analogue of microservices bulkheads.

Production case study (high volume)

Context: SaaS multi-tenant API: one shared pool for Stripe calls, Salesforce sync, and PDF generation; Salesforce degraded and consumed all workers. Why seniors care: No bulkhead → one dependency’s latency is everyone else’s outage; noisy neighbor tenants amplify. Failure / symptom: Checkout timeouts while PDF queue looks “busy”; thread dump: all workers in Salesforce SDK IO. Resolution: Separate pools/semaphores per dependency; tenant concurrency caps; shed non-critical work first; dashboard pool utilization per bulkhead. Seen at / similar to: Netflix Hystrix/Resilience4j bulkheads; AWS cell-based architecture themes; Shopify multi-tenant isolation patterns.


Common pool antipatterns

  1. Shared unbounded executor for everything
  2. Deadlock: block on Future from same pool (starvation & deadlock)
  3. Ignoring InterruptedException / poor cancellation
  4. ThreadLocal leaks across reused workers
  5. Creating pool per request
  6. Mismatched DB pool: 200 app threads → 10 DB connections → wait storms

Metrics that matter

  • Active threads / pool size
  • Queue depth & age (time in queue)
  • Task execution time histogram
  • Rejection count
  • Downstream pool wait time

SLOs should include queueing delay, not only handler time.


Virtual threads note

With virtual threads, many servers move toward huge numbers of cheap blocking tasks and smaller dedicated pools for CPU-bound or pinning-prone sections. You still need bounding (semaphores, rate limits, connection pools) — unbounded concurrency remains unsafe.

Production case study (high volume)

Context: After Loom, a logistics tracking API removed platform pool limits; Postgres max connections exhausted in minutes at peak truck-ping volume. Why seniors care: Virtual threads shift the bottleneck to connection pools and downstream; seniors size admission control, not just executors. Failure / symptom: HikariPool timeout storms; DB too many connections; app CPU fine. Resolution: Semaphore ≈ DB pool size; timeout budgets; per-tenant limits; treat rejection as first-class 503 with client backoff. Seen at / similar to: Early JDK 21 adopters’ postmortems; classic “node-pg pool exhausted” / Rails Puma+DB stories in other stacks.


What interviewers probe

  1. How do you size a pool for a given service?
  2. Unbounded queue trap in ThreadPoolExecutor.
  3. CallerRunsPolicy as backpressure — pros/cons (latency on request thread, risk of reentrancy).
  4. Design pools for: incoming HTTP, DB, outbound HTTP, CPU-heavy encryption.
  5. What happens under overload? — shed load explicitly vs melt down.
  6. Diagnose thread pool exhaustion from symptoms (high queue, timeouts, 503s).
  7. Bulkheads under dependency failure at millions of req/day.

Senior-level expectation: Pool design as part of capacity planning; isolation; explicit overload behavior.


Pitfalls

  • Equating “async” with “no need to bound concurrency”
  • Silent discard policies on money/path-critical tasks
  • Huge pools to “fix timeouts” without fixing slow dependencies
  • Missing thread names → unreadable dumps
  • Not propagating MDC/context into workers
  • Unbounded queues that never reject (false sense of health)

Cross-references

Interview reference — explanation quality and judgment, not syntax memorization.