Threads

On this page 14

Part II — Concurrency · Interview reference

Threads are the primary unit of concurrent execution on the JVM and most server stacks. Seniors must compare concurrency models, quantify scheduling cost, and know when more threads hurt.


Definitions

TermDefinition
ProcessIsolated address space; heavyweight creation; IPC needed to share
ThreadSchedulable execution path sharing process memory
Fiber / virtual thread / green threadUser-space or lightweight scheduled task multiplexed onto OS threads (e.g. Loom, Go goroutines — different implementations)
ConcurrencyStructure for dealing with multiple tasks (may be interleaved on 1 core)
ParallelismSimultaneous execution on multiple cores
Context switchSaving/restoring registers, stack, TLB effects — not free

Mental model

CPU cores  ⊂  runnable threads  ⊂  existing threads
                 ↑
         blocked on IO / locks / timers
  • Too few threads → underutilize CPU or stall on IO
  • Too many threads → memory (stacks), contention, context-switch thrash
  • Blocking IO model: thread-per-request scales poorly without pools / async / virtual threads

Production case study (high volume)

Context: Legacy thread-per-request payments API behind an ALB; Black Friday open connections jump from 5K to 80K concurrent. Why seniors care: Each platform thread costs stack + scheduler weight; unbounded accept → native OOM / runnable queue explosion; p99 collapses from context switches before CPU is “100% useful work.” Failure / symptom: java.lang.OutOfMemoryError: unable to create new native thread; load average spikes; thread dump shows tens of thousands RUNNABLE/WAITING on JDBC. Resolution: Bound Tomcat/server.tomcat.threads.max; queue + 503 at edge; move to virtual threads with DB pool caps; metrics: active threads, accept queue, pool wait. Seen at / similar to: Classic servlet outages; Netflix/Amazon “shed load” culture; Cloudflare/Fastly edge concurrency lessons applied to origin.


Thread lifecycle (typical)

NEW → RUNNABLE ⇄ BLOCKED/WAITING/TIMED_WAITING → TERMINATED

Know the difference:

  • BLOCKED — waiting to enter a monitor (synchronized)
  • WAITING — wait, join, park without timeout
  • TIMED_WAITING — sleep, poll with timeout

Sharing & isolation

SharedNot shared
Heap objectsCall stack / local vars
Static fieldsThread-local storage (by design)
File descriptors (process)Program counter

Implication: almost all concurrency bugs are about shared mutable heap state → thread safety.


Concurrency models (compare in interviews)

ModelIdeaStrengthsWeaknesses
OS threads + locksShared memorySimple mental model for small critical sectionsContention; scaling limits
Thread poolsBound concurrencyControl resourcesQueueing/latency; sizing hard
Async / event loopFew threads, nonblocking IOHigh connection countsCallback hell; CPU hogs block loop
Actors / message passingNo shared stateIsolationSemantics of delivery; debugging
CSP / channelsExplicit communicationClear pipelinesStill need care for shared refs
Virtual threadsCheap blocking styleScale blocking IO codePinning; not magic for CPU-bound

Cost drivers

  • Stack memory (OS threads: often ~1MB reserved order — platform dependent)
  • Context switches under oversubscription
  • Cache/TLB disruption when bouncing across cores
  • Synchronization (see safety / deadlock chapters)
  • Thread creation — prefer pools for short tasks

Thread-per-request vs async vs virtual threads

ApproachBest fit
Thread-per-request (bounded pool)Traditional servlet apps; moderate concurrency
Reactive / NIOHuge concurrent connections; skilled team
Virtual threadsHigh blocking IO concurrency with simpler code

Interview-ready line: Virtual threads make blocking cheap to schedule; they do not add CPU cores. CPU-bound work still needs ~#cores parallelism and careful pooling.

Production case study (high volume)

Context: Chat/messaging backend (Discord/Slack-class fan-out) and a separate CPU-heavy media-transcode fleet. Why seniors care: One concurrency model does not fit both; event-loop blocked by CPU work takes down all connections on that thread; actors help isolation but need delivery/ops story. Failure / symptom: Event-loop latency histogram goes vertical when JSON parse/crypto runs inline; or virtual-thread explosion without bulkheads hits Redis hard. Resolution: Split pools/models by workload; CPU on fixed pools; IO on virtual threads/async; load-test fan-out separately from compute. Seen at / similar to: Discord gateway discussions; Netty-based gateways (Netflix Zuul eras); Go services at Twitch/Uber with GOMAXPROCS discipline.


Java under the hood

ConceptJDK reality
Platform Thread≈ OS pthread; start() creates native thread
Virtual threadContinuation + mount on carrier (ForkJoinPool by default); Thread.ofVirtual() / Executors.newVirtualThreadPerTaskExecutor()
Thread.StateNEW, RUNNABLE, BLOCKED (monitor enter), WAITING, TIMED_WAITING, TERMINATED — maps to OS schedule + JVM park
ThreadLocalPer-thread ThreadLocalMap (open addressing, weak keys); leaks on pooled threads if not remove()
InterruptFlag + optional InterruptedException on wait/sleep/join/lock interruptibly; swallowing breaks cancellation
DaemonJVM exits when only daemon threads remain — non-daemon workers keep process alive
Dumpsjstack, jcmd Thread.print, JFR — carrier + virtual thread frames differ on Loom

Pinning: long synchronized or JNI on a virtual thread pins its carrier — prefer ReentrantLock and short critical sections. See JVM.

Production case study (high volume)

Context: High-QPS SaaS after Loom migration; ThreadLocal auth context on pooled platform threads historically, then virtual threads. Why seniors care: ThreadLocal on pooled workers leaks classloaders/buffers across requests; on virtual threads, ThreadLocal semantics/cost surprise teams; interrupt swallowing breaks cancel of payment authorizations. Failure / symptom: Metaspace/heap creep; wrong tenant identity on a request (cross-tenant!); stuck cancels during deploy. Resolution: Prefer request-scoped context objects / StructuredTaskScope patterns; always remove() ThreadLocals on platform pools; never swallow interrupts; verify with dumps + tenant-id assertions in canary. Seen at / similar to: Multi-tenant JVM SaaS (Salesforce-like); Tomcat pool ThreadLocal incidents widely documented.


What interviewers probe

  1. Process vs thread isolation and shared memory consequences.
  2. Why unbounded thread creation fails under traffic spikes.
  3. Blocking vs non-blocking and how each saturates resources differently.
  4. How you’d find a stuck thread (thread dump, jstack, flight recorder).
  5. Daemon vs non-daemon (JVM exit behavior).
  6. ThreadLocal use cases (request context) and memory leak pitfalls with pools.
  7. Black Friday concurrency — what you bound first (accept, app, DB).

Senior-level expectation: Map model choice to workload (IO vs CPU), failure modes, and operability (dumps, metrics: active threads, queue depth).


Pitfalls

  • Starting a thread per task without bounds
  • Using ThreadLocal without cleanup on pooled threads
  • Doing heavy CPU work on event-loop threads
  • Swallowing InterruptedException (breaks cancellation)
  • Assuming more threads ⇒ more throughput
  • Virtual threads without downstream admission control

Cross-references

Interview reference — explanation quality and judgment, not syntax memorization.