Threads
On this page 14
Part II — Concurrency · Interview reference
Threads are the primary unit of concurrent execution on the JVM and most server stacks. Seniors must compare concurrency models, quantify scheduling cost, and know when more threads hurt.
Definitions
| Term | Definition |
|---|---|
| Process | Isolated address space; heavyweight creation; IPC needed to share |
| Thread | Schedulable execution path sharing process memory |
| Fiber / virtual thread / green thread | User-space or lightweight scheduled task multiplexed onto OS threads (e.g. Loom, Go goroutines — different implementations) |
| Concurrency | Structure for dealing with multiple tasks (may be interleaved on 1 core) |
| Parallelism | Simultaneous execution on multiple cores |
| Context switch | Saving/restoring registers, stack, TLB effects — not free |
Mental model
CPU cores ⊂ runnable threads ⊂ existing threads
↑
blocked on IO / locks / timers
- Too few threads → underutilize CPU or stall on IO
- Too many threads → memory (stacks), contention, context-switch thrash
- Blocking IO model: thread-per-request scales poorly without pools / async / virtual threads
Production case study (high volume)
Context: Legacy thread-per-request payments API behind an ALB; Black Friday open connections jump from 5K to 80K concurrent.
Why seniors care: Each platform thread costs stack + scheduler weight; unbounded accept → native OOM / runnable queue explosion; p99 collapses from context switches before CPU is “100% useful work.”
Failure / symptom: java.lang.OutOfMemoryError: unable to create new native thread; load average spikes; thread dump shows tens of thousands RUNNABLE/WAITING on JDBC.
Resolution: Bound Tomcat/server.tomcat.threads.max; queue + 503 at edge; move to virtual threads with DB pool caps; metrics: active threads, accept queue, pool wait.
Seen at / similar to: Classic servlet outages; Netflix/Amazon “shed load” culture; Cloudflare/Fastly edge concurrency lessons applied to origin.
Thread lifecycle (typical)
NEW → RUNNABLE ⇄ BLOCKED/WAITING/TIMED_WAITING → TERMINATED
Know the difference:
- BLOCKED — waiting to enter a monitor (
synchronized) - WAITING —
wait,join,parkwithout timeout - TIMED_WAITING — sleep, poll with timeout
Sharing & isolation
| Shared | Not shared |
|---|---|
| Heap objects | Call stack / local vars |
| Static fields | Thread-local storage (by design) |
| File descriptors (process) | Program counter |
Implication: almost all concurrency bugs are about shared mutable heap state → thread safety.
Concurrency models (compare in interviews)
| Model | Idea | Strengths | Weaknesses |
|---|---|---|---|
| OS threads + locks | Shared memory | Simple mental model for small critical sections | Contention; scaling limits |
| Thread pools | Bound concurrency | Control resources | Queueing/latency; sizing hard |
| Async / event loop | Few threads, nonblocking IO | High connection counts | Callback hell; CPU hogs block loop |
| Actors / message passing | No shared state | Isolation | Semantics of delivery; debugging |
| CSP / channels | Explicit communication | Clear pipelines | Still need care for shared refs |
| Virtual threads | Cheap blocking style | Scale blocking IO code | Pinning; not magic for CPU-bound |
Cost drivers
- Stack memory (OS threads: often ~1MB reserved order — platform dependent)
- Context switches under oversubscription
- Cache/TLB disruption when bouncing across cores
- Synchronization (see safety / deadlock chapters)
- Thread creation — prefer pools for short tasks
Thread-per-request vs async vs virtual threads
| Approach | Best fit |
|---|---|
| Thread-per-request (bounded pool) | Traditional servlet apps; moderate concurrency |
| Reactive / NIO | Huge concurrent connections; skilled team |
| Virtual threads | High blocking IO concurrency with simpler code |
Interview-ready line: Virtual threads make blocking cheap to schedule; they do not add CPU cores. CPU-bound work still needs ~#cores parallelism and careful pooling.
Production case study (high volume)
Context: Chat/messaging backend (Discord/Slack-class fan-out) and a separate CPU-heavy media-transcode fleet. Why seniors care: One concurrency model does not fit both; event-loop blocked by CPU work takes down all connections on that thread; actors help isolation but need delivery/ops story. Failure / symptom: Event-loop latency histogram goes vertical when JSON parse/crypto runs inline; or virtual-thread explosion without bulkheads hits Redis hard. Resolution: Split pools/models by workload; CPU on fixed pools; IO on virtual threads/async; load-test fan-out separately from compute. Seen at / similar to: Discord gateway discussions; Netty-based gateways (Netflix Zuul eras); Go services at Twitch/Uber with GOMAXPROCS discipline.
Java under the hood
| Concept | JDK reality |
|---|---|
Platform Thread | ≈ OS pthread; start() creates native thread |
| Virtual thread | Continuation + mount on carrier (ForkJoinPool by default); Thread.ofVirtual() / Executors.newVirtualThreadPerTaskExecutor() |
Thread.State | NEW, RUNNABLE, BLOCKED (monitor enter), WAITING, TIMED_WAITING, TERMINATED — maps to OS schedule + JVM park |
ThreadLocal | Per-thread ThreadLocalMap (open addressing, weak keys); leaks on pooled threads if not remove() |
| Interrupt | Flag + optional InterruptedException on wait/sleep/join/lock interruptibly; swallowing breaks cancellation |
| Daemon | JVM exits when only daemon threads remain — non-daemon workers keep process alive |
| Dumps | jstack, jcmd Thread.print, JFR — carrier + virtual thread frames differ on Loom |
Pinning: long synchronized or JNI on a virtual thread pins its carrier — prefer ReentrantLock and short critical sections. See JVM.
Production case study (high volume)
Context: High-QPS SaaS after Loom migration; ThreadLocal auth context on pooled platform threads historically, then virtual threads.
Why seniors care: ThreadLocal on pooled workers leaks classloaders/buffers across requests; on virtual threads, ThreadLocal semantics/cost surprise teams; interrupt swallowing breaks cancel of payment authorizations.
Failure / symptom: Metaspace/heap creep; wrong tenant identity on a request (cross-tenant!); stuck cancels during deploy.
Resolution: Prefer request-scoped context objects / StructuredTaskScope patterns; always remove() ThreadLocals on platform pools; never swallow interrupts; verify with dumps + tenant-id assertions in canary.
Seen at / similar to: Multi-tenant JVM SaaS (Salesforce-like); Tomcat pool ThreadLocal incidents widely documented.
What interviewers probe
- Process vs thread isolation and shared memory consequences.
- Why unbounded thread creation fails under traffic spikes.
- Blocking vs non-blocking and how each saturates resources differently.
- How you’d find a stuck thread (thread dump, jstack, flight recorder).
- Daemon vs non-daemon (JVM exit behavior).
- ThreadLocal use cases (request context) and memory leak pitfalls with pools.
- Black Friday concurrency — what you bound first (accept, app, DB).
Senior-level expectation: Map model choice to workload (IO vs CPU), failure modes, and operability (dumps, metrics: active threads, queue depth).
Pitfalls
- Starting a thread per task without bounds
- Using
ThreadLocalwithout cleanup on pooled threads - Doing heavy CPU work on event-loop threads
- Swallowing
InterruptedException(breaks cancellation) - Assuming more threads ⇒ more throughput
- Virtual threads without downstream admission control