Thread Safety
On this page 18
Part II — Concurrency · Interview reference
Thread safety means shared mutable state remains correct under concurrent execution. Seniors must explain visibility and atomicity separately, choose synchronization primitives deliberately, and design to minimize shared mutability.
Definitions
| Term | Precise meaning |
|---|---|
| Race condition | Correctness depends on timing / interleaving |
| Data race | Conflicting accesses to same location without happens-before (JMM); undefined behavior in C++) |
| Atomicity | Operation appears indivisible |
| Visibility | Updates become visible to other threads as required by the memory model |
| Ordering | Reordering constraints between operations |
| Thread-safe | Class/component behaves correctly when used by multiple threads per its contract |
| Immutable | State cannot change after construction → inherently thread-safe to share |
Critical distinction: “It works on my machine” under low contention does not prove safety.
Happens-before (Java-centric mental model)
Establish ordering with:
synchronizedunlock → subsequent lock of same monitorvolatilewrite → subsequent read of same variable- Thread
start/ successfuljoin - Concurrent utilities with documented HB (e.g.
BlockingQueuehandoff)
Without HB, compilers/CPUs may reorder; reads may see stale cached values forever.
Failure modes
| Bug | Symptom |
|---|---|
| Lost update | Counters skip; inventory wrong |
| Dirty read / stale read | Config flip invisible |
| Check-then-act | if (!map.contains) put races |
| Unsafe publication | Partially constructed object escaped |
| Deadlock / livelock | See starvation & deadlock chapter |
Production case study (high volume)
Context: Inventory reservation for e-commerce SKUs (~millions of checkouts/day) with if (stock > 0) stock-- style updates and a “feature flag” boolean published without volatile/locking.
Why seniors care: Lost updates oversell inventory (revenue + trust); stale flag reads mean half the fleet runs old pricing for minutes; blast radius is money, not a flaky unit test.
Failure / symptom: Oversell spikes under flash sale; config change “didn’t apply” until bounce; jcstress-style races only reproduce at high QPS.
Resolution: Atomic conditional updates / DB conditional writes / Lua in Redis; volatile or immutable republish for config snapshots; load-test with contended SKUs (hot keys).
Seen at / similar to: Amazon/Shopify inventory; ticketmaster-style onsale; Redis DECR/WATCH patterns; LinkedIn/Netflix dynamic config libraries.
Tools & techniques (decision table)
| Technique | Provides | Cost / notes |
|---|---|---|
| Immutability / copy-on-write | Safety by design | Allocation; versioning |
| Confinement (stack / thread-local / single owner) | No sharing | Limits scalability patterns |
synchronized / ReentrantLock | Mutual exclusion + visibility | Contention; deadlock risk |
| ReadWriteLock / StampedLock | Read scalability | Write bias / complexity |
volatile | Visibility + atomicity of reference write / narrow types | Not compound actions |
Atomics (AtomicInteger, CAS) | Lock-free updates to single vars | ABA; retry loops |
| Concurrent collections | Thread-safe containers | Compound sequences still need external sync |
| Actors / message passing | Avoid shared mutability | Different programming model |
Patterns seniors must know
Safe initialization / publication
- Static initializers,
finalfields (initialization safety), volatile holder, enums-as-singleton - Double-checked locking:
volatilerequired on the field
Check-then-act → single atomic op
// unsafe
if (!map.containsKey(k)) map.put(k, v);
// safer
map.putIfAbsent(k, v);
// or compute / merge APIs
Atomic compound state
When invariants span multiple fields, one lock (or one atomic structure / immutable snapshot) must protect the invariant — not “volatile each field.”
Lock ordering
Document and globally order locks → prevent deadlock (starvation & deadlock).
Production case study (high volume)
Context: Idempotent payment intent API: clients retry; servers did containsKey + put on a local ConcurrentHashMap before charging Stripe/Adyen.
Why seniors care: CHM does not make check-then-act safe; double charge under at-least-once retries is a senior-level incident.
Failure / symptom: Duplicate captures for same idempotency key during LB retries; metrics show two side effects per key.
Resolution: putIfAbsent / DB unique constraint on idempotency key; single-flight; treat map as cache of outcomes, not the source of truth.
Seen at / similar to: Stripe Idempotency-Key header model; AWS API idempotency; Uber/Airbnb payment platforms.
Java under the hood
| Primitive | Implementation sketch |
|---|---|
synchronized | monitorenter / monitorexit; mark word thin → inflated monitor under contention |
volatile | Store-load barriers per JMM; visibility + atomicity of the write itself only |
ReentrantLock | AQS CLH-style wait queue; fair vs unfair; tryLock, interruptible lock |
ReadWriteLock / StampedLock | Shared/exclusive modes; StampedLock optimistic read (validate stamp) |
| Atomics | AtomicInteger etc. CAS via VarHandle; LongAdder / LongAccumulator stripe cells |
ConcurrentHashMap | Table of bins; CAS empty bin; lock bin head on contended updates; treeify long bins; CounterCells for size; no null keys/values; iterators weakly consistent |
CopyOnWriteArrayList | Fresh array copy on mutate; snapshot iterators |
ConcurrentLinkedQueue | Michael–Scott lock-free linked queue (CAS) |
Critical: CHM makes single ops thread-safe; containsKey + put is still a race — use putIfAbsent / compute. Compound multi-map invariants need an external lock or redesign.
Performance vs safety
- Prefer smaller critical sections but not so small that atomicity breaks
- Prefer stripíng / sharding over one global lock when contended
- Prefer immutable snapshots for read-mostly configs
- Measure: lock contention metrics, CPU in
park/futex, throughput collapse under load
False sharing on atomics → CPU architecture.
Production case study (high volume)
Context: Global metrics lock and a single synchronized on a shared “rate limiter state” object in an ad-bid path (~hundreds of K RPS).
Why seniors care: Correctness without striping becomes a throughput ceiling; seniors must quantify contention (park/futex, JFR lock profiles) and redesign, not “add servers.”
Failure / symptom: Throughput flat while CPU cores idle waiting; lock profiles show one hot monitor; p99 rises with QPS.
Resolution: Shard limiters by key; LongAdder; lock-free where invariants allow; immutable config snapshots for read-mostly data.
Seen at / similar to: Guava/Resilience4j rate limiters; Cloudflare edge limiting; Twitter early “fail whale” lock/contention lore as cautionary tale.
Testing & detection
- Stress tests with high parallelism; Torché / jcstress-style thinking
- Thread sanitizers (C/C++); JVM: careful review + concurrency test libraries
- Production: thread dumps, Flight Recorder lock profiles
Never claim “we tested with 2 threads once.”
What interviewers probe
- volatile vs synchronized — when each suffices.
- Why ConcurrentHashMap alone doesn’t make multi-step business logic safe.
- Implement a thread-safe counter / sequence — lock vs atomic tradeoffs.
- Explain a real race you’ve fixed; root cause in visibility vs atomicity terms.
- Immutable design for a shared cache entry.
- Memory model question: can a thread loop forever seeing
flag == falsewithout volatile? - Idempotency + races under payment retries at scale.
Senior-level expectation: Design that reduces shared mutability; precise vocabulary; awareness of language memory model.
Pitfalls
- Synchronizing only writers (readers need visibility too)
- Locking on
thisof a publicly accessible instance - Using
ConcurrentHashMapsize/iteration as consistent snapshots without care - Non-atomic “read modify write” on volatiles (
count++) - Escaping
thisfrom constructors - Using CHM check-then-act for money side effects
- One global lock on a multi-hundred-K-RPS path