Thread Safety

On this page 18

Part II — Concurrency · Interview reference

Thread safety means shared mutable state remains correct under concurrent execution. Seniors must explain visibility and atomicity separately, choose synchronization primitives deliberately, and design to minimize shared mutability.


Definitions

TermPrecise meaning
Race conditionCorrectness depends on timing / interleaving
Data raceConflicting accesses to same location without happens-before (JMM); undefined behavior in C++)
AtomicityOperation appears indivisible
VisibilityUpdates become visible to other threads as required by the memory model
OrderingReordering constraints between operations
Thread-safeClass/component behaves correctly when used by multiple threads per its contract
ImmutableState cannot change after construction → inherently thread-safe to share

Critical distinction: “It works on my machine” under low contention does not prove safety.


Happens-before (Java-centric mental model)

Establish ordering with:

  • synchronized unlock → subsequent lock of same monitor
  • volatile write → subsequent read of same variable
  • Thread start / successful join
  • Concurrent utilities with documented HB (e.g. BlockingQueue handoff)

Without HB, compilers/CPUs may reorder; reads may see stale cached values forever.


Failure modes

BugSymptom
Lost updateCounters skip; inventory wrong
Dirty read / stale readConfig flip invisible
Check-then-actif (!map.contains) put races
Unsafe publicationPartially constructed object escaped
Deadlock / livelockSee starvation & deadlock chapter

Production case study (high volume)

Context: Inventory reservation for e-commerce SKUs (~millions of checkouts/day) with if (stock > 0) stock-- style updates and a “feature flag” boolean published without volatile/locking. Why seniors care: Lost updates oversell inventory (revenue + trust); stale flag reads mean half the fleet runs old pricing for minutes; blast radius is money, not a flaky unit test. Failure / symptom: Oversell spikes under flash sale; config change “didn’t apply” until bounce; jcstress-style races only reproduce at high QPS. Resolution: Atomic conditional updates / DB conditional writes / Lua in Redis; volatile or immutable republish for config snapshots; load-test with contended SKUs (hot keys). Seen at / similar to: Amazon/Shopify inventory; ticketmaster-style onsale; Redis DECR/WATCH patterns; LinkedIn/Netflix dynamic config libraries.


Tools & techniques (decision table)

TechniqueProvidesCost / notes
Immutability / copy-on-writeSafety by designAllocation; versioning
Confinement (stack / thread-local / single owner)No sharingLimits scalability patterns
synchronized / ReentrantLockMutual exclusion + visibilityContention; deadlock risk
ReadWriteLock / StampedLockRead scalabilityWrite bias / complexity
volatileVisibility + atomicity of reference write / narrow typesNot compound actions
Atomics (AtomicInteger, CAS)Lock-free updates to single varsABA; retry loops
Concurrent collectionsThread-safe containersCompound sequences still need external sync
Actors / message passingAvoid shared mutabilityDifferent programming model

Patterns seniors must know

Safe initialization / publication

  • Static initializers, final fields (initialization safety), volatile holder, enums-as-singleton
  • Double-checked locking: volatile required on the field

Check-then-act → single atomic op

// unsafe
if (!map.containsKey(k)) map.put(k, v);

// safer
map.putIfAbsent(k, v);
// or compute / merge APIs

Atomic compound state

When invariants span multiple fields, one lock (or one atomic structure / immutable snapshot) must protect the invariant — not “volatile each field.”

Lock ordering

Document and globally order locks → prevent deadlock (starvation & deadlock).

Production case study (high volume)

Context: Idempotent payment intent API: clients retry; servers did containsKey + put on a local ConcurrentHashMap before charging Stripe/Adyen. Why seniors care: CHM does not make check-then-act safe; double charge under at-least-once retries is a senior-level incident. Failure / symptom: Duplicate captures for same idempotency key during LB retries; metrics show two side effects per key. Resolution: putIfAbsent / DB unique constraint on idempotency key; single-flight; treat map as cache of outcomes, not the source of truth. Seen at / similar to: Stripe Idempotency-Key header model; AWS API idempotency; Uber/Airbnb payment platforms.


Java under the hood

PrimitiveImplementation sketch
synchronizedmonitorenter / monitorexit; mark word thin → inflated monitor under contention
volatileStore-load barriers per JMM; visibility + atomicity of the write itself only
ReentrantLockAQS CLH-style wait queue; fair vs unfair; tryLock, interruptible lock
ReadWriteLock / StampedLockShared/exclusive modes; StampedLock optimistic read (validate stamp)
AtomicsAtomicInteger etc. CAS via VarHandle; LongAdder / LongAccumulator stripe cells
ConcurrentHashMapTable of bins; CAS empty bin; lock bin head on contended updates; treeify long bins; CounterCells for size; no null keys/values; iterators weakly consistent
CopyOnWriteArrayListFresh array copy on mutate; snapshot iterators
ConcurrentLinkedQueueMichael–Scott lock-free linked queue (CAS)

Critical: CHM makes single ops thread-safe; containsKey + put is still a race — use putIfAbsent / compute. Compound multi-map invariants need an external lock or redesign.


Performance vs safety

  • Prefer smaller critical sections but not so small that atomicity breaks
  • Prefer stripíng / sharding over one global lock when contended
  • Prefer immutable snapshots for read-mostly configs
  • Measure: lock contention metrics, CPU in park/futex, throughput collapse under load

False sharing on atomics → CPU architecture.

Production case study (high volume)

Context: Global metrics lock and a single synchronized on a shared “rate limiter state” object in an ad-bid path (~hundreds of K RPS). Why seniors care: Correctness without striping becomes a throughput ceiling; seniors must quantify contention (park/futex, JFR lock profiles) and redesign, not “add servers.” Failure / symptom: Throughput flat while CPU cores idle waiting; lock profiles show one hot monitor; p99 rises with QPS. Resolution: Shard limiters by key; LongAdder; lock-free where invariants allow; immutable config snapshots for read-mostly data. Seen at / similar to: Guava/Resilience4j rate limiters; Cloudflare edge limiting; Twitter early “fail whale” lock/contention lore as cautionary tale.


Testing & detection

  • Stress tests with high parallelism; Torché / jcstress-style thinking
  • Thread sanitizers (C/C++); JVM: careful review + concurrency test libraries
  • Production: thread dumps, Flight Recorder lock profiles

Never claim “we tested with 2 threads once.”


What interviewers probe

  1. volatile vs synchronized — when each suffices.
  2. Why ConcurrentHashMap alone doesn’t make multi-step business logic safe.
  3. Implement a thread-safe counter / sequence — lock vs atomic tradeoffs.
  4. Explain a real race you’ve fixed; root cause in visibility vs atomicity terms.
  5. Immutable design for a shared cache entry.
  6. Memory model question: can a thread loop forever seeing flag == false without volatile?
  7. Idempotency + races under payment retries at scale.

Senior-level expectation: Design that reduces shared mutability; precise vocabulary; awareness of language memory model.


Pitfalls

  • Synchronizing only writers (readers need visibility too)
  • Locking on this of a publicly accessible instance
  • Using ConcurrentHashMap size/iteration as consistent snapshots without care
  • Non-atomic “read modify write” on volatiles (count++)
  • Escaping this from constructors
  • Using CHM check-then-act for money side effects
  • One global lock on a multi-hundred-K-RPS path

Cross-references

Interview reference — explanation quality and judgment, not syntax memorization.