Microservices Rules

On this page 24

Part IV — System design & architecture · Interview reference

Microservices are an organizational and failure-isolation architecture, not a default virtue. Seniors are graded on when to split, data ownership, consistency, and operability — and on knowing when a modular monolith wins.


Definition

Microservice: independently deployable service owning a business capability and its data, communicating over the network with explicit contracts.

Contrast: modular monolith — strong module boundaries, single deployable.


Rules that matter (senior checklist)

1) Split along business capabilities, not technical layers

  • Prefer “Orders”, “Billing”, “Inventory” over “DaoService” / “UtilsService”
  • Mirror team ownership (Conway’s law — use deliberately)

2) Database per service (logical ownership)

  • No shared mutable tables across services
  • Integration via APIs/events, not foreign keys across service DBs
  • Shared read-only replicas / data products are explicit exceptions with governance

3) Contracts are first-class

  • Versioned APIs/events; compatibility rules
  • Consumer-driven contract tests where appropriate
  • See spec-driven development

4) Design for partial failure

Network calls fail. Timeouts, retries with jitter, idempotency, circuit breakers, bulkheads, degradation paths are mandatory vocabulary.

5) Idempotency for write side-effects

At-least-once delivery is the default reality (HTTP retries, broker redelivery). Use idempotency keys / dedupe stores.

6) Avoid distributed transactions (2PC) as default

Prefer:

  • Saga (choreography/orchestration) with compensating actions
  • Outbox pattern for reliable event publishing with DB commits
  • Orchestration carefully to avoid god-services

7) Observability is part of the architecture

Correlation IDs, metrics (RED/USE), structured logs, traces across service hops. Without this, microservices are undebuggable.

8) Autonomy has a cost ceiling

Each service adds: deploy pipeline, runtime, oncall, versioning, latency. Optimize for team scale and failure isolation, not service count vanity.

Production case study (high volume)

Context: Marketplace platform split “Orders”, “Payments”, “Inventory”, “Shipping” across teams after a modular monolith hit release contention (~tens of M orders/year growing to peak day millions). Why seniors care: Wrong split (by layer: “DaoService”) creates a distributed monolith; shared DB keeps the coupling; seniors justify split with Conway + failure domains + scale profiles. Failure / symptom: Cross-team schema locks; inventory oversell from dual writers; deploys blocked by unrelated services. Resolution: Capability boundaries + DB-per-service; events for propagation; freeze shared tables; strangler migration with metrics on independent deploy rate. Seen at / similar to: Amazon service-oriented evolution; Shopify modular monolith → services selectively; Uber early microservice sprawl cautionary tales.


When microservices are justified

SignalWhy
Multiple teams need independent release cadenceReduce coordination
Different scale/storage profilesScale parts independently
Strong isolation / blast radius needsFault containment
Polyglot necessity provenRarely the first reason

When they are not

SignalPrefer
Small team / early productModular monolith
Tight consistency across a write pathSingle service / single DB transaction
No ops maturity (no tracing, weak CI)Fix platform first
“Nanoservices” by endpointRe-aggregate

Communication styles

StyleUseRisks
Sync request/response (HTTP/gRPC)Queries, user-facing need immediate resultCoupling, cascading failures
Async events/messagesState change propagation, decouplingEventual consistency complexity
Batch / streamsAnalytics, syncFreshness

Default rule: don’t make a sync call on a write path if an event can finish the workflow — but don’t create event spaghetti either.

Production case study (high volume)

Context: Checkout wrote order → sync called payments → sync called inventory → sync called email; any hop timeout failed the purchase after card auth. Why seniors care: Sync write chains amplify p99 and create distributed failure coupling; money paths need idempotent async completion with clear UX. Failure / symptom: Elevated checkout failures when email/SMS provider slows; nested timeout budgets exceeded; double charges on client retry. Resolution: Auth payment with idempotency; persist order; outbox events for inventory/email; sagas with compensations; sync only for user-visible authorization step. Seen at / similar to: Amazon checkout/event stories; Stripe payment intents; Airbnb/Uber order state machines.


Data & consistency

ApproachNotes
Single-service ACIDBest when possible
SagaLong-running business tx across services
CQRS / read modelsScale reads; accept lag
Shared kernel / duplicate dataCache with clear invalidation ownership

Be ready to explain read-your-writes, causal consistency, and user-visible lag.

Production case study (high volume)

Context: Social feed CQRS: write path to graph service; read models via Kafka → Elasticsearch (~billions of events/day ingest). Why seniors care: Users notice lag; dual-write without outbox loses events; seniors set freshness SLOs and repair paths. Failure / symptom: Missing posts in feed; consumer lag hours; inconsistent counts across regions. Resolution: Outbox/CDC; lag alerts as product SLO; rebuild-from-log; read-your-writes via sticky read or version tokens. Seen at / similar to: Twitter/X fan-out; LinkedIn feed; Netflix recommendation pipelines; Facebook TAO-style read paths (conceptual).


Resilience patterns (name + when)

PatternPurpose
TimeoutBound wait
Retry + jitterTransient faults only; idempotent ops
Circuit breakerFail fast when dependency sick
BulkheadIsolate pools/resources
Rate limit / load shedProtect core capacity
HedgingDuplicate requests carefully (costly)

Align with thread pools and LB/gateway.

Production case study (high volume)

Context: Black Friday: fraud vendor latency 2s→30s; without bulkheads, all checkout threads blocked; circuit breaker absent. Why seniors care: One dependency’s brownout is a site-wide outage without isolation; seniors design degradation (skip non-critical fraud tier) explicitly. Failure / symptom: Thread pool exhaustion; 503s on cart; cascading timeouts upstream. Resolution: Bulkhead + circuit breaker + deadline propagation; degrade to rules-only fraud; shed load at gateway; game-day the vendor failure. Seen at / similar to: Netflix Hystrix → Resilience4j; Amazon “avoid cascading failures”; Walmart/Target peak retail postmortems (public themes).


Java under the hood

ConcernCommon Java stack
Sync RPCjava.net.http.HttpClient, Spring WebClient, gRPC stubs
Timeouts / retry / CBResilience4j (TimeLimiter, Retry, CircuitBreaker) or Failsafe
BulkheadResilience4j semaphore vs thread-pool bulkhead; separate Executors per dependency
IdempotencyDB unique constraint / Redis SETNX keyed by idempotency header
ObservabilityMicrometer metrics; OpenTelemetry Java agent — trace propagation (traceparent)
ContractsOpenAPI clients, protobuf stubs — see spec-driven

Judgment: framework choice is secondary to timeout budgets, isolation, and idempotent writes.


Deployment & versioning rules

  • Blue/green or canary with health + metrics gates
  • Expand/contract for schema and API changes
  • Backward-compatible event evolution (add fields; don’t rename casually)

What interviewers probe

  1. How do you decide service boundaries? (capability + data + change frequency)
  2. Shared database across services — why is it bad? Exceptions?
  3. Saga vs 2PC
  4. Design order → payment → inventory with failure at each step
  5. Fan-out failure and cascading timeout storms — mitigations
  6. “Would you start with microservices?” — usually no for greenfield small teams; explain migration path
  7. Contract testing / API versioning strategy
  8. Cascading failure from one vendor under peak — what you cut first.

Senior-level expectation: Judgment and tradeoffs; operational realism; consistency story that matches UX.


Pitfalls

  • Distributed monolith (sync mesh of calls + shared DB)
  • Chatty interfaces (N+1 RPCs)
  • Ignoring idempotency
  • Gateway or orchestration service becomes new monolith
  • Per-service snowflake infra without platform
  • Sync email/SMS on the payment critical path
  • Dual-write to DB and Kafka without outbox

Cross-references

Interview reference — explanation quality and judgment, not syntax memorization.