Microservices Rules
On this page 24
Part IV — System design & architecture · Interview reference
Microservices are an organizational and failure-isolation architecture, not a default virtue. Seniors are graded on when to split, data ownership, consistency, and operability — and on knowing when a modular monolith wins.
Definition
Microservice: independently deployable service owning a business capability and its data, communicating over the network with explicit contracts.
Contrast: modular monolith — strong module boundaries, single deployable.
Rules that matter (senior checklist)
1) Split along business capabilities, not technical layers
- Prefer “Orders”, “Billing”, “Inventory” over “DaoService” / “UtilsService”
- Mirror team ownership (Conway’s law — use deliberately)
2) Database per service (logical ownership)
- No shared mutable tables across services
- Integration via APIs/events, not foreign keys across service DBs
- Shared read-only replicas / data products are explicit exceptions with governance
3) Contracts are first-class
- Versioned APIs/events; compatibility rules
- Consumer-driven contract tests where appropriate
- See spec-driven development
4) Design for partial failure
Network calls fail. Timeouts, retries with jitter, idempotency, circuit breakers, bulkheads, degradation paths are mandatory vocabulary.
5) Idempotency for write side-effects
At-least-once delivery is the default reality (HTTP retries, broker redelivery). Use idempotency keys / dedupe stores.
6) Avoid distributed transactions (2PC) as default
Prefer:
- Saga (choreography/orchestration) with compensating actions
- Outbox pattern for reliable event publishing with DB commits
- Orchestration carefully to avoid god-services
7) Observability is part of the architecture
Correlation IDs, metrics (RED/USE), structured logs, traces across service hops. Without this, microservices are undebuggable.
8) Autonomy has a cost ceiling
Each service adds: deploy pipeline, runtime, oncall, versioning, latency. Optimize for team scale and failure isolation, not service count vanity.
Production case study (high volume)
Context: Marketplace platform split “Orders”, “Payments”, “Inventory”, “Shipping” across teams after a modular monolith hit release contention (~tens of M orders/year growing to peak day millions). Why seniors care: Wrong split (by layer: “DaoService”) creates a distributed monolith; shared DB keeps the coupling; seniors justify split with Conway + failure domains + scale profiles. Failure / symptom: Cross-team schema locks; inventory oversell from dual writers; deploys blocked by unrelated services. Resolution: Capability boundaries + DB-per-service; events for propagation; freeze shared tables; strangler migration with metrics on independent deploy rate. Seen at / similar to: Amazon service-oriented evolution; Shopify modular monolith → services selectively; Uber early microservice sprawl cautionary tales.
When microservices are justified
| Signal | Why |
|---|---|
| Multiple teams need independent release cadence | Reduce coordination |
| Different scale/storage profiles | Scale parts independently |
| Strong isolation / blast radius needs | Fault containment |
| Polyglot necessity proven | Rarely the first reason |
When they are not
| Signal | Prefer |
|---|---|
| Small team / early product | Modular monolith |
| Tight consistency across a write path | Single service / single DB transaction |
| No ops maturity (no tracing, weak CI) | Fix platform first |
| “Nanoservices” by endpoint | Re-aggregate |
Communication styles
| Style | Use | Risks |
|---|---|---|
| Sync request/response (HTTP/gRPC) | Queries, user-facing need immediate result | Coupling, cascading failures |
| Async events/messages | State change propagation, decoupling | Eventual consistency complexity |
| Batch / streams | Analytics, sync | Freshness |
Default rule: don’t make a sync call on a write path if an event can finish the workflow — but don’t create event spaghetti either.
Production case study (high volume)
Context: Checkout wrote order → sync called payments → sync called inventory → sync called email; any hop timeout failed the purchase after card auth. Why seniors care: Sync write chains amplify p99 and create distributed failure coupling; money paths need idempotent async completion with clear UX. Failure / symptom: Elevated checkout failures when email/SMS provider slows; nested timeout budgets exceeded; double charges on client retry. Resolution: Auth payment with idempotency; persist order; outbox events for inventory/email; sagas with compensations; sync only for user-visible authorization step. Seen at / similar to: Amazon checkout/event stories; Stripe payment intents; Airbnb/Uber order state machines.
Data & consistency
| Approach | Notes |
|---|---|
| Single-service ACID | Best when possible |
| Saga | Long-running business tx across services |
| CQRS / read models | Scale reads; accept lag |
| Shared kernel / duplicate data | Cache with clear invalidation ownership |
Be ready to explain read-your-writes, causal consistency, and user-visible lag.
Production case study (high volume)
Context: Social feed CQRS: write path to graph service; read models via Kafka → Elasticsearch (~billions of events/day ingest). Why seniors care: Users notice lag; dual-write without outbox loses events; seniors set freshness SLOs and repair paths. Failure / symptom: Missing posts in feed; consumer lag hours; inconsistent counts across regions. Resolution: Outbox/CDC; lag alerts as product SLO; rebuild-from-log; read-your-writes via sticky read or version tokens. Seen at / similar to: Twitter/X fan-out; LinkedIn feed; Netflix recommendation pipelines; Facebook TAO-style read paths (conceptual).
Resilience patterns (name + when)
| Pattern | Purpose |
|---|---|
| Timeout | Bound wait |
| Retry + jitter | Transient faults only; idempotent ops |
| Circuit breaker | Fail fast when dependency sick |
| Bulkhead | Isolate pools/resources |
| Rate limit / load shed | Protect core capacity |
| Hedging | Duplicate requests carefully (costly) |
Align with thread pools and LB/gateway.
Production case study (high volume)
Context: Black Friday: fraud vendor latency 2s→30s; without bulkheads, all checkout threads blocked; circuit breaker absent. Why seniors care: One dependency’s brownout is a site-wide outage without isolation; seniors design degradation (skip non-critical fraud tier) explicitly. Failure / symptom: Thread pool exhaustion; 503s on cart; cascading timeouts upstream. Resolution: Bulkhead + circuit breaker + deadline propagation; degrade to rules-only fraud; shed load at gateway; game-day the vendor failure. Seen at / similar to: Netflix Hystrix → Resilience4j; Amazon “avoid cascading failures”; Walmart/Target peak retail postmortems (public themes).
Java under the hood
| Concern | Common Java stack |
|---|---|
| Sync RPC | java.net.http.HttpClient, Spring WebClient, gRPC stubs |
| Timeouts / retry / CB | Resilience4j (TimeLimiter, Retry, CircuitBreaker) or Failsafe |
| Bulkhead | Resilience4j semaphore vs thread-pool bulkhead; separate Executors per dependency |
| Idempotency | DB unique constraint / Redis SETNX keyed by idempotency header |
| Observability | Micrometer metrics; OpenTelemetry Java agent — trace propagation (traceparent) |
| Contracts | OpenAPI clients, protobuf stubs — see spec-driven |
Judgment: framework choice is secondary to timeout budgets, isolation, and idempotent writes.
Deployment & versioning rules
- Blue/green or canary with health + metrics gates
- Expand/contract for schema and API changes
- Backward-compatible event evolution (add fields; don’t rename casually)
What interviewers probe
- How do you decide service boundaries? (capability + data + change frequency)
- Shared database across services — why is it bad? Exceptions?
- Saga vs 2PC
- Design order → payment → inventory with failure at each step
- Fan-out failure and cascading timeout storms — mitigations
- “Would you start with microservices?” — usually no for greenfield small teams; explain migration path
- Contract testing / API versioning strategy
- Cascading failure from one vendor under peak — what you cut first.
Senior-level expectation: Judgment and tradeoffs; operational realism; consistency story that matches UX.
Pitfalls
- Distributed monolith (sync mesh of calls + shared DB)
- Chatty interfaces (N+1 RPCs)
- Ignoring idempotency
- Gateway or orchestration service becomes new monolith
- Per-service snowflake infra without platform
- Sync email/SMS on the payment critical path
- Dual-write to DB and Kafka without outbox