Load Balancing vs API Gateway
On this page 18
Part IV — System design & architecture · Interview reference
Interviewers expect a clean separation: load balancers distribute traffic across instances; API gateways are an application entry policy layer. They often compose; they are not synonyms.
Definitions
| Component | Primary job |
|---|---|
| Load balancer (LB) | Spread connections/requests across healthy backends; improve availability & scale |
| Reverse proxy | Terminate client connection; forward to upstream (often overlapping with L7 LB) |
| API gateway | Edge application facade: routing by API, authn/authz, rate limit, request shaping, sometimes aggregation |
| Service mesh sidecar | East-west policy/mTLS/retries (complementary; not a replacement for edge gateway in all designs) |
Comparison matrix
| Concern | Load balancer | API gateway |
|---|---|---|
| L4 TCP/UDP distribution | Core | Unusual as primary |
| L7 HTTP routing | Common (path/host) | Core (path, header, method, version) |
| Health checks | Core | Often delegates / also checks |
| TLS termination | Common | Common |
| AuthN (JWT, API keys, OIDC) | Limited / add-ons | Core expectation |
| Rate limiting / quotas | Basic or WAF | Core product feature |
| Request/response transform | Limited | Common |
| API composition / BFF | No | Sometimes (carefully) |
| Circuit breaking / retries | Some L7 products | Common |
| WAF / bot | Often paired (CloudFront+WAF, etc.) | Sometimes integrated |
Load balancing depth
L4 vs L7
| L4 | L7 | |
|---|---|---|
| Sees | IP/port/TCP | HTTP/gRPC metadata |
| Routing | Connection-based | URL, headers, cookies |
| TLS | Pass-through or terminate | Usually terminate (or re-encrypt) |
| Cost | Fast, simple | More CPU; richer policy |
Algorithms
| Algorithm | Notes |
|---|---|
| Round robin | Simple; uneven if workloads differ |
| Least connections | Better for long-lived uneven connections |
| Least response time | Needs good metrics; can oscillate |
| Consistent hashing | Session affinity / cache locality; mind peer change rehash |
| Maglev / rendezvous | Large-scale consistent mapping |
Health checking
- Shallow (TCP accept) vs deep (HTTP
/healththat checks deps) - Deep checks can cascade fail entire fleet if shared dependency dies — prefer shallow for LB + separate readiness vs liveness in orchestrators
- Slow start / connection draining on deploy
Sticky sessions
Prefer stateless services + external session store. Stickiness reduces failure isolation and complicates scaling.
Production case study (high volume)
Context: Ride-hailing API fleet behind NLB/ALB (~millions of trips/day); deep /health checked shared Redis+Postgres; sticky sessions for “in-memory cart” leftovers.
Why seniors care: Deep health checks cascade multi-AZ outages when a dependency blips; stickiness concentrates load and complicates deploys; Maglev/consistent hash peer changes reshuffle caches.
Failure / symptom: All targets marked unhealthy during Redis blip; one sticky AZ hot; cache hit ratio falls after scale event.
Resolution: Shallow LB checks + K8s readiness separate from liveness; externalize session; connection draining; slow start; cell-based / shard routing where needed.
Seen at / similar to: Uber/Lyft edge; Google Maglev paper; AWS ALB target group health pitfalls; Netflix active-active stories.
API gateway depth
Typical responsibilities
- External API routing / versioning (
/v1, header versioning) - Authentication & coarse authorization
- Rate limiting, API keys, quotas
- Request validation / schema enforcement
- Observability: edge metrics, tracing headers
- Protocol adaptation (REST↔gRPC) — optional
Anti-responsibilities (senior judgment)
- Heavy business logic / domain rules (becomes a monolith bottleneck)
- Large fan-out aggregations without caching (latency + failure coupling)
- Replacing service-to-service auth entirely if mesh/mTLS already covers east-west
Production case study (high volume)
Context: Public fintech OpenAPI via Amazon API Gateway / Kong: JWT auth, per-key rate limits, request validation; gateway team started embedding “smart” multi-service aggregation for mobile. Why seniors care: Gateway as BFF-without-bounds becomes latency and deploy bottleneck; rate limits misaligned with thread pools → fair clients starved or backends melted. Failure / symptom: Gateway p99 owns the error budget; fan-out partial failures return inconsistent payloads; 429 storms after marketing email. Resolution: Keep gateway thin; BFF as separate service; align gateway limits with bulkheads; cache aggregations carefully; chaos partial upstreams. Seen at / similar to: Netflix Zuul/Edge; Stripe API versioning+limits; Cloudflare Workers/API Shield patterns; Apigee enterprise gateways.
Composition patterns
Client → (CDN/WAF) → LB → API Gateway → Services
↘︎ sometimes combined appliances / cloud managed
| Pattern | When |
|---|---|
| LB only | Simple services; internal apps; gRPC with mesh |
| Gateway + LB | Public APIs needing policy; multi-service facade |
| Gateway per BFF | Mobile/web-specific aggregation |
| Ingress controller as “gateway-ish” | K8s edge; still separate concerns carefully |
Cloud examples (illustrative): ALB/NLB + API Gateway; Cloud Load Balancing + Apigee; nginx/Envoys playing both roles with different configs.
Failure & overload
- LB must fail away from bad instances quickly without flapping
- Gateway rate limits protect backends — coordinate with thread pools and bulkheads
- Timeouts: client → gateway → service must be aligned (outer ≥ inner with budget)
- Retries amplify load — idempotency required (microservices)
Production case study (high volume)
Context: During a payments provider brownout, clients + gateway + service each retried 3× without budgets on a non-idempotent charge path. Why seniors care: Retry storms turn a dependency incident into a self-inflicted outage; timeout misalignment guarantees amplification. Failure / symptom: 10× traffic to sick dependency; thread pools exhausted; successful charges duplicated. Resolution: Single retry budget end-to-end; idempotency keys; hedged requests only where safe; fail fast with circuit breakers; outer timeout ≥ sum of inner budgets. Seen at / similar to: Amazon Builders’ Library timeouts/retries; Google SRE handling overload; Stripe Idempotency-Key culture.
Decision criteria
| If you need… | Prefer |
|---|---|
| Scale identical stateless instances | LB |
| Cross-cutting API policy at edge | Gateway |
| Raw TCP / non-HTTP | L4 LB |
| Central auth for many public APIs | Gateway |
| Minimal hops / ultra-low latency internal | Maybe skip extra gateway hop |
Java under the hood
No JDK load balancer. In Java services you usually consume LBs/gateways and implement client-side policy:
| Concern | Typical Java piece |
|---|---|
| Client LB | gRPC Java round_robin / pick_first / custom LoadBalancerProvider; service discovery (Consul, K8s) |
| Consistent hash | Guava Hashing.consistentHash / ring libraries for cache affinity |
| In-process gateway | Spring Cloud Gateway (Netty) or servlet filters — routing, auth, rate limit at edge JVM |
| Resilience | Resilience4j: retry, circuit breaker, rate limiter, bulkhead (semaphore vs thread-pool) |
| Timeouts | HttpClient connect/request timeouts; gRPC deadlines — align with outer LB idle timeouts |
Edge policy still lives in ALB/API Gateway/Envoy; the JVM owns outbound timeouts, retries, and bulkheads so one dependency cannot exhaust the pool (thread pools).
What interviewers probe
- “What’s the difference between LB and API gateway?” — crisp table-level answer.
- Design public API edge for 20 microservices — where auth, limit, TLS live.
- L4 vs L7 choice for gRPC / WebSockets.
- Health check design that doesn’t take down all AZs on DB blip.
- Retry + timeout budgets across LB/gateway/service.
- Sticky sessions — why avoid.
- Retry storm narrative with metrics you’d watch (RPS amplification, pool wait, 429/503).
Senior-level expectation: Compose components; place policy at the right layer; discuss blast radius.
Pitfalls
- Calling every reverse proxy an “API gateway” in design docs
- Putting business workflows in the gateway
- Inconsistent timeouts causing retry storms
- LB pointing at pods without readiness discipline
- Double rate-limiting that confuses clients
- Deep health checks on shared dependencies