Load Balancing vs API Gateway

On this page 18

Part IV — System design & architecture · Interview reference

Interviewers expect a clean separation: load balancers distribute traffic across instances; API gateways are an application entry policy layer. They often compose; they are not synonyms.


Definitions

ComponentPrimary job
Load balancer (LB)Spread connections/requests across healthy backends; improve availability & scale
Reverse proxyTerminate client connection; forward to upstream (often overlapping with L7 LB)
API gatewayEdge application facade: routing by API, authn/authz, rate limit, request shaping, sometimes aggregation
Service mesh sidecarEast-west policy/mTLS/retries (complementary; not a replacement for edge gateway in all designs)

Comparison matrix

ConcernLoad balancerAPI gateway
L4 TCP/UDP distributionCoreUnusual as primary
L7 HTTP routingCommon (path/host)Core (path, header, method, version)
Health checksCoreOften delegates / also checks
TLS terminationCommonCommon
AuthN (JWT, API keys, OIDC)Limited / add-onsCore expectation
Rate limiting / quotasBasic or WAFCore product feature
Request/response transformLimitedCommon
API composition / BFFNoSometimes (carefully)
Circuit breaking / retriesSome L7 productsCommon
WAF / botOften paired (CloudFront+WAF, etc.)Sometimes integrated

Load balancing depth

L4 vs L7

L4L7
SeesIP/port/TCPHTTP/gRPC metadata
RoutingConnection-basedURL, headers, cookies
TLSPass-through or terminateUsually terminate (or re-encrypt)
CostFast, simpleMore CPU; richer policy

Algorithms

AlgorithmNotes
Round robinSimple; uneven if workloads differ
Least connectionsBetter for long-lived uneven connections
Least response timeNeeds good metrics; can oscillate
Consistent hashingSession affinity / cache locality; mind peer change rehash
Maglev / rendezvousLarge-scale consistent mapping

Health checking

  • Shallow (TCP accept) vs deep (HTTP /health that checks deps)
  • Deep checks can cascade fail entire fleet if shared dependency dies — prefer shallow for LB + separate readiness vs liveness in orchestrators
  • Slow start / connection draining on deploy

Sticky sessions

Prefer stateless services + external session store. Stickiness reduces failure isolation and complicates scaling.

Production case study (high volume)

Context: Ride-hailing API fleet behind NLB/ALB (~millions of trips/day); deep /health checked shared Redis+Postgres; sticky sessions for “in-memory cart” leftovers. Why seniors care: Deep health checks cascade multi-AZ outages when a dependency blips; stickiness concentrates load and complicates deploys; Maglev/consistent hash peer changes reshuffle caches. Failure / symptom: All targets marked unhealthy during Redis blip; one sticky AZ hot; cache hit ratio falls after scale event. Resolution: Shallow LB checks + K8s readiness separate from liveness; externalize session; connection draining; slow start; cell-based / shard routing where needed. Seen at / similar to: Uber/Lyft edge; Google Maglev paper; AWS ALB target group health pitfalls; Netflix active-active stories.


API gateway depth

Typical responsibilities

  1. External API routing / versioning (/v1, header versioning)
  2. Authentication & coarse authorization
  3. Rate limiting, API keys, quotas
  4. Request validation / schema enforcement
  5. Observability: edge metrics, tracing headers
  6. Protocol adaptation (REST↔gRPC) — optional

Anti-responsibilities (senior judgment)

  • Heavy business logic / domain rules (becomes a monolith bottleneck)
  • Large fan-out aggregations without caching (latency + failure coupling)
  • Replacing service-to-service auth entirely if mesh/mTLS already covers east-west

Production case study (high volume)

Context: Public fintech OpenAPI via Amazon API Gateway / Kong: JWT auth, per-key rate limits, request validation; gateway team started embedding “smart” multi-service aggregation for mobile. Why seniors care: Gateway as BFF-without-bounds becomes latency and deploy bottleneck; rate limits misaligned with thread pools → fair clients starved or backends melted. Failure / symptom: Gateway p99 owns the error budget; fan-out partial failures return inconsistent payloads; 429 storms after marketing email. Resolution: Keep gateway thin; BFF as separate service; align gateway limits with bulkheads; cache aggregations carefully; chaos partial upstreams. Seen at / similar to: Netflix Zuul/Edge; Stripe API versioning+limits; Cloudflare Workers/API Shield patterns; Apigee enterprise gateways.


Composition patterns

Client → (CDN/WAF) → LB → API Gateway → Services
                 ↘︎ sometimes combined appliances / cloud managed
PatternWhen
LB onlySimple services; internal apps; gRPC with mesh
Gateway + LBPublic APIs needing policy; multi-service facade
Gateway per BFFMobile/web-specific aggregation
Ingress controller as “gateway-ish”K8s edge; still separate concerns carefully

Cloud examples (illustrative): ALB/NLB + API Gateway; Cloud Load Balancing + Apigee; nginx/Envoys playing both roles with different configs.


Failure & overload

  • LB must fail away from bad instances quickly without flapping
  • Gateway rate limits protect backends — coordinate with thread pools and bulkheads
  • Timeouts: client → gateway → service must be aligned (outer ≥ inner with budget)
  • Retries amplify load — idempotency required (microservices)

Production case study (high volume)

Context: During a payments provider brownout, clients + gateway + service each retried 3× without budgets on a non-idempotent charge path. Why seniors care: Retry storms turn a dependency incident into a self-inflicted outage; timeout misalignment guarantees amplification. Failure / symptom: 10× traffic to sick dependency; thread pools exhausted; successful charges duplicated. Resolution: Single retry budget end-to-end; idempotency keys; hedged requests only where safe; fail fast with circuit breakers; outer timeout ≥ sum of inner budgets. Seen at / similar to: Amazon Builders’ Library timeouts/retries; Google SRE handling overload; Stripe Idempotency-Key culture.


Decision criteria

If you need…Prefer
Scale identical stateless instancesLB
Cross-cutting API policy at edgeGateway
Raw TCP / non-HTTPL4 LB
Central auth for many public APIsGateway
Minimal hops / ultra-low latency internalMaybe skip extra gateway hop

Java under the hood

No JDK load balancer. In Java services you usually consume LBs/gateways and implement client-side policy:

ConcernTypical Java piece
Client LBgRPC Java round_robin / pick_first / custom LoadBalancerProvider; service discovery (Consul, K8s)
Consistent hashGuava Hashing.consistentHash / ring libraries for cache affinity
In-process gatewaySpring Cloud Gateway (Netty) or servlet filters — routing, auth, rate limit at edge JVM
ResilienceResilience4j: retry, circuit breaker, rate limiter, bulkhead (semaphore vs thread-pool)
TimeoutsHttpClient connect/request timeouts; gRPC deadlines — align with outer LB idle timeouts

Edge policy still lives in ALB/API Gateway/Envoy; the JVM owns outbound timeouts, retries, and bulkheads so one dependency cannot exhaust the pool (thread pools).


What interviewers probe

  1. “What’s the difference between LB and API gateway?” — crisp table-level answer.
  2. Design public API edge for 20 microservices — where auth, limit, TLS live.
  3. L4 vs L7 choice for gRPC / WebSockets.
  4. Health check design that doesn’t take down all AZs on DB blip.
  5. Retry + timeout budgets across LB/gateway/service.
  6. Sticky sessions — why avoid.
  7. Retry storm narrative with metrics you’d watch (RPS amplification, pool wait, 429/503).

Senior-level expectation: Compose components; place policy at the right layer; discuss blast radius.


Pitfalls

  • Calling every reverse proxy an “API gateway” in design docs
  • Putting business workflows in the gateway
  • Inconsistent timeouts causing retry storms
  • LB pointing at pods without readiness discipline
  • Double rate-limiting that confuses clients
  • Deep health checks on shared dependencies

Cross-references

Interview reference — explanation quality and judgment, not syntax memorization.