AWS Services

On this page 19

Part VI — Cloud · Interview reference

Senior AWS fluency is selection criteria and failure modes, not memorizing every product name. Be able to pick a default building block, justify alternatives, and discuss IAM, networking, and cost.


How to answer AWS questions

Structure every choice as:

  1. Workload shape (latency, throughput, state, consistency)
  2. Ops model (managed vs you run)
  3. Failure domain (AZ, region, blast radius)
  4. Cost drivers (data transfer, requests, idle capacity)
  5. Security boundary (IAM, VPC, encryption)

Compute

ServiceUse whenWatch outs
EC2Full control; specialized kernels; lift-and-shiftYou patch; capacity planning
ECS / FargateContainers without kube complexityTask networking/IAM nuance
EKSNeed Kubernetes ecosystemControl plane + ops cost
LambdaEvent-driven, spiky, short jobsCold start; 15m limit; VPC cold; at-least-once
Elastic BeanstalkSimple web appsLess common in greenfield senior designs

Decision: long-running always-on APIs → ECS/EKS/EC2; bursty async → Lambda (or mix).

Production case study (high volume)

Context: Retail flash sale: API on ECS/Fargate multi-AZ; async order enrichment on Lambda from SQS; peak hundreds of K RPS at edge. Why seniors care: Lambda concurrency + cold starts interact with downstream RDS; ECS capacity planning vs cost; wrong compute choice fails either latency or wallet. Failure / symptom: RDS connection storms from Lambda scale-out; cold-start p99 on auth Lambda; Fargate task thrash on bad health checks. Resolution: RDS Proxy / pooled data API; provisioned concurrency for critical Lambdas; scale ECS on queue depth + RPS; load-test with Multi-AZ failover. Seen at / similar to: Amazon.com/AWS reference architectures; Coca-Cola/enterprise serverless case studies; Netflix on AWS (EC2/containers historically).


Storage & data

ServiceUse whenWatch outs
S3Object storage; data lake; backupsConsistency model fine for most; request patterns/cost; public access
EBSBlock for EC2AZ-bound; snapshot strategy
EFSShared POSIX-ish filePerformance modes; cost
RDSManaged SQLSingle-primary write scaling; failover RPO/RTO
AuroraMySQL/Postgres-compatible, faster storage layerCost; still not infinite write scale
DynamoDBKey-value / document; huge scale; serverless opsAccess patterns must be designed up front; hot partitions
ElastiCacheRedis/Memcached cache/sessionEviction; stampede; durability expectations
OpenSearchSearch/logsOps & cost; not primary system of record

DynamoDB interview must: PK/SK design, GSIs, single-table vs multi, capacity modes, idempotency.

Production case study (high volume)

Context: Social app stored all writes with PK=USER#<id> and a celebrity account became a hot partition; ElastiCache stampede on session expiry at top of hour. Why seniors care: DynamoDB throttling is often design, not “need more tables”; cache stampede takes down origin; S3 request patterns (small chatty GETs) blow cost. Failure / symptom: ProvisionedThroughputExceeded on one partition key; Redis CPU cliff; S3 bill surprise. Resolution: Write sharding / salt hot keys; DAX or singleflight; TTL jitter; S3 prefix design + CloudFront; revisit GSIs for access patterns. Seen at / similar to: Amazon DynamoDB best-practice posts; Discord/Slack presence on cloud KV; Netflix EVCache; Lyft/Uber Redis usage.


Networking

ServiceRole
VPCIsolation; subnets public/private
ALB / NLB / CLBL7 / L4 / legacy load balancing
API GatewayHTTP/REST/WebSocket edge; auth; throttling
CloudFrontCDN; edge caching; TLS
Route 53DNS; health checks; routing policies
NAT GatewayEgress from private subnets — cost hotspot
PrivateLink / endpointsPrivate access to AWS APIs / services

Link to LB vs API gateway.

Production case study (high volume)

Context: Multi-AZ microservices egressing to public SaaS APIs via NAT Gateway; data processing fees dominated the bill after traffic 5×. Why seniors care: NAT cost and cross-AZ transfer are senior-level architecture; PrivateLink/endpoints cut both risk and spend. Failure / symptom: Monthly bill spike unexplained by EC2; latency via congested NAT; single NAT AZ failure impacts egress. Resolution: VPC endpoints; per-AZ NATs consciously; CloudFront for downloads; track DataProcessing-Bytes; prefer PrivateLink for partners. Seen at / similar to: Widespread AWS cost postmortems; enterprise PrivateLink adoptions; Cloudflare as alternative edge.


Messaging & streaming

ServiceUse when
SQSDecouple workers; at-least-once queue; DLQ
SNSFan-out pub/sub
EventBridgeEvent bus; SaaS/AWS integrations; rules
Kinesis Data StreamsStreaming ingest (shards)
MSK (Kafka)Kafka compatibility / ecosystem

Compare to Kafka chapter: SQS simpler ops; Kafka stronger for multi-consumer log/replay.

Production case study (high volume)

Context: Payments webhooks → SNS → multiple SQS queues (ledger, fraud, analytics); visibility timeout shorter than processing → duplicate receives; Lambda retries doubled side effects. Why seniors care: At-least-once is default; DLQ without alarms is silent loss; Kinesis shard count is parallelism ceiling like Kafka partitions. Failure / symptom: Duplicate ledger entries; DLQ depth ignored; Kinesis IteratorAge hours. Resolution: Idempotent consumers; visibility ≥ p99 processing + buffer; DLQ alarms; scale shards; prefer MSK when replay/multi-subscribe dominate. Seen at / similar to: Amazon.com async architectures; Stripe webhook + queue patterns; Netflix (Kafka) vs AWS-native SQS shops.


Security (non-negotiable in senior answers)

TopicExpectation
IAMLeast privilege; roles not long-lived keys; condition keys
KMSEncryption at rest; key policies
Secrets Manager / SSMNo secrets in images/env committed to git
WAF / ShieldEdge protection
CloudTrail / ConfigAudit

Production case study (high volume)

Context: Compromised long-lived IAM access key in a CI AMI used by hundreds of build agents; exfil via S3. Why seniors care: Blast radius of identity dwarfs instance count; seniors design least privilege + short-lived roles (IRSA/instance profiles) unprompted. Failure / symptom: CloudTrail GetObject anomalies; unexpected regions; GuardDuty findings. Resolution: Kill keys; SCPs; IRSA/task roles; KMS CMKs; break-glass only; org-level boundaries. Seen at / similar to: Capital One-adjacent IAM lessons (public); AWS Well-Architected Security pillar; Uber/Lyft credential leak postmortems industry-wide.


Observability

ServiceRole
CloudWatchMetrics, logs, alarms
X-Ray / OpenTelemetryTracing
CloudWatch Logs InsightsAd hoc queries

Seniors mention SLOs + alarms on user pain, not only CPU.


Common architecture pairings

NeedTypical combo
Public HTTP APIRoute53 → CloudFront/WAF → ALB or API GW → ECS/Lambda → RDS/Dynamo
Async processingAPI → SQS → workers (ECS/Lambda) → DLQ
Fan-out notifyDomain event → SNS → SQS subscriptions
Static + APIS3 + CloudFront; API separate origin
Multi-AZ SQLRDS Multi-AZ; app in private subnets

Cost & ops landmines

  • NAT Gateway data processing fees
  • Idle multi-AZ / overprovisioned RDS
  • Cross-AZ and egress data transfer
  • Lambda in VPC without provisioned concurrency (latency)
  • CloudWatch logs retention unbounded
  • One giant account without org/SCPs

Production case study (high volume)

Context: Logging every request body to CloudWatch Logs at 100M req/day; retention never set; multi-region “DR” active-active attempted before single-region HA was solid. Why seniors care: Observability cost can exceed compute; multi-region without idempotent data plane is a correctness hazard. Failure / symptom: Log bill > EC2; incomplete failover runbooks; split-brain writes. Resolution: Sample/structured logs; retention policies; perfect Multi-AZ + backups first; multi-region only with clear RPO/RTO and conflict story. Seen at / similar to: AWS cost optimization cases; Netflix multi-region carefully; Slack/Discord regional strategies (public themes).


Java under the hood

Service touchpointJava reality
AWS SDK v2Sync or Netty-based async clients; default credentials provider chain (env, profile, IMDS/IRSA) — no long-lived keys in code
DynamoDBLow-level client or Enhanced Client (beans ↔ attributes); design PK/SK in code to match access patterns
S3Multipart upload APIs; stream to avoid huge byte[] on heap
RDS / AuroraJDBC + pool (Hikari); same SQL/index story as query optimisation
LambdaCold start = classload + JIT; SnapStart / Graal native-image tradeoffs; handler thread model + at-least-once retries → idempotency
SQS / SNSSDK receive/delete; visibility timeout ↔ processing time; DLQ

Interview: justify the AWS building block and how the JVM client is configured (timeouts, retries, pool sizes, credentials).


What interviewers probe

  1. Design a URL shortener / timeline / e-commerce using AWS — justify each pick.
  2. SQS vs Kafka/MSK vs Kinesis
  3. RDS vs DynamoDB decision criteria (access patterns, transactions, scale).
  4. Multi-AZ vs multi-region — RPO/RTO, active-active difficulty.
  5. IAM role for EC2/ECS task accessing S3 — no access keys.
  6. How Lambda retries interact with idempotency / DLQ.
  7. Blast radius of VPC, account, IAM.
  8. Hot partition / NAT bill / retry+DLQ production narratives.

Senior-level expectation: Tradeoff-aware defaults; security and cost included unprompted; clear consistency/failure story.


Pitfalls

  • Reciting services without workload fit
  • Putting databases in public subnets
  • Overusing Lambda for long-lived high-throughput APIs without reason
  • Ignoring idempotency with at-least-once queues
  • Designing multi-region before single-region HA is solid
  • Long-lived access keys in images/CI
  • Unbounded CloudWatch retention at high QPS

Cross-references

Interview reference — explanation quality and judgment, not syntax memorization.