AWS Services
On this page 19
Part VI — Cloud · Interview reference
Senior AWS fluency is selection criteria and failure modes, not memorizing every product name. Be able to pick a default building block, justify alternatives, and discuss IAM, networking, and cost.
How to answer AWS questions
Structure every choice as:
- Workload shape (latency, throughput, state, consistency)
- Ops model (managed vs you run)
- Failure domain (AZ, region, blast radius)
- Cost drivers (data transfer, requests, idle capacity)
- Security boundary (IAM, VPC, encryption)
Compute
| Service | Use when | Watch outs |
|---|---|---|
| EC2 | Full control; specialized kernels; lift-and-shift | You patch; capacity planning |
| ECS / Fargate | Containers without kube complexity | Task networking/IAM nuance |
| EKS | Need Kubernetes ecosystem | Control plane + ops cost |
| Lambda | Event-driven, spiky, short jobs | Cold start; 15m limit; VPC cold; at-least-once |
| Elastic Beanstalk | Simple web apps | Less common in greenfield senior designs |
Decision: long-running always-on APIs → ECS/EKS/EC2; bursty async → Lambda (or mix).
Production case study (high volume)
Context: Retail flash sale: API on ECS/Fargate multi-AZ; async order enrichment on Lambda from SQS; peak hundreds of K RPS at edge. Why seniors care: Lambda concurrency + cold starts interact with downstream RDS; ECS capacity planning vs cost; wrong compute choice fails either latency or wallet. Failure / symptom: RDS connection storms from Lambda scale-out; cold-start p99 on auth Lambda; Fargate task thrash on bad health checks. Resolution: RDS Proxy / pooled data API; provisioned concurrency for critical Lambdas; scale ECS on queue depth + RPS; load-test with Multi-AZ failover. Seen at / similar to: Amazon.com/AWS reference architectures; Coca-Cola/enterprise serverless case studies; Netflix on AWS (EC2/containers historically).
Storage & data
| Service | Use when | Watch outs |
|---|---|---|
| S3 | Object storage; data lake; backups | Consistency model fine for most; request patterns/cost; public access |
| EBS | Block for EC2 | AZ-bound; snapshot strategy |
| EFS | Shared POSIX-ish file | Performance modes; cost |
| RDS | Managed SQL | Single-primary write scaling; failover RPO/RTO |
| Aurora | MySQL/Postgres-compatible, faster storage layer | Cost; still not infinite write scale |
| DynamoDB | Key-value / document; huge scale; serverless ops | Access patterns must be designed up front; hot partitions |
| ElastiCache | Redis/Memcached cache/session | Eviction; stampede; durability expectations |
| OpenSearch | Search/logs | Ops & cost; not primary system of record |
DynamoDB interview must: PK/SK design, GSIs, single-table vs multi, capacity modes, idempotency.
Production case study (high volume)
Context: Social app stored all writes with PK=USER#<id> and a celebrity account became a hot partition; ElastiCache stampede on session expiry at top of hour.
Why seniors care: DynamoDB throttling is often design, not “need more tables”; cache stampede takes down origin; S3 request patterns (small chatty GETs) blow cost.
Failure / symptom: ProvisionedThroughputExceeded on one partition key; Redis CPU cliff; S3 bill surprise.
Resolution: Write sharding / salt hot keys; DAX or singleflight; TTL jitter; S3 prefix design + CloudFront; revisit GSIs for access patterns.
Seen at / similar to: Amazon DynamoDB best-practice posts; Discord/Slack presence on cloud KV; Netflix EVCache; Lyft/Uber Redis usage.
Networking
| Service | Role |
|---|---|
| VPC | Isolation; subnets public/private |
| ALB / NLB / CLB | L7 / L4 / legacy load balancing |
| API Gateway | HTTP/REST/WebSocket edge; auth; throttling |
| CloudFront | CDN; edge caching; TLS |
| Route 53 | DNS; health checks; routing policies |
| NAT Gateway | Egress from private subnets — cost hotspot |
| PrivateLink / endpoints | Private access to AWS APIs / services |
Link to LB vs API gateway.
Production case study (high volume)
Context: Multi-AZ microservices egressing to public SaaS APIs via NAT Gateway; data processing fees dominated the bill after traffic 5×.
Why seniors care: NAT cost and cross-AZ transfer are senior-level architecture; PrivateLink/endpoints cut both risk and spend.
Failure / symptom: Monthly bill spike unexplained by EC2; latency via congested NAT; single NAT AZ failure impacts egress.
Resolution: VPC endpoints; per-AZ NATs consciously; CloudFront for downloads; track DataProcessing-Bytes; prefer PrivateLink for partners.
Seen at / similar to: Widespread AWS cost postmortems; enterprise PrivateLink adoptions; Cloudflare as alternative edge.
Messaging & streaming
| Service | Use when |
|---|---|
| SQS | Decouple workers; at-least-once queue; DLQ |
| SNS | Fan-out pub/sub |
| EventBridge | Event bus; SaaS/AWS integrations; rules |
| Kinesis Data Streams | Streaming ingest (shards) |
| MSK (Kafka) | Kafka compatibility / ecosystem |
Compare to Kafka chapter: SQS simpler ops; Kafka stronger for multi-consumer log/replay.
Production case study (high volume)
Context: Payments webhooks → SNS → multiple SQS queues (ledger, fraud, analytics); visibility timeout shorter than processing → duplicate receives; Lambda retries doubled side effects.
Why seniors care: At-least-once is default; DLQ without alarms is silent loss; Kinesis shard count is parallelism ceiling like Kafka partitions.
Failure / symptom: Duplicate ledger entries; DLQ depth ignored; Kinesis IteratorAge hours.
Resolution: Idempotent consumers; visibility ≥ p99 processing + buffer; DLQ alarms; scale shards; prefer MSK when replay/multi-subscribe dominate.
Seen at / similar to: Amazon.com async architectures; Stripe webhook + queue patterns; Netflix (Kafka) vs AWS-native SQS shops.
Security (non-negotiable in senior answers)
| Topic | Expectation |
|---|---|
| IAM | Least privilege; roles not long-lived keys; condition keys |
| KMS | Encryption at rest; key policies |
| Secrets Manager / SSM | No secrets in images/env committed to git |
| WAF / Shield | Edge protection |
| CloudTrail / Config | Audit |
Production case study (high volume)
Context: Compromised long-lived IAM access key in a CI AMI used by hundreds of build agents; exfil via S3.
Why seniors care: Blast radius of identity dwarfs instance count; seniors design least privilege + short-lived roles (IRSA/instance profiles) unprompted.
Failure / symptom: CloudTrail GetObject anomalies; unexpected regions; GuardDuty findings.
Resolution: Kill keys; SCPs; IRSA/task roles; KMS CMKs; break-glass only; org-level boundaries.
Seen at / similar to: Capital One-adjacent IAM lessons (public); AWS Well-Architected Security pillar; Uber/Lyft credential leak postmortems industry-wide.
Observability
| Service | Role |
|---|---|
| CloudWatch | Metrics, logs, alarms |
| X-Ray / OpenTelemetry | Tracing |
| CloudWatch Logs Insights | Ad hoc queries |
Seniors mention SLOs + alarms on user pain, not only CPU.
Common architecture pairings
| Need | Typical combo |
|---|---|
| Public HTTP API | Route53 → CloudFront/WAF → ALB or API GW → ECS/Lambda → RDS/Dynamo |
| Async processing | API → SQS → workers (ECS/Lambda) → DLQ |
| Fan-out notify | Domain event → SNS → SQS subscriptions |
| Static + API | S3 + CloudFront; API separate origin |
| Multi-AZ SQL | RDS Multi-AZ; app in private subnets |
Cost & ops landmines
- NAT Gateway data processing fees
- Idle multi-AZ / overprovisioned RDS
- Cross-AZ and egress data transfer
- Lambda in VPC without provisioned concurrency (latency)
- CloudWatch logs retention unbounded
- One giant account without org/SCPs
Production case study (high volume)
Context: Logging every request body to CloudWatch Logs at 100M req/day; retention never set; multi-region “DR” active-active attempted before single-region HA was solid. Why seniors care: Observability cost can exceed compute; multi-region without idempotent data plane is a correctness hazard. Failure / symptom: Log bill > EC2; incomplete failover runbooks; split-brain writes. Resolution: Sample/structured logs; retention policies; perfect Multi-AZ + backups first; multi-region only with clear RPO/RTO and conflict story. Seen at / similar to: AWS cost optimization cases; Netflix multi-region carefully; Slack/Discord regional strategies (public themes).
Java under the hood
| Service touchpoint | Java reality |
|---|---|
| AWS SDK v2 | Sync or Netty-based async clients; default credentials provider chain (env, profile, IMDS/IRSA) — no long-lived keys in code |
| DynamoDB | Low-level client or Enhanced Client (beans ↔ attributes); design PK/SK in code to match access patterns |
| S3 | Multipart upload APIs; stream to avoid huge byte[] on heap |
| RDS / Aurora | JDBC + pool (Hikari); same SQL/index story as query optimisation |
| Lambda | Cold start = classload + JIT; SnapStart / Graal native-image tradeoffs; handler thread model + at-least-once retries → idempotency |
| SQS / SNS | SDK receive/delete; visibility timeout ↔ processing time; DLQ |
Interview: justify the AWS building block and how the JVM client is configured (timeouts, retries, pool sizes, credentials).
What interviewers probe
- Design a URL shortener / timeline / e-commerce using AWS — justify each pick.
- SQS vs Kafka/MSK vs Kinesis
- RDS vs DynamoDB decision criteria (access patterns, transactions, scale).
- Multi-AZ vs multi-region — RPO/RTO, active-active difficulty.
- IAM role for EC2/ECS task accessing S3 — no access keys.
- How Lambda retries interact with idempotency / DLQ.
- Blast radius of VPC, account, IAM.
- Hot partition / NAT bill / retry+DLQ production narratives.
Senior-level expectation: Tradeoff-aware defaults; security and cost included unprompted; clear consistency/failure story.
Pitfalls
- Reciting services without workload fit
- Putting databases in public subnets
- Overusing Lambda for long-lived high-throughput APIs without reason
- Ignoring idempotency with at-least-once queues
- Designing multi-region before single-region HA is solid
- Long-lived access keys in images/CI
- Unbounded CloudWatch retention at high QPS