TCP/IP Stack
On this page 19
Part III — Networking & protocols · Interview reference
Seniors debug latency, timeouts, and partial failures using a layered model. Interviews probe whether you know what each layer guarantees — and what it does not.
Layered model (practical view)
| Layer | Role | Examples |
|---|---|---|
| Application | App protocols | HTTP, gRPC, Kafka protocol, DNS |
| Transport | Process-to-process channels | TCP, UDP, TLS (often discussed here/above) |
| Network | Host-to-host routing | IP, ICMP |
| Link | Local network segment | Ethernet, Wi-Fi |
| Physical | Bits on medium | — |
OSI 7-layer naming is fine; be accurate about guarantees.
IP essentials
- Unreliable, best-effort packet delivery; can drop/reorder/duplicate
- IPv4 vs IPv6 addressing; NAT common on IPv4 paths
- ICMP for diagnostics (
ping,traceroute— often filtered)
TCP — what you must explain
Guarantees
- Reliable byte stream: retransmission, ordering, checksums
- Connection-oriented: handshake before data
- Flow control: receiver window
- Congestion control: protect the network (AIMD variants, CUBIC, BBR, …)
Three-way handshake / teardown
SYN → SYN-ACK → ACK (establish)
FIN/ACK exchange (or RST) (terminate)
TIME_WAIT: intentional; don’t “fix” by wildly reducing without understanding port exhaustion.
States worth knowing
LISTEN, SYN_SENT, ESTABLISHED, CLOSE_WAIT, TIME_WAIT, FIN_WAIT_*
CLOSE_WAIT pileup → application not closing sockets (bug).
Windows & performance
- Nagle / delayed ACK interactions → latency surprises for small RPCs (often disable Nagle for chatty RPC:
TCP_NODELAY) - Bandwidth-delay product: window must be large enough for high BDP paths
- Head-of-line blocking within a TCP connection (HTTP/2 vs HTTP/3/QUIC angle)
Timeouts & failure modes
| Symptom | Often means |
|---|---|
| Connect timeout | SYN not accepted / filtered / wrong route |
| Read idle timeout | App-level; peer stuck or middlebox |
| Connection reset | Peer RST; crash; backlog overflow; LB |
| Half-open | One side dead; need keepalive / app heartbeat |
TCP keepalive ≠ application health. Use app-level heartbeats for RPCs.
Production case study (high volume)
Context: Microservices mesh for a fintech (~50M API calls/day): gRPC between ledger, fraud, and notifications; ALB idle timeout 60s; app read timeout 120s.
Why seniors care: Mismatched timeouts cause connection resets attributed to “network flakes”; half-open connections corrupt pools; Nagle + small RPCs inflate p99.
Failure / symptom: Intermittent Connection reset / read timeouts; ss shows CLOSE_WAIT pileup on one service (app not closing); retransmit metrics up on noisy AZ path.
Resolution: Align LB idle < app idle with heartbeat; TCP_NODELAY for chatty RPC; fix missing close paths causing CLOSE_WAIT; tcpdump only after metrics point to a hop.
Seen at / similar to: AWS ALB+gRPC timeout incidents; Google SRE timeout budget thinking; Cloudflare/Fastly origin connect issues.
UDP (contrast)
- No connection, no reliability — app must handle loss/order if needed
- Low latency / streaming / QUIC underpins HTTP/3
- Good interview contrast: when you’d choose UDP/QUIC vs TCP
Sockets & servers
- Listen backlog: SYN queue / accept queue overflows → client fails under load
- Ephemeral ports: client-side exhaustion under many short connections → prefer pools / HTTP keep-alive / HTTP/2
- Connection pooling: critical for microservices
Production case study (high volume)
Context: E-commerce edge talking to dozens of backends without HTTP keep-alive / connection pools during a flash sale (hundreds of K RPS at peak).
Why seniors care: Ephemeral port exhaustion and TIME_WAIT storms look like random connect failures; SYN backlog overflows drop new clients; seniors diagnose with ss -s not app logs alone.
Failure / symptom: Cannot assign requested address; elevated connect timeouts; accept queue overflows on origins.
Resolution: Pool + keep-alive / HTTP/2 multiplexing; raise backlog carefully; tune TIME_WAIT only with understanding; scale out clients; prefer fewer longer connections.
Seen at / similar to: Classic TIME_WAIT outages at high-churn JVM shops; Netflix Zuul/Ribbon pool tuning; HAProxy/NGINX backlog docs.
Middleboxes & cloud reality
- Load balancers, NAT, firewalls: idle connection kills
- Security groups / NACLs: connect timeouts vs refused
- MTU / fragmentation path issues (VPN especially)
Debugging toolkit (name confidently)
| Tool | Use |
|---|---|
ss / netstat | States, queues |
tcpdump / Wireshark | Packets, retransmits |
curl -v / openssl s_client | App/TLS layer |
| traceroute / mtr | Path loss/latency |
| Metrics | Retransmit rates, connect failures, pool wait |
Java under the hood
| API | What it wraps |
|---|---|
Socket / ServerSocket | Blocking BIO; one thread per connection typical |
SocketChannel + Selector | NIO non-blocking; OS epoll/kqueue/select under Selector |
ByteBuffer | Heap or direct; flip/clear discipline; framing is app responsibility (TCP is a byte stream) |
| Options | StandardSocketOptions.TCP_NODELAY, SO_TIMEOUT (read timeout), keepalive — map to OS socket opts |
java.net.http.HttpClient (11+) | Connection pool + HTTP/1.1 and HTTP/2; prefer over HttpURLConnection |
| Exceptions | ConnectException (connect failed), SocketTimeoutException (connect/read timeout), Connection reset |
Production servers often use Netty (event loop, direct buffers, pooled allocators) rather than raw NIO — same mental model, richer pipeline. TLS: see SSL/TLS (SSLEngine with channels).
Production case study (high volume)
Context: Netty-based API gateway for streaming video control plane; byte-stream framing bugs under partial reads at multi-GB/s aggregate.
Why seniors care: TCP has no message boundaries — length-prefix bugs cause rare corruption at scale; direct buffer leaks → native OOM.
Failure / symptom: Occasional parse errors; OutOfMemoryError: Direct buffer memory; Selector threads hot.
Resolution: Explicit framing codecs; leak detection in Netty; bound allocators; load-test with partial/chunked reads.
Seen at / similar to: Netty at Apple Push / LinkedIn / Discord-class gateways; gRPC Java (Netty transport).
What interviewers probe
- What TCP guarantees vs IP.
- Explain handshake and why TIME_WAIT exists.
- Difference between connect timeout and read timeout — where configured.
- Why connection pooling matters; TIME_WAIT / ephemeral port exhaustion.
- TCP_NODELAY / small message latency.
- How you’d diagnose “service intermittently times out” across DNS, TCP, TLS, app.
- HTTP/2 HOL blocking vs HTTP/3.
- CLOSE_WAIT vs TIME_WAIT ownership (app vs stack).
Senior-level expectation: Layered debugging narrative; correct blame (app vs network vs LB idle timeout).
Pitfalls
- Treating TCP as a message boundary (it’s a byte stream — need framing)
- Relying only on TCP keepalive for liveness
- Ignoring LB idle timeouts shorter than app timeouts
- Opening new TCP+TLS connection per request in hot paths
- Misreading
CLOSE_WAITas a network provider issue - “Fixing” TIME_WAIT with sysctls before fixing connection churn