TCP/IP Stack

On this page 19

Part III — Networking & protocols · Interview reference

Seniors debug latency, timeouts, and partial failures using a layered model. Interviews probe whether you know what each layer guarantees — and what it does not.


Layered model (practical view)

LayerRoleExamples
ApplicationApp protocolsHTTP, gRPC, Kafka protocol, DNS
TransportProcess-to-process channelsTCP, UDP, TLS (often discussed here/above)
NetworkHost-to-host routingIP, ICMP
LinkLocal network segmentEthernet, Wi-Fi
PhysicalBits on medium—

OSI 7-layer naming is fine; be accurate about guarantees.


IP essentials

  • Unreliable, best-effort packet delivery; can drop/reorder/duplicate
  • IPv4 vs IPv6 addressing; NAT common on IPv4 paths
  • ICMP for diagnostics (ping, traceroute — often filtered)

TCP — what you must explain

Guarantees

  • Reliable byte stream: retransmission, ordering, checksums
  • Connection-oriented: handshake before data
  • Flow control: receiver window
  • Congestion control: protect the network (AIMD variants, CUBIC, BBR, …)

Three-way handshake / teardown

SYN → SYN-ACK → ACK          (establish)
FIN/ACK exchange (or RST)    (terminate)

TIME_WAIT: intentional; don’t “fix” by wildly reducing without understanding port exhaustion.

States worth knowing

LISTEN, SYN_SENT, ESTABLISHED, CLOSE_WAIT, TIME_WAIT, FIN_WAIT_*

CLOSE_WAIT pileup → application not closing sockets (bug).

Windows & performance

  • Nagle / delayed ACK interactions → latency surprises for small RPCs (often disable Nagle for chatty RPC: TCP_NODELAY)
  • Bandwidth-delay product: window must be large enough for high BDP paths
  • Head-of-line blocking within a TCP connection (HTTP/2 vs HTTP/3/QUIC angle)

Timeouts & failure modes

SymptomOften means
Connect timeoutSYN not accepted / filtered / wrong route
Read idle timeoutApp-level; peer stuck or middlebox
Connection resetPeer RST; crash; backlog overflow; LB
Half-openOne side dead; need keepalive / app heartbeat

TCP keepalive ≠ application health. Use app-level heartbeats for RPCs.

Production case study (high volume)

Context: Microservices mesh for a fintech (~50M API calls/day): gRPC between ledger, fraud, and notifications; ALB idle timeout 60s; app read timeout 120s. Why seniors care: Mismatched timeouts cause connection resets attributed to “network flakes”; half-open connections corrupt pools; Nagle + small RPCs inflate p99. Failure / symptom: Intermittent Connection reset / read timeouts; ss shows CLOSE_WAIT pileup on one service (app not closing); retransmit metrics up on noisy AZ path. Resolution: Align LB idle < app idle with heartbeat; TCP_NODELAY for chatty RPC; fix missing close paths causing CLOSE_WAIT; tcpdump only after metrics point to a hop. Seen at / similar to: AWS ALB+gRPC timeout incidents; Google SRE timeout budget thinking; Cloudflare/Fastly origin connect issues.


UDP (contrast)

  • No connection, no reliability — app must handle loss/order if needed
  • Low latency / streaming / QUIC underpins HTTP/3
  • Good interview contrast: when you’d choose UDP/QUIC vs TCP

Sockets & servers

  • Listen backlog: SYN queue / accept queue overflows → client fails under load
  • Ephemeral ports: client-side exhaustion under many short connections → prefer pools / HTTP keep-alive / HTTP/2
  • Connection pooling: critical for microservices

Production case study (high volume)

Context: E-commerce edge talking to dozens of backends without HTTP keep-alive / connection pools during a flash sale (hundreds of K RPS at peak). Why seniors care: Ephemeral port exhaustion and TIME_WAIT storms look like random connect failures; SYN backlog overflows drop new clients; seniors diagnose with ss -s not app logs alone. Failure / symptom: Cannot assign requested address; elevated connect timeouts; accept queue overflows on origins. Resolution: Pool + keep-alive / HTTP/2 multiplexing; raise backlog carefully; tune TIME_WAIT only with understanding; scale out clients; prefer fewer longer connections. Seen at / similar to: Classic TIME_WAIT outages at high-churn JVM shops; Netflix Zuul/Ribbon pool tuning; HAProxy/NGINX backlog docs.


Middleboxes & cloud reality

  • Load balancers, NAT, firewalls: idle connection kills
  • Security groups / NACLs: connect timeouts vs refused
  • MTU / fragmentation path issues (VPN especially)

Debugging toolkit (name confidently)

ToolUse
ss / netstatStates, queues
tcpdump / WiresharkPackets, retransmits
curl -v / openssl s_clientApp/TLS layer
traceroute / mtrPath loss/latency
MetricsRetransmit rates, connect failures, pool wait

Java under the hood

APIWhat it wraps
Socket / ServerSocketBlocking BIO; one thread per connection typical
SocketChannel + SelectorNIO non-blocking; OS epoll/kqueue/select under Selector
ByteBufferHeap or direct; flip/clear discipline; framing is app responsibility (TCP is a byte stream)
OptionsStandardSocketOptions.TCP_NODELAY, SO_TIMEOUT (read timeout), keepalive — map to OS socket opts
java.net.http.HttpClient (11+)Connection pool + HTTP/1.1 and HTTP/2; prefer over HttpURLConnection
ExceptionsConnectException (connect failed), SocketTimeoutException (connect/read timeout), Connection reset

Production servers often use Netty (event loop, direct buffers, pooled allocators) rather than raw NIO — same mental model, richer pipeline. TLS: see SSL/TLS (SSLEngine with channels).

Production case study (high volume)

Context: Netty-based API gateway for streaming video control plane; byte-stream framing bugs under partial reads at multi-GB/s aggregate. Why seniors care: TCP has no message boundaries — length-prefix bugs cause rare corruption at scale; direct buffer leaks → native OOM. Failure / symptom: Occasional parse errors; OutOfMemoryError: Direct buffer memory; Selector threads hot. Resolution: Explicit framing codecs; leak detection in Netty; bound allocators; load-test with partial/chunked reads. Seen at / similar to: Netty at Apple Push / LinkedIn / Discord-class gateways; gRPC Java (Netty transport).


What interviewers probe

  1. What TCP guarantees vs IP.
  2. Explain handshake and why TIME_WAIT exists.
  3. Difference between connect timeout and read timeout — where configured.
  4. Why connection pooling matters; TIME_WAIT / ephemeral port exhaustion.
  5. TCP_NODELAY / small message latency.
  6. How you’d diagnose “service intermittently times out” across DNS, TCP, TLS, app.
  7. HTTP/2 HOL blocking vs HTTP/3.
  8. CLOSE_WAIT vs TIME_WAIT ownership (app vs stack).

Senior-level expectation: Layered debugging narrative; correct blame (app vs network vs LB idle timeout).


Pitfalls

  • Treating TCP as a message boundary (it’s a byte stream — need framing)
  • Relying only on TCP keepalive for liveness
  • Ignoring LB idle timeouts shorter than app timeouts
  • Opening new TCP+TLS connection per request in hot paths
  • Misreading CLOSE_WAIT as a network provider issue
  • “Fixing” TIME_WAIT with sysctls before fixing connection churn

Cross-references

Interview reference — explanation quality and judgment, not syntax memorization.