Spec-Driven Development

On this page 16

Part VII — Engineering practices · Interview reference

Spec-driven development treats an explicit, versioned specification as the source of truth for behavior — before (or tightly alongside) implementation. For seniors, this is about reducing ambiguity, aligning contracts, and making change reviewable, including in AI-assisted workflows.


Definition

Spec-driven development (SDD): write or update a precise description of intended behavior (API contracts, acceptance criteria, data schemas, event contracts, UI states) and use that artifact to drive implementation, tests, and review.

Related / overlapping practices:

PracticeRelationship
API-first / contract-firstOpenAPI/AsyncAPI/Protobuf as the contract
TDDTests as executable spec; SDD often broader (human + machine readable)
BDDExamples in domain language; can feed specs
Design docs / RFCsNarrative rationale; SDD focuses on normative behavior
ADRDecision record; complements specs

Why seniors care

ProblemSpec helps by
Ambiguous ticketsForces inputs/outputs/errors/edge cases
Microservice misintegrationShared contracts + versioning
AI-generated code driftMachine-checkable constraints
Review theaterDiff the spec + tests, not only code
Knowledge silosSpec becomes onboarding artifact

What a good spec contains

For an API/feature, normative sections:

  1. Purpose / non-goals
  2. Actors & authZ
  3. Inputs / outputs — schemas, headers, idempotency
  4. Invariants & error model — status codes, error codes
  5. Consistency / side effects — sync vs async; events emitted
  6. Acceptance scenarios — given/when/then or examples
  7. SLOs / limits — rate limits, payload sizes
  8. Observability — metrics, traces, logs fields
  9. Migration / compatibility — expand-contract plan

Vague adjectives (“fast”, “secure”, “handle scale”) are not a spec until quantified.

Production case study (high volume)

Context: Payments platform launching a new “Capture” API used by hundreds of merchants at millions of captures/day; ticket said “make it reliable and fast.” Why seniors care: Without normative error codes, idempotency, and SLOs, every team invents incompatible clients; incidents become contract disputes. Failure / symptom: Merchants disagree on retry behavior; duplicate captures; mobile/web diverge; oncall cannot tell bug vs undefined behavior. Resolution: Spec PR first — schemas, Idempotency-Key, error model, rate limits, observability fields; review before code; publish OpenAPI as source of truth. Seen at / similar to: Stripe API versioning discipline; Twilio/GitHub public API specs; Amazon API Gateway + OpenAPI shops.


Typical artifacts

ArtifactUse
OpenAPI / GraphQL schemaHTTP API contract
Protobuf / gRPC protoRPC contract
AsyncAPI / Avro / JSON SchemaEvents
JSON Schema / DB migration planData
State diagramsLifecycle-heavy domains
Executable acceptance testsLiving verification
Pact / contract testsConsumer-driven

Production case study (high volume)

Context: Event-driven order system (Kafka + Avro) with 20 consumers; producer renamed a field in a “compatible” deploy without registry checks. Why seniors care: Schema breaks poison all consumers at once — blast radius ≫ a single HTTP endpoint; seniors gate compatibility in CI. Failure / symptom: Consumer DLQ flood; lag explosion; mobile push/email pipelines stall during peak. Resolution: Schema Registry BACKWARD/FULL modes; contract tests in CI; expand/contract; canary consumers; treat AsyncAPI/Avro like public API. Seen at / similar to: Confluent/LinkedIn schema discipline; Netflix/Uber event evolution; Shopify Kafka contracts.


Workflow (interview-ready)

Problem → Spec PR (review) → Generate stubs/tests → Implement → Verify vs spec → Version & publish contract

Rules of thumb

  • Spec merges before or in the same change set as the first consumer-visible behavior
  • Breaking changes require version bump + migration notes
  • Generated clients/servers reduce hand-written drift — still review generators’ outputs

Spec-driven + AI-assisted development

Seniors should articulate:

DoDon’t
Feed specs as constraints to codegen agentsAccept code that contradicts the spec
Require tests generated from scenariosLet the model invent undocumented endpoints
Diff against OpenAPI in CITreat chat transcript as the contract
Keep specs in repo, reviewedKeep “truth” only in a ticket comment

Interview angle: Specs improve AI leverage by narrowing the solution space; they don’t remove the need for human judgment on tradeoffs.

Production case study (high volume)

Context: Team used an LLM to “implement the refunds endpoint” from a Slack thread; generated code invented status codes and skipped idempotency; shipped behind a feature flag to 5% traffic. Why seniors care: At high volume, invented contracts cause client breakages and money bugs faster than humans can review prose; specs make AI output checkable. Failure / symptom: Client 4xx/5xx mismatch; duplicate refunds; OpenAPI drift vs running service; AI tests asserted the wrong behavior. Resolution: Feed OpenAPI + acceptance scenarios to codegen; CI fails on spec drift; human review of money paths; chat is never the contract. Seen at / similar to: Industry AI-assist adoption at Stripe/Shopify-scale engineering orgs (process themes); Google/Microsoft API linter cultures.


Java under the hood

Spec artifactJava toolchain
OpenAPIopenapi-generator / Springdoc → interfaces, models, stubs
Protobuf / gRPCprotoc + protobuf-java / gRPC Java stubs
JSON Schema / beansJackson + Bean Validation (jakarta.validation) as executable constraints
Consumer contractsPact JVM; Spring Cloud Contract
Architecture rulesArchUnit tests as enforceable “spec” for package/layer deps
EventsAvro/Protobuf schemas + Schema Registry compatibility checks in CI

Generated code is still reviewed; CI should fail on OpenAPI/protobuf breaking changes, not only on unit tests.


Tradeoffs

BenefitCost
Clarity & alignmentUpfront time
Parallel FE/BE workSpec churn early on
Better test oraclesOver-specification of unknowns
Safer evolutionProcess can become bureaucracy if every spike needs a novel

Avoid: writing a 40-page spec for a throwaway spike. Use lightweight specs that grow with certainty.


Compatibility strategies

  • Additive changes preferred (new optional fields)
  • Explicit versioning (/v2, package versions, schema registry compatibility modes)
  • Consumer-driven contracts when many consumers evolve at different speeds
  • Feature flags for behavioral rollout after contract publish

Production case study (high volume)

Context: Mobile apps with slow release cadence consumed a gRPC banking API; backend removed a field still read by old apps during a peak payroll week. Why seniors care: Compatibility is an availability property at scale; expand/contract and versioning prevent SEVs that look like “random mobile crashes.” Failure / symptom: Old app versions fail parse; support volume spikes; rollback doesn’t help already-published server. Resolution: Additive fields first; min supported version telemetry; consumer-driven Pact tests; kill switches / feature flags for behavior after contract publish. Seen at / similar to: Stripe API version pins; Google protobuf compatibility rules; Apple App Store slow-client reality for banks.


What interviewers probe

  1. How do you start a cross-team API? — contract-first story.
  2. Breaking change process for a public API or event.
  3. How specs interact with TDD/CI
  4. Outbox/event schema evolution (link to Kafka).
  5. Experience with OpenAPI/protobuf — tooling, linting, breaking-change checks.
  6. When you’d skip heavy SDD — exploration vs platform API.
  7. AI coding — how you keep generated code honest.
  8. Production incident caused by spec drift — detection and prevention.

Senior-level expectation: Specs as engineering leverage; proportional ceremony; contract testing awareness; clear ownership of the source of truth.


Pitfalls

  • Spec written once then abandoned (“doc rot”)
  • Speculating far beyond known requirements
  • Multiple conflicting sources of truth (Confluence ≠ repo OpenAPI)
  • Generated code committed without review
  • Specifying implementation details instead of behavior (over-constrains)
  • No error model — happy path only
  • AI output accepted without contract/CI gates
  • Breaking event fields without registry compatibility checks

Cross-references

Interview reference — explanation quality and judgment, not syntax memorization.