Home › Guides › Spec-Driven Development
Tech Explained · 2026What Is Spec-Driven Development in 2026? How It Works, SDD vs Vibe Coding and 6 Steps to Ship With It
Spec-driven development is a workflow where a written, version-controlled specification, not the prompt, is the source of truth an AI coding agent builds from. You write a constitution, a spec, a plan and a task list first, then let the agent implement. GitHub's Spec Kit supports 30 or more coding agents.
- The spec is the artefact, not the chat. In spec-driven development your reviewable output is a markdown file in git, which means a teammate can diff it and argue with it before any code exists.
-
Five commands carry the whole method. Spec Kit runs
/constitutiononce per repo, then/specify,/plan,/tasksand/implementper feature. - The tooling is cheap, the discipline is not. Spec Kit is free and open source; Kiro's pricing page lists a free tier of 50 credits and Pro at USD 20 per user per month.
- AGENTS.md is the standard worth adopting first. Reported across 60,000 or more open source repositories, and now governed by the Agentic AI Foundation under the Linux Foundation.
- It slows down small work. For a one-file bug fix, writing a spec costs more than it saves, and most teams that abandon SDD abandoned it on exactly that kind of task.
- Hiring signal is already there. JetBrains research published in August 2026 reported 90% of professional developers using AI coding agents weekly, which is why "can you supervise an agent" is becoming an interview question.
You describe a feature to your coding agent in three sentences, it produces 900 lines across eleven files in four minutes, and everything looks plausible. Two days later a reviewer asks why the permissions check runs after the database write, and nobody can answer, because the reasoning existed only in a chat window that has since scrolled away. That gap between "the agent produced code" and "the team understands the code" is the problem spec-driven development was invented to solve, and it is why the practice went from a niche GitHub repo to a default workflow across most major coding tools inside about a year.
What Is Spec-Driven Development, and Why It Appeared Now
Spec-driven development, usually shortened to SDD, treats an executable, version-controlled specification as the single source of truth for a feature. You write down what the system must do, in structured markdown, with acceptance criteria and explicit non-goals. The plan and the task breakdown are derived from that spec. Code is derived from the tasks. When behaviour and spec disagree later, the spec wins and the code changes.
The reason it arrived in 2026 rather than 2016 is that the bottleneck moved. When humans typed every line, the specification was the slow part and the code was slower still, so teams shipped two-sentence Jira tickets and let senior engineers hold the design in their heads. Agents inverted that. Generating 900 lines is now the cheap step. Deciding what those 900 lines should do, and proving afterwards that they do it, is the expensive step.
Andrej Karpathy coined "vibe coding" in February 2025 for the practice of accepting agent output without reading it closely. By 2026 the same community that adopted the phrase had largely moved past it for anything headed to production, and spec-driven development is the discipline that replaced it.
Three numbers that explain the shift
Agent use is near-universal; the instruction files that steer agents are now an ecosystem, not a hack.
Figures from JetBrains' developer research post, the AGENTS.md project and the Spec Kit repository, checked 26 September 2026.
How Spec-Driven Development Works: The Constitution to Implement Pipeline
Every flavour of SDD implements roughly the same five stages. GitHub's Spec Kit names them most explicitly, so it is the useful reference implementation even if you end up using something else.
GitHub Spec Kit: the five stages and what each one produces
Constitution. Run once per repository. It writes constitution.md, which captures the rules the agent must never break: your test framework, your layering rules, whether raw SQL is allowed, how errors surface. This is the file that stops the agent from helpfully introducing a second ORM.
Specify. Produces spec.md: what the feature does and for whom, in user-visible terms, with acceptance criteria and explicit non-goals. No technology choices here. If your spec mentions Redis, you have skipped ahead.
Plan. Produces plan.md: the technical approach, data model changes, interfaces, and the trade-offs you accepted. This is where Redis belongs.
Tasks. Produces tasks.md: atomic, independently reviewable steps, each small enough that a failed one is cheap to throw away.
Implement. The agent works the task list. You review per task, not per feature.
A detail that trips people up: these are not compiled CLI binaries. Spec Kit's commands are markdown prompt templates that your agent reads, which is exactly why one toolkit can drive Copilot, Claude Code, Cursor, Gemini CLI and Codex CLI without writing an adapter for each. Installation is a single command, uv tool install specify-cli, followed by specify init my-project --integration copilot.
The spec-driven development pipeline, with its review gates
Note where the human sits: between the stages, editing markdown, not reading diffs at the end.
Stage names follow GitHub Spec Kit's documented command pipeline, read 26 September 2026.
If this pattern feels familiar, it should: it is the same separation of intent, orchestration and execution that shows up in agent frameworks generally. Also read: What Is an AI Agent Harness in 2026? for how that layering works at runtime rather than at authoring time. Teams building agents rather than merely using them practise exactly this decomposition in the live AI Engineer course on generative AI, RAG and agents.
Spec-Driven Development vs Vibe Coding: What Actually Changes
The honest comparison is not "careless versus careful". Vibe coding is genuinely faster for a class of work, and pretending otherwise is how you lose the argument with your team. What changes is where the reviewable artefact lives and how failures surface.
| Dimension | Vibe coding | Spec-driven development |
|---|---|---|
| Source of truth | The chat transcript, which is not in git |
spec.md, committed and diffable |
| Where you review | The final diff, after 900 lines exist | The spec and plan, before code exists |
| Cost of a wrong assumption | Rewrite the implementation | Edit two paragraphs of markdown |
| Best on | Prototypes, spikes, throwaway scripts, single-file fixes | Multi-file features, permissions, billing, anything with compliance exposure |
| Onboarding a new teammate | They read code and guess intent | They read the spec, then the code |
| Handling agent context loss | Re-explain from memory each session | Agent re-reads the spec and plan files |
| Typical failure mode | Code that works but solves the wrong problem | A 40-page spec nobody reads, written for a two-day task |
| Test strategy | Written after, if at all | Acceptance criteria exist before implementation starts |
| Time to first running code | Minutes | Often an hour or more of spec work first |
| Auditability for a regulator or client | Weak; reasoning is not recorded | Strong; intent and trade-offs are in version control |
One structural point people miss. SDD does not require an agent at all. A two-person team writing every line by hand gets most of the rework reduction from the spec-first habit alone. The agent just makes skipping the spec much more tempting, because the punishment for skipping it arrives later and looks like someone else's bug.
Spec-Driven Development Tools in 2026 and What They Cost
By 2026 most major coding tools ship a flavour of this: GitHub Spec Kit, AWS Kiro, Claude Code, Cursor, OpenSpec, BMAD, Tessl and Google Antigravity among them. The practical choice is narrower than that list suggests, because they split into three groups: a toolkit that sits on top of whatever agent you already use, an IDE that bakes the method in, and the plain-markdown approach with no toolkit at all.
AGENTS.md: the smallest version of spec-driven development that works
Before you install anything, write an AGENTS.md file. It is a plain markdown file at your repository root telling any compliant agent how your project builds, tests and behaves. First released in August 2025, it is now reported across more than 60,000 open source repositories and its governance sits with the Agentic AI Foundation under the Linux Foundation, alongside MCP. Twenty useful lines in AGENTS.md will change more agent behaviour than a week of prompt tweaking.
| Option | Shape | Cost, as listed September 2026 | Pick it when |
|---|---|---|---|
| GitHub Spec Kit | Open source toolkit, agent-agnostic, installed with uv tool install specify-cli
|
Free; you still pay for whichever agent runs it | You already have a coding agent and want the method without switching editors |
| AWS Kiro | Agentic IDE with specs as a first-class object; generally available since 17 November 2025 | Free tier of 50 credits per month; Pro USD 20, Pro+ USD 40, Pro Max USD 100, Power USD 200 per user per month; add-on credits USD 0.04 each | You want the spec workflow enforced by the tool rather than by team discipline |
| Plain AGENTS.md plus a specs folder | Convention, no tooling; reported across 60,000 or more repositories | Free | You are introducing this to a sceptical team and need the smallest possible first step |
| Vendor-specific instruction files (CLAUDE.md and similar) | Agent-specific project memory, often alongside a plan mode | Included with the agent subscription | Your whole team standardised on one agent and you want its deepest integration |
My recommendation for a team of five or fewer: start with AGENTS.md and a specs/ directory, run it for three features, and only then reach for Spec Kit. Adding a toolkit before the habit exists gives you scaffolding nobody fills in. I would not put Kiro's Pro tier on a company card until at least two engineers can point at a feature the spec-first pass visibly saved, because the credit model makes it easy to spend USD 40 a month discovering you did not want the workflow.
The facts worth memorising before your first spec
Four figures that decide whether this costs you anything to try.
From the Spec Kit repository, Kiro's pricing page and the Linux Foundation's AAIF announcement, checked 26 September 2026.
A Worked Example: One Permissions Rewrite, Two Ways
Take an illustrative scenario. A four-person product team at a mid-size Bengaluru B2B SaaS company maintains a Django codebase started in 2019. Their task this sprint is replacing an ad hoc if user.is_admin pattern, scattered across 40 view functions, with proper role-based permissions. Three roles, two of which can be scoped to a single customer account.
Done the fast way, the lead prompts an agent, gets a permissions module and 40 edited views in an afternoon, and opens a pull request. QA then finds that the scoped roles behave differently on two endpoints, because "scoped" was never defined anywhere. The reviewer cannot tell which behaviour was intended. Everyone re-litigates the design in pull request comments, which is the most expensive place to hold a design discussion.
Done spec-first, the lead spends 50 minutes writing a spec that names the three roles, defines scoping precisely as "visible rows are filtered by account_id, and writes are rejected with 403", and lists a non-goal: no per-field permissions in this release. That last line is the one that saves the sprint, because it is the thing an agent will otherwise invent. The plan step then forces a decision the fast version never surfaced: does the permission check live in a middleware, a decorator, or the query layer? Pick the query layer and the scoping falls out for free.
There is a real published version of this pattern too. XB Software's write-up on applying Spec Kit inside a mature legacy codebase reports a task they had estimated at a week finishing in roughly half the time, with the spec pass surfacing requirement gaps the team had not noticed. Their own caveat matters as much as the result: it worked because experienced engineers reviewed each stage. Treat that as one team's documented experience rather than a benchmark.
-
The spec becomes a design document. If your
spec.mdnames a database, you have merged specify and plan, and you have lost the ability to change the technology without rewriting the requirements. - Nobody updates the spec after implementation. Two sprints later the spec is fiction, which is worse than no spec, because people trust it. Make "spec updated" part of your pull request checklist or do not bother.
- The task list is too coarse. "Implement permissions" is not a task. If a task cannot fail in a way you can see in one test run, split it.
-
The constitution is aspirational. Writing "all code must have 100% coverage" in
constitution.mdwhen your repo sits at 40% teaches the agent to ignore the file.
Where Spec-Driven Development Is Not Worth It
Here is the caveat the tool vendors will not lead with. For a large share of daily work, spec-driven development is net negative. A one-file bug fix, a copy change, a dependency bump, a spike you intend to throw away: writing a constitution and a spec for any of these costs more attention than it returns, and forcing the ceremony is how you get a team that quietly stops doing it for the features that actually needed it. Karpathy's fast loop is still the right loop for exploration.
The second limit is that a spec cannot make a vague requirement precise on your behalf. If the product owner genuinely does not know whether scoped admins should see deleted records, the agent will produce a confident spec section about it, and that section will be a guess wearing formal clothing. SDD surfaces ambiguity earlier; it does not resolve it.
The rework-reduction numbers circulating in community write-ups, often quoted as 60% to 80% fewer rework cycles, should be read as self-reported and unaudited. The direction is well supported, the magnitude is not. Industry commentary also repeats the older software-engineering finding that defects caught at planning cost several times less than defects caught in implementation, which is plausible and consistent, but not something to quote as a precise multiple.
6 Steps to Start Spec-Driven Development This Week
A first pass you can finish in one working week
Each step ends in a file committed to your repository, so progress is visible to your team.
Write AGENTS.md by hand
Twenty lines, no toolkit. Build and test commands, directory conventions, the two or three rules you are tired of repeating in code review. Commit it at the repo root.
Day 1, 30 minutesPick a feature that already hurts
Choose something multi-file with at least one ambiguous requirement. Permissions, pricing rules and anything touching an audit trail are ideal. Do not pick a bug fix.
Day 1Write the spec before opening the agent
User-visible behaviour, acceptance criteria, and a non-goals list. Force yourself to write three non-goals. Then have one teammate read it and mark every sentence they could interpret two ways.
Day 2, 1 hourDerive the plan and argue about it
Ask the agent for two plans, not one, and make it state the trade-off each accepts. Choosing between two written options is far easier than critiquing a single confident answer.
Day 2Break work into tasks you can reject cheaply
Each task should be reviewable in under ten minutes and revertible in one commit. If the agent hands you four tasks for a two-week feature, ask again.
Day 3Add specs to your definition of done
One line in the pull request template: "spec updated, or state why not". Then run Spec Kit on the next feature, now that you know what the files are for.
Day 4 to 5Sequence built around the Spec Kit documented pipeline and the AGENTS.md convention, both read 26 September 2026.
Step 3 is the one that transfers to every other part of AI engineering, because writing unambiguous instructions for a non-human reader is the same skill whether the output is code, a retrieval pipeline or a tool-calling agent. It is the skill behind context engineering, and it is examined directly in the CCDV-F developer foundations prep course, which drills building and debugging with Claude rather than just prompting it. If your interest is the architecture side of the same problem, the CCAR-F architect foundations track covers how these decisions are documented and defended.
Specs are half the job. Building the agent that reads them is the other half.
100+ hours of live training over 16 weeks on generative AI, RAG and agents that plan, call tools and act. Certified on both Microsoft Copilot Studio and Claude Code, the two stacks Indian job postings actually name.
Explore the course
Is Spec-Driven Development Worth Learning in 2026?
For anyone whose code reaches production, yes, and the reason is a hiring reason as much as a craft one. When 90% of professional developers use coding agents weekly, writing code stops being the differentiator and supervising output becomes it. Interviewers have noticed. "Show me how you would specify this feature for an agent, then how you would verify it did what you asked" is a question that separates candidates far more cleanly than a reversed linked list.
It matters more in some roles than others. If you are heading toward forward deployed engineering, where you build inside a client's constraints and have to defend every design choice in a room, a written spec is not overhead, it is your evidence. In MLOps the same instinct shows up as pipeline contracts and model cards. Teams rolling agents out across a whole organisation hit it as policy rather than practice, which is the ground covered by the Microsoft 365 Copilot and Agent Administrator (AB-900) course. If you want to see how these tracks fit together before committing, the full certifications overview lays out the sequence.
The skill is also cheap to acquire relative to its payoff, which is rare. You can learn the method in a week from free tooling and a public repository. What takes longer is building the taste to know which features deserve a spec, and that only comes from running it on real work. Also read: What Are Agent Skills in 2026?, which covers the closely related question of how you package reusable instructions for an agent instead of re-explaining them.
In your position I would do exactly two things this month, and skip everything else being written about this. Write AGENTS.md today, because it costs half an hour and improves every agent session afterwards. Then run one spec-first feature, pick the messiest thing in your backlog, and notice specifically whether the spec review caught something QA would have caught. If it did, you have your answer and you do not need anyone's benchmark. The harder half is not the spec, it is building the agent that has to live inside it, and that is the work the AI Engineer course takes you through live over 16 weeks; if you want to judge the teaching before you commit, sit in on one of the free webinars first.
Related guides
- AI Engineer Roadmap 2026: 7 Steps to Land Your First Role in India where spec discipline sits in the wider sequence of skills you need to get hired.
- AI Engineer Interview Questions 2026: 18 Real Questions With Model Answers because "how do you verify agent output" now shows up in real interview loops.
- LLM Evaluation in 2026: How Evals Work and a Pipeline You Can Ship the verification half of the story, once your acceptance criteria exist.
- AI Engineer Salary in India 2026: Pay by Experience, City and the Skills That Move It what the agent-supervision skill set is currently advertised at.
- Agentic AI Jobs in India 2026: How a Hiring Surge Is Creating 6 New Careers which roles this workflow is becoming a stated requirement for.
Frequently asked questions
What is spec-driven development in simple terms?
Spec-driven development means you write and commit a specification describing what a feature must do before any code is written, then derive the technical plan, the task list and finally the code from that specification. The spec lives in version control, so it can be reviewed and diffed like code. When the code and the spec disagree later, you update the spec first and let the change flow forward.
Is spec-driven development the same as waterfall?
No, and the difference is scope. Waterfall specified an entire system up front and then forbade changes. SDD specifies one feature at a time, usually in one or two pages, and expects the spec to be edited as you learn. The loop is a sprint, not a year, and the spec is a living file rather than a signed-off document.
Do I need GitHub Spec Kit to do spec-driven development?
No. The method is a discipline, not a dependency. Most teams get the majority of the benefit from an AGENTS.md file at the repo root plus a specs/ folder of plain markdown. Spec Kit is useful once the habit exists, because it standardises the file structure and the command pipeline across your repositories and works with 30 or more agents.
What does spec-driven development cost to try?
Nothing beyond the coding agent you already pay for. Spec Kit is open source and free, and the AGENTS.md convention is free. If you want the method enforced by an IDE instead, Kiro's pricing page as of September 2026 lists a free tier of 50 credits per month and paid tiers from USD 20 up to USD 200 per user per month, with add-on credits at USD 0.04 each.
How long does writing a spec actually take?
For a normal multi-file feature, budget 45 to 90 minutes for the spec and another 30 for the plan once you are practised. That feels expensive on day one. It stops feeling expensive the first time the spec review catches a requirement gap that would otherwise have appeared in QA, because the fix at that point is editing a paragraph rather than reworking eleven files.
Does spec-driven development work on a legacy codebase?
Yes, and this is where it tends to pay off fastest, because legacy repositories are exactly where an agent's assumptions go wrong. The constitution file matters most here: write down the patterns the codebase already uses so the agent extends them rather than inventing a parallel approach. XB Software's published Spec Kit case study on a mature codebase reports a week-long task completing in about half the time, with experienced review at each stage as the stated condition.
What is AGENTS.md and how does it relate to CLAUDE.md?
AGENTS.md is an open convention for project-level instructions any compliant coding agent can read: build commands, conventions, constraints. It was first released in August 2025 and its governance later moved to the Agentic AI Foundation under the Linux Foundation, the same body that stewards MCP. CLAUDE.md and similar files serve the same purpose for one specific agent. Practically, many repositories keep AGENTS.md as the shared file and a short vendor file that points at it.
Which 360DT course covers building agents that use these workflows?
The AI Engineer course is the closest fit: 100+ hours live over 16 weeks on generative AI, RAG and agents that plan and call tools, taught Saturday and Sunday from 8 to 11 PM IST. If your interest is passing a developer credential on the Claude stack specifically, the CCDV-F prep course runs 40+ hours over 6 weeks.
About this guide. 360 Digital Transformation is an Authorized Training Partner of Anthropic and Microsoft. Other certification bodies, vendors and employers named here are not affiliated with us. Product features and pricing change often; figures cited were checked on 26 September 2026.




