What Are Agent Skills in 2026? How SKILL.md Works, Skills vs MCP and 6 Steps to Build One
Agent Skills are folders of written instructions an AI agent loads only when a request matches them. Each folder holds a SKILL.md file with a name, a description and a Markdown procedure. The description sits in context always; the body loads on match. Claude Code truncates that listing text at 1,536 characters.
-
A skill is a folder, not a feature. One SKILL.md file with YAML frontmatter and Markdown steps, dropped in
.claude/skills/<name>/, is a complete skill. - Progressive disclosure is the whole trick. Only names and descriptions stay resident; the procedure loads on match, and bundled files load only when the procedure asks for them.
- Skills and MCP solve different problems. MCP gives an agent reach into live systems. Skills give it your team's procedure. Most production agents need both.
- The benchmark evidence is real but uneven. SkillsBench (arXiv 2602.12670) reports average pass rates moving from 33.9% to 50.5% with curated skills, and only +4.5 percentage points on software engineering tasks.
- The description field does the hiring. If the agent never picks your skill, the body never runs, and 90% of failures start there.
- It is a portable standard now. The spec was published as an open, vendor-neutral standard on 18 December 2025, and around 40 skills-compatible products were listed on the official showcase by June 2026.
You have written the same eleven-line instruction into your agent's prompt for the fourth month running: pull the claim PDF, check the policy number against the ledger, flag anything over Rs 2 lakh for a human, never auto-approve. It works. It also sits in every single request you send, burning tokens on conversations that have nothing to do with claims. That specific annoyance is what Agent Skills were built to remove, and understanding why the fix looks the way it does tells you more about production agent design in 2026 than any framework comparison will.
What Are Agent Skills, and Why They Appeared in 2026
Anthropic shipped Agent Skills for Claude in October 2025 and published the specification as an open, vendor-neutral standard at agentskills.io on 18 December 2025. By June 2026 the official showcase listed roughly 40 compatible products, including OpenAI Codex, GitHub Copilot, Cursor, Gemini CLI and VS Code. Microsoft documents skills support in its own Agent Framework. That breadth matters more than the launch did: a skill you write this quarter is not locked to one vendor's runtime.
Mechanically, a skill is a directory. Inside it sits SKILL.md: YAML frontmatter at the top between --- markers, then Markdown instructions below. Optional subfolders hold anything the instructions need, conventionally scripts/, references/ and assets/. Personal skills live at ~/.claude/skills/<skill-name>/SKILL.md and load across every project on your machine. Project skills live at .claude/skills/<skill-name>/SKILL.md, and because they sit in the repository, committing one ships it to the whole team the same way you ship a linter config.
The loading behaviour is the part worth internalising. Every skill's name and description are always in the agent's context so it knows what exists. The body is not. When a request matches a description, the rendered SKILL.md enters the conversation as a single message and stays there for later turns. Files in scripts/ or references/ are read only if the body tells the agent to read them. Three tiers, three different price points.
The three loading tiers of an Agent Skill
What you pay for on every request, what you pay for on a match, and what you pay for only on demand.
Tier behaviour and the 1,536 character listing cap are from the Claude Code skills documentation, checked 24 September 2026.
Inside SKILL.md: The Frontmatter Fields That Change Behaviour
Most SKILL.md tutorials stop at name and description. The official Claude Code reference lists more than a dozen fields, and four of them change what the agent is actually allowed to do. Here is the set worth knowing before you write your first one.
| Field | What it does | When you need it |
|---|---|---|
name |
Display name in the skill listing, defaults to the directory name | Only when the folder name reads badly to a human |
description |
What the skill does and when to apply it, used for trigger matching | Always. This is the field the agent reads to decide |
when_to_use |
Extra trigger context, appended to the description | When the trigger conditions are longer than the description |
allowed-tools |
Pre-approves tools for the skill's turn so the agent is not prompted | Unattended runs where permission prompts would stall the job |
disallowed-tools |
Removes tools from the pool while the skill is active | When you actually want to restrict, rather than pre-approve |
disable-model-invocation |
Stops the agent auto-triggering it, leaving manual invocation only | Destructive procedures such as a deploy or a data purge |
context: fork |
Runs the skill in an isolated subagent context | Long, noisy procedures you do not want polluting the main thread |
paths |
Glob patterns limiting when the skill activates | Monorepos where a skill applies to one service only |
model and effort
|
Override the model or reasoning effort for this skill's turn | Cheap model for extraction, expensive one for review |
compatibility |
Spec field for cross-platform notes, up to 500 characters | Skills you publish for other teams or other runtimes |
Read the allowed-tools row twice. It pre-approves, it does not sandbox. Teams reach for it expecting a permission boundary, then discover that removing a capability is a different field. If your compliance story depends on an agent never touching a shell, disallowed-tools is the field you want, and a proper execution boundary outside the agent is the control you actually want. Designing those boundaries is core CCAR-F architect territory rather than a frontmatter problem.
How Agent Skills Keep Context Costs Down
Return to the Pune logistics team from the top of this page: four people, one internal agent that reconciles proof-of-delivery scans against the billing ledger. Before skills, their system prompt carried the claims procedure, the escalation rules, the vendor code lookup table and the tone guidance for customer replies. Roughly 6,000 tokens, sent on every request, including the ones where someone just asked the agent to summarise a spreadsheet.
Split into four skills, the resident cost drops to four descriptions. The claims procedure loads only when a claim is mentioned. The vendor lookup table moves into references/vendor-codes.md and loads perhaps once a week. Same behaviour, a fraction of the standing cost.
There is a second-order effect that nobody mentions in the launch posts. Claude Code's auto-compaction re-attaches invoked skills after a conversation is summarised, keeping the first 5,000 tokens of each, with re-attached skills sharing a combined budget of 25,000 tokens. So a bloated SKILL.md does not just cost you once. It competes for a fixed budget every time the conversation compacts, and the skill that loses that competition is the one that silently stops steering the agent. Keep bodies short and push detail into referenced files, which is the same discipline covered in our guide to context engineering.
Where the context budget goes on a 12-skill agent
Skills are rarely the expensive part. Tool output usually is.
Illustrative split for a mid-size internal agent, not a measured average. The 5,000 and 25,000 token compaction budgets it is built around come from the Claude Code documentation, checked 24 September 2026.
Agent Skills vs MCP: Which One Your Agent Actually Needs
The agent skills vs MCP question gets framed as a rivalry, and it is not one. MCP is a client-server protocol that connects an agent to live systems: a Postgres database, a Jira instance, a payments API. A skill is a text file describing how your organisation wants a job done. One is reach, the other is judgement.
| Criterion | Agent Skills | MCP servers |
|---|---|---|
| What it supplies | Procedure, standards, worked examples | Live access to data and systems |
| Artefact | A folder with SKILL.md and optional files | A running server speaking the MCP protocol |
| Always in context? | Name and description only | Tool definitions are typically always present |
| Where it executes | The agent's own environment | Its own process or container |
| Credential handling | None of its own, inherits the agent's | Holds the credential, agent never sees it |
| Freshness of data | Static until you edit the file | Live on every call |
| Who can author it | Anyone who can write Markdown | Someone who can write and host a server |
| Versioning | Git, like any other repo file | Server deploys and protocol versions |
| Typical failure mode | Never triggers because the description is vague | Latency, auth expiry, schema drift |
| Cost to try | Minutes, zero infrastructure | Hours, plus something to run it on |
Choose a skill when
The knowledge is procedural and stable: your SQL naming conventions, the seven checks before a release, the exact structure of a board update, how to redact a customer document. If a competent new hire could learn it from a one-page wiki article, it is a skill.
Choose MCP when
The answer changes between invocations, or a credential must not reach the model. Current inventory, an open ticket queue, a customer's payment status. If the data changes between invocations, you need MCP, and no amount of Markdown fixes that. Our MCP tutorial walks the protocol from zero if that side is new to you.
In practice the serious 2026 builds use both, and the division is clean: MCP fetches the claim record, the skill decides what counts as a red flag and who signs off. That pairing, plus the evaluation loop that keeps it honest, is what 360DT's AI Engineer course has students build live rather than read about, using the same Claude Code and Copilot Studio stacks that job descriptions name.
Also read: What is an AI agent harness, for where skills sit inside a full orchestration layer.
Do Agent Skills Actually Improve Results?
Here is a rare thing in agent tooling: a public benchmark. SkillsBench (arXiv 2602.12670) evaluates 87 tasks across 8 domains, each paired with curated skills and deterministic verifiers. Curated skills raised the average pass rate from 33.9% to 50.5%, a gain of 16.6 percentage points, with configuration-level gains reported between +4.1 and +25.7 points. The paper also reports that skills let smaller models match or beat larger models running without them on procedural work, which is the cost argument in one sentence.
Now the part you should sit with. The gains were wildly uneven by domain: healthcare +51.9 points, manufacturing +41.9, cybersecurity +23.2, and software engineering just +4.5. That pattern is not noise. Frontier models have absorbed enormous quantities of public code and already know how to write a Python service. They have absorbed very little about how your hospital codes a discharge summary or how your factory logs a batch deviation.
The honest caveat. If you are writing skills to make an agent better at generic coding, you will probably be disappointed, and you should spend that effort on evaluation and context design instead. Skills pay off where your procedure is specific, written down nowhere public, and currently lives in three people's heads. Test that assumption on your own tasks before you write thirty of them.
How to Build an Agent Skill in 6 Steps
You can do the whole thing in an afternoon. The order below matters more than the tooling, because step two is where most skills are won or lost.
From blank folder to a skill your team uses
Each step ends in something you can check, not something you can feel good about.
Pick one repeated instruction
Find a paragraph you have pasted into prompts more than three times this month. Deliverable: that paragraph, in a file, verbatim.
Write the description like a job ad
State what it does, when to use it, and the words a user would actually type. Deliverable: a description under 1,024 characters that names at least four trigger phrases.
Create the folder
Make .claude/skills/claim-review/SKILL.md, frontmatter between --- markers, Markdown steps below. Deliverable: a file the agent lists on startup.
Move the bulk out
Push tables, long examples and code into references/ and scripts/, then point at them by filename. Deliverable: a body under 500 lines.
Validate and trigger-test
Run claude plugin validate on the skills directory, then ask for the task in five different phrasings. Deliverable: five out of five triggers, or a rewritten description.
Commit it and set a review date
Check it into the repo so the team inherits it, and diary a re-read for the day your process changes. Deliverable: a merged pull request.
Paths, the validate command and the 500-line guidance follow the Claude Code skills documentation and the published spec, checked 24 September 2026.
Step two deserves its own paragraph. The description is the only thing the agent sees when deciding, so write it in the user's vocabulary, not yours. "Reviews insurance claim documents for policy mismatches, missing signatures and amounts above the approval threshold. Use when someone mentions a claim, a POD scan, a settlement or a reimbursement." That triggers. "Claims helper utility" does not. If you want structured practice at writing agent instructions that hold up under real prompts, that is the core of the CCDV-F developer foundations track, and the associate-level CCAO-F path covers the same discipline for non-engineers who write the procedures.
Skills plus MCP is the 2026 agent stack. Build both, live, in 16 weeks.
Most courses still teach you to chat with a model. This one teaches you to build agents that plan, use tools and act, certified on both Microsoft Copilot Studio and Claude Code.
Explore the course
What Usually Goes Wrong With Agent Skills
Six failures account for almost everything teams hit in the first month. None of them are exotic.
Six ways a skill quietly stops working
Most of these show up as "the agent ignored it", which is a symptom, not a cause.
The description is written for you
You described the implementation, not the request. The agent matches on user language, so a description full of internal system names never fires.
Trigger failureTwo skills claim the same ground
A "report writer" and a "board update" skill with overlapping descriptions make the choice a coin toss. Split by trigger phrase, not by topic.
OverlapThe body became a manual
Three thousand lines of policy in SKILL.md competes for the compaction budget and loses. Reference files exist for exactly this.
Bloatallowed-tools mistaken for a sandbox
It pre-approves tools for the turn. It is not a restriction mechanism, and treating it as one puts a hole in your threat model.
SecurityNobody owns the file
The process changed in March, the skill still describes February, and the agent now confidently enforces a rule your team dropped.
DriftIt was never evaluated
Skills change behaviour, so they need the same regression tests as a prompt change. Ship it with three golden cases or do not ship it.
No evalsFailure patterns compiled from the published spec guidance and vendor documentation, checked 24 September 2026.
- Treating a skill as a security control. A skill is text the model reads, and text the model reads can be argued with. Enforce approval limits, spend caps and destructive-action gates in code outside the agent, then use the skill to explain the rule.
- Importing skills you have not read. A shared skill can instruct an agent to run commands. Claude Code blocks command execution in skills synced from claude.ai on local machines from version 2.1.228, but a skill pulled from a random repository into your own directory carries no such protection. Read it like you would read a shell script.
Governance, once more than two people write them
Once skills spread past a single repo, they become policy artefacts: they encode who approves what and at which threshold. Review them in pull requests, keep an owner on each, and treat a skill edit as a change to how the agent behaves in production, because it is. The control set for that is the same one in our AI agent governance guide, and on the Microsoft side the equivalent admin work over Copilot agents is what the AB-900 Copilot administrator course covers.
Are Agent Skills Worth Learning in India in 2026?
Take a 2019-vintage Java developer in Bengaluru with about 30 spare hours a month. Writing skills is not, on its own, a hireable skill. It is a Markdown file. What is hireable is the surrounding judgement: knowing when a procedure belongs in a skill and when it belongs in a tool, sizing a context budget, wiring evaluation so a skill edit cannot quietly degrade output, and designing the permission boundary that a frontmatter field does not give you.
That bundle shows up in agentic AI job descriptions under titles like AI engineer, applied AI engineer and forward deployed engineer, where getting an agent working inside a client's messy environment is the entire job. Indian postings for these roles typically advertise wide bands, and industry reports suggest agentic AI hiring has grown sharply over the past two years, so treat any single number you see with suspicion and check live listings yourself.
If you want the fastest honest route: learn the skill format in a weekend, because it genuinely is that small, then spend the next three months on evaluation, retrieval and deployment. The MLOps engineer track covers the evaluation and monitoring half, the generative AI developer course covers the building half, and the full certifications overview shows how the Claude and Microsoft credentials stack. All live cohorts run Saturday and Sunday, 8:00 to 11:00 PM IST, which is the only part of this that a working developer usually cares about.
Also read: AI Engineer Roadmap 2026, for the seven-step sequence that sits around this.
The Verdict: Write Three Skills, Not Thirty
If you run an agent that anyone other than you depends on, write skills. The format costs nothing to adopt, the spec is open and portable across roughly 40 products, and the benchmark evidence says curated skills move pass rates meaningfully on procedural work. That is an easy call.
What I would not do is what most teams do in week one, which is convert every prompt they own into a skill and end up with forty overlapping descriptions that make triggering a lottery. Start with three: your most repeated procedure, your most dangerous one (with disable-model-invocation: true on it), and one reference-heavy lookup that proves the tier-three pattern to your team. Measure trigger accuracy on real phrasings before you write the fourth. The trade-off you are accepting is slower coverage in exchange for an agent whose behaviour you can still predict in month six, and that is a trade worth making every time. If you want that judgement built properly rather than picked up in fragments, the AI Engineer course is the place to start.
Related guides
- LLM Evaluation in 2026 the regression tests that stop a skill edit from silently degrading your agent.
- Enterprise AI Agents in 2026 why pilots stall at the boundary between demo and production, which is where skills earn their keep.
- CCDV-F Exam Prep 2026 a six-week plan if you want the developer credential behind this work.
- AI Engineer Interview Questions 2026 the questions that test whether you understand agent context, not just agent syntax.
- AI Engineer Salary in India 2026 what the roles that use this stack actually advertise, by experience and city.
- A2A Protocol Explained the layer above, for when your agents need to talk to each other.
Frequently asked questions
What are agent skills in AI?
Agent skills are folders of written instructions that an AI agent loads only when a request matches them. Each folder contains a SKILL.md file with YAML frontmatter holding a name and a description, followed by Markdown steps. Optional subfolders hold scripts, reference documents and assets that load only when the instructions call for them.
What is the difference between agent skills and MCP?
MCP is a protocol that gives an agent live access to systems such as a database, a ticket tracker or a payments API, and the MCP server holds the credentials. A skill is a text file that tells the agent how your team wants a job done. Use MCP when the data changes between invocations; use a skill when the procedure is stable. Production agents usually run both.
How many tokens does an agent skill use?
Only the name and description stay resident, which is roughly 100 tokens per skill. The body loads on match and should be kept well under 5,000 tokens: the Claude Code documentation notes that after auto-compaction, re-attached skills keep the first 5,000 tokens each and share a combined 25,000 token budget. Files in reference folders cost nothing until they are read.
How do I write a good SKILL.md description?
Write it in the words a user would type, not in your internal system names. Cover what the skill does, when to use it, and three or four trigger phrases. Keep it under 1,024 characters, and remember that Claude Code truncates the combined description and when_to_use text at 1,536 characters in the listing. Then test five different phrasings of the same request and confirm it triggers on all five.
Do agent skills work outside Claude?
Yes. The specification was published as an open, vendor-neutral standard in December 2025, and by June 2026 the official showcase listed around 40 compatible products, including OpenAI Codex, GitHub Copilot, Cursor, Gemini CLI and VS Code. Microsoft documents skills support in its Agent Framework. Runtime-specific frontmatter fields will not always carry across, but the format does.
Can an agent skill run code?
It can instruct the agent to run scripts bundled alongside it, and skills execute in the agent's own environment rather than a separate container. That is why you should read any skill you import as carefully as a shell script. Claude Code blocks command execution in skills synced from claude.ai on local machines from version 2.1.228, but a skill you copy into your own directory carries no such protection.
Are agent skills worth learning for an AI job in India in 2026?
The format itself takes a weekend. What employers pay for is the judgement around it: deciding what belongs in a skill versus a tool, sizing context budgets, building evaluation so a skill edit cannot degrade output, and enforcing permissions outside the agent. Those appear in AI engineer and forward deployed engineer postings, which typically advertise wide bands in India.
Do agent skills replace RAG or fine-tuning?
No, they sit alongside both. RAG retrieves facts that change; fine-tuning shifts default behaviour across every request; a skill injects a procedure only when a task matches. If your problem is that the model does not know your data, that is retrieval. If the problem is that it does not follow your process, that is a skill.
About this guide. 360 Digital Transformation is an Authorized Training Partner of Anthropic and Microsoft. Other certification bodies, vendors and employers named here are not affiliated with us. Product features and pricing change often; figures cited were checked on 24 September 2026.




