Blog
Building with AI September 30, 2026 · EvalyMe

Your AGENTS.md Is a Product Spec, Not a Prompt

AI coding agents read your instruction file more often than any human reads your docs. Here is how to write AGENTS.md like an onboarding spec that actually changes agent behavior.

AGENTS.md AI agents context engineering documentation

The most-read document in your repository is written for a machine

Sometime in 2025, AGENTS.md went from an OpenAI Codex convention to a de facto cross-tool standard, now recognized by dozens of coding agents across tens of thousands of repositories. If you use Claude Code, the same role is played by CLAUDE.md. Whatever the filename, the function is identical: it is the standing brief every AI agent reads before it touches your code.

Most of these files are bad in a specific way: they are written like prompts (“write clean, maintainable code, follow best practices”) instead of like specifications (“there is no root package.json; run commands from the relevant subdirectory”).

The distinction matters because of what the file actually is. It is not a wish list of virtues. It is the onboarding document for a tireless, amnesiac contractor with excellent general skills and zero knowledge of your product. And unlike your README — which a junior developer reads once — your instruction file is read before every single task, by every agent, forever. That makes it the highest-leverage document in an AI-assisted repository, and it deserves engineering effort proportional to that.

What actually changes agent behavior

After a year of maintaining instruction files for a codebase spanning an iOS app, an Android app, an Express API, and two web surfaces, the pattern is consistent: agents follow concrete, situated instructions and ignore abstract ones. Four categories earn their place in the file.

Machine-checkable facts about your layout. The single most useful line in our file is that there is no root package.json and commands must run from the relevant subdirectory. Without it, agents repeatedly attempt root-level installs and invent workflows that do not exist. File facts beat adjectives.

Tool gotchas with a reason. Example from our file, verbatim in spirit: bun run test invokes the configured Vitest runner, while bun test invokes Bun’s different built-in runner, so append -t 'test name' to select one Vitest case. Two commands that look interchangeable and are not. Every ecosystem has several of these; each one you write down saves a dozen confused agent sessions.

Invariants with their why. “Never read the legacy pro_users table to authorize paid features — RevenueCat is the sole entitlement authority.” An agent told only “don’t” will route around the rule when the cleaner-looking option appears. An agent told why will recognize the class of mistake. The same applies to “never edit an applied migration” and “never commit ProductionConfig.xcconfig” — each carries a one-line consequence so the agent can generalize.

Behavioral guardrails discovered the hard way. Our deploy script’s backend check runs the test suite, but landing deployment only runs the build — so the file instructs running the relevant checks explicitly before deploying the landing site. Nobody would know that from the tool names. It is exactly the kind of tacit operational knowledge that used to live in one engineer’s head and now belongs in the file agents read.

What to leave out

Instruction files fail from bloat more often than from omission. Every token in the file is loaded into context on every task, competing with your actual code for attention — the practice now called context engineering. Three things to cut:

  • Virtue words. “Clean, maintainable, well-tested code” is noise; it changes nothing an agent does.
  • Anything derivable from the code. If package.json says it, the agent will read it. Do not spend context restating your dependency list.
  • Task-specific directions. The file is standing policy. Per-task intent belongs in your conversation, not frozen into every future context window.

A useful heuristic: if a competent new contractor would need to be told the fact on day one, and it is not obvious from the repository structure, it belongs in the file. Everything else is prompt poetry.

Tell, then show: an example of the difference

Incorrect example — an instruction that changes nothing:

Write high-quality TypeScript and follow established project conventions.

An agent reading this behaves identically with or without it. It is a mood, not a constraint.

Correct example — an instruction that changes behavior on the next task:

All three JavaScript packages use Vitest, but invoke it with bun run test, not bun test. There is no root package.json; run commands from backend/, app/, or landing/. Route tests inject in-memory dependencies and need no live database.

Same file, same reader — but the second version contains facts a model cannot infer from vibes and will apply within the minute.

Where this breaks: the spec nobody updates

The failure mode of instruction files is rot. The codebase evolves weekly; the file is edited rarely; eventually it confidently describes a repository that no longer exists — and a wrong instruction is worse than none, because agents trust it.

Two defenses work. First, treat instruction updates as part of definition-of-done: any change that alters commands, invariants, or workflows updates the file in the same commit. Second, keep the file short enough that a human can re-read the whole thing in five minutes during review; a file nobody re-reads is a file nobody corrects.

This mirrors the broader movement toward spec-driven development — GitHub’s Spec Kit, Amazon’s Kiro, the pattern Martin Fowler analyzed in late 2025 — where the written spec becomes the source of truth that both humans and agents work from. You do not need a formal methodology to benefit. Your instruction file, kept honest, is the lightweight version of the same idea.

A checklist for your instruction file

  • Does it state how to run and test the project, exactly, including the commands that look right but are wrong?
  • Does every invariant include its consequence, so agents can generalize instead of workaround?
  • Does it name the things that must never happen (secrets committed, applied migrations edited, legacy tables read) in actionable language?
  • Does it encode at least one piece of operational knowledge that exists nowhere else in the repository?
  • Is it short enough to re-read in five minutes?
  • Was it updated in the same commit as the last workflow change that affected it?

If your honest answer is that your file contains aspirations but no machine-checkable facts, it is not yet a spec. Your agents are onboarding themselves — and the quality of their decisions is showing it.

We maintain ours for the EvalyMe codebase the way we maintain the product: the instructions are the spec, the spec is tested, and both get updated when reality moves. It is the least glamorous document we own and the one we would rebuild first.