Agents-as-Code With YAML Configuration in Git Repositories
Agent configuration moves from scattered prompts into version-controlled YAML and Markdown files.

The move from hand-configured servers to declarative infrastructure-as-code did more than speed up provisioning. It changed who could review a change, who could audit it after the fact, and who could undo it if it went wrong, and that governance shift is what let infrastructure scale past what any one team could hold in its head. Before tools like Terraform, Ansible, and CloudFormation took hold, a server's actual state lived in the memory of whichever ops engineer had last touched it, scattered across bash scripts nobody had fully documented. A bad change often couldn't be undone unless someone happened to have taken a snapshot beforehand. Terraform, Ansible, and CloudFormation changed that by making infrastructure intent something you could read: reviewable in a pull request, diffable across versions, and revertable to a known good state.
The organizational consequence mattered as much as the technical one. Once the config file, not a particular engineer's memory, became the source of truth, infrastructure ownership could spread across teams without each new contributor needing years of tribal knowledge first. AI agent configuration is now approaching the same turn. Agent behavior, which has mostly lived buried inside application code or scattered across prompts and scripts, is moving into version-controlled, declarative files that sit in the repository next to the software those agents operate on. Files like CLAUDE.md, AGENTS.md, and agents.yaml are starting to do for AI agents what Terraform configs did for cloud infrastructure: they separate what a system is supposed to do from how it happens to be built. The mechanism driving both shifts is identical. Intent gets written down declaratively, a tool enforces it, and a human reviews it in a pull request before it ships.
Agents-as-code: the files and their roles
A common structure appears repeatedly once the branding differences between agent frameworks are stripped away: a YAML manifest handles model selection and runtime configuration, a Markdown file carries behavioral identity, and a separate rules file holds constraints, all of it checked into the repository like any other source file. The specific filenames vary by tool. CLAUDE.md, AGENTS.md, CrewAI's agents.yaml and tasks.yaml, and Gitagent's SOUL.md, RULES.md, and agent.yaml are all formats in active use, differing in name but not in underlying shape.
YAML won the configuration half of that shape for reasons that have nothing to do with AI specifically. It's human-readable, it diffs cleanly in version control, every CI/CD pipeline already parses it without extra tooling, and it's already the language engineers use for Kubernetes manifests, GitHub Actions workflows, and Terraform itself, so adopting it for agent config asks nothing new of a team that already runs modern infrastructure. Markdown won the behavioral half for a different but equally practical reason: behavioral rules read as prose, not as structured data, and Markdown keeps that prose legible to the humans who have to review it when it changes in a pull request. Splitting the two formats this way lets a reviewer scan a YAML diff for a changed rate limit in seconds, and read a Markdown diff for a changed behavioral rule the way they'd read a comment change in code.
The standard furthest along this path is AGENTS.md, stewarded by the Agentic AI Foundation under the Linux Foundation. Since its release in August 2025, more than 60,000 open-source repositories have adopted it, and as of June 2026 at least 28 tools list native support for the format. That adoption curve points to a real operational problem the format solves. Every AI coding session starts from zero: without a config file injecting project context the moment a session begins, an agent has no idea what conventions the team follows, what architecture decisions already got made, or which libraries are preferred over which alternatives. A checked-in config file fixes that the same way a README or a style guide fixes it for a new human hire, except the agent reads it automatically, every time, with no chance of skipping it out of habit.
What a real agents-as-code repository looks like
A fully built-out agents-as-code repository tries to hand an agent everything a senior engineer would otherwise carry around in their head: who the agent is, what rules bind it, what reusable skills it can draw on, what the team already knows, and how the codebase it's working in connects to everything around it. All of that lives version-controlled and open to pull-request review. One documented production structure built around this idea organizes a repository with CLAUDE.md read first at every session start to establish identity and behavioral rules, then branches out into functional directories. A skills/ directory holds more than 20 reusable task definitions covering things like bug analysis, workspace cloning, and Jira triage. A workflows/ directory defines scheduled or triggered multi-step processes, such as a daily bug triage run or a weekly pull request report. A repos/repos.yaml file maps core repositories, their dependencies, and cross-team repositories, letting the agent trace call chains and dependency relationships across more than 20 projects. A team-members/ directory keeps a roster with component ownership and Slack IDs.
That repos.yaml file deserves particular attention because it solves a problem common in real engineering organizations: almost no task lives inside a single repository. A bug might surface in a request to the API layer, but its root cause sits in the auth service, and fixing it properly requires a change to a shared SDK. A map of that dependency graph checked into the agent's own configuration tells an agent working a ticket that the fix it needs may live somewhere else.
The runtime side of this pattern appears clearly in the.agents.yaml file used in AndrewAltimit's template-repo, where agent selection and operational policy sit in the same document. It lists which agents are enabled, sets per-agent rate limits by requests per minute and per hour, names trusted sources, and specifies agent_admins who are authorized to make changes, while security settings like autonomous mode and sandbox requirements appear as comment blocks rather than discrete keys. What stands out in that file isn't just what it permits, it's what it preserves: a disabled agent isn't deleted from the config, it's commented out with a dated note explaining why, the same discipline engineers already apply when they deprecate a block of application code rather than silently delete it.
That discipline is the real payoff of treating agent configuration as a repository instead of a runtime black box. Every skill added, every workflow changed, every solution later deprecated has a name and a timestamp attached to it, because git log and git blame can answer those questions the same way they answer them for any other code change. The agent's entire operational knowledge exists as files, and files have histories.
Governance and agent config as code
Version-controlling agent configuration is the precondition for governance for developers who need more than clean repositories. It's the precondition for governance, because nobody can audit, scope, or revert behavior they can't first see written down somewhere. The visibility gap closed by that step is large, by the report's own numbers. Research behind the 2026 CISO AI Risk Report found that a large majority of enterprise security leaders lack full visibility into their own organizations' AI identities, and most don't enforce access policies for those identities at all, even while those same AI systems hold access to core business platforms including ERP, CRM, and financial systems. An agent with standing access to a company's financial system and no enforced access policy is not a hypothetical risk; it is, by that survey's own numbers, closer to the norm than the exception.
Agents-as-code attacks that visibility gap directly, because every behavioral constraint, every tool the agent is permitted to touch, every rate limit, and every human authorized to change any of it becomes a line in a file that passed through review before it ever reached production. That matters beyond internal governance, too. The EU AI Act's high-risk track became fully operative on August 2, 2026, and agentic systems deployed in critical infrastructure, education and vocational training, employment, and access to essential private and public services, including credit scoring and health insurance, now carry mandatory logging, human oversight requirements, and conformity assessments. Every one of those obligations depends on the agent's behavior being legible and traceable to begin with. A config file nobody reviewed and nobody can diff cannot satisfy a conformity assessment; a config file that moved through a pull request can at least produce one.
The 2026 International Scientific Exchange on AI Safety arrived at a similar conclusion from the research side, identifying ten foundational principles for managing agentic risk: least privilege, traceable identity, auditability, interruptibility, validated deployment, adversarial resilience, multi-agent stability, runtime assurance, legibility, and human oversight. Every one of those principles is easier to enforce when an agent's permissions and behavioral constraints live in a file that already went through code review. Least privilege becomes a line item in a YAML rate limit. Traceable identity becomes a commit author. Interruptibility becomes a feature flag someone can flip and revert.
The pull request itself becomes the actual governance checkpoint. A change to what tools an agent can call, or to the rules that bound its behavior, gets reviewed by a person before it reaches production, using the exact workflow the team already runs for every other code change. When an agent starts behaving in a way nobody expected, the first diagnostic step is the same one engineers already know: git diff the config history and find what changed. Rolling back is a git revert, not an emergency meeting and a manual reconstruction of what the system used to look like.
The objection that an agent's behavior isn't fully determined by its config, since the same file can still produce different outcomes on different runs, is an argument for pairing configuration with better observability, not an argument against writing the configuration down. The next two sections take up what that observability actually looks like.
Where the IaC analogy breaks down
The parallel to infrastructure-as-code holds up until it runs into a basic property of language models that infrastructure never had to deal with. Applying the same Terraform config twice produces identical infrastructure, every time, by design. Applying the same CLAUDE.md twice may produce different agent behavior, because LLM non-determinism, model version drift, and the accumulated effects of context are simply not captured anywhere in a config file. Research on this question states the problem directly: the tension between declarative configuration and emergent behavior is, in its own words, "the central challenge of reproducible AI agent deployments".
The tension has concrete mechanisms. An agent can run against an identical configuration on two separate occasions and reason its way to two different conclusions because the underlying model received a silent update between runs, because effects within the context window shifted what the model attended to, or because stochastic sampling sent the model down a different branch of reasoning entirely. None of those three causes appears as a line in a YAML file. A config can tell an agent which model to call and which tools it may use; it cannot guarantee what that model will decide to do with them on any given run.
This limitation does not undercut the case for agents-as-code so much as it clarifies what that case actually claims. A declarative config establishes intent: what the agent is supposed to do, what it's allowed to touch, and who is accountable for changing those boundaries. It does not, by itself, guarantee that execution matched that intent on any particular run. Knowing whether it did requires watching the agent work.
Observability and cost controls for agents running from version-controlled configuration
Observability and cost control become tractable in a specific way once agent configuration is version-controlled and declarative: every session can be tied back to a named config version, every tool call can be traced to a permission that was explicitly declared, and every dollar of spend can be budgeted at the config level before anything actually runs. That last point is the practical payoff of the whole pattern. Because sessions are attributable to a specific config and a specific trigger, rather than existing as anonymous API traffic, spend can be governed per session, per developer, or per period, instead of only being watched in aggregate after the fact.
The telemetry side of this has settled on a shared standard rather than a patchwork of proprietary dashboards. The field has converged on OpenTelemetry as the vendor-neutral telemetry standard, so trace data coming out of an agent session can feed directly into the same monitoring infrastructure a team already runs for everything else. On the financial side, the stakes are not abstract. LLM API spend was tracked doubling over a six-month span as of mid-2025, and a single agent loop that gets mis-deployed can burn through a quarter's budget overnight if nothing is watching it in real time. The operational fix that's emerged in response is threshold alerts set at multiple spend levels before the budget runs out, arriving ahead of a monthly bill.
That kind of control can be written directly into the same config file that defines the agent. The.agents.yaml pattern from the AndrewAltimit template-repo bakes budget discipline into the configuration itself: per-agent rate limits set by the minute and by the hour, a subprocess_timeout, a memory_limit_mb cap, and a max_prompt_length, all declared before a single agent session ever starts. Cost and quality don't have to be tracked separately, either. Braintrust, used by companies including Notion, Stripe, Vercel, and Instacart, ties cost data directly to quality evaluation: when it finds a given step in an agent's workflow burning a large share of the token budget for only marginal gains in quality, it lets a team swap in a cheaper model and check the quality impact before that change ever reaches production.
Sandbox and infrastructure requirements for agents running from repo-based configuration
A YAML file can declare every constraint a team wants an agent to respect, but the file itself enforces nothing. An agent still needs to run inside an execution environment built to actually hold those constraints, because without an isolated sandbox, everything written in the config is aspirational rather than real. The production.agents.yaml configs referenced earlier make this explicit: require_sandbox: true and autonomous_mode: true consistently appear together. Autonomous operation is only granted alongside sandboxed operation, never as an independent setting that a team can turn on without the other.
The engineering lesson behind that pairing is straightforward once it's stated. An agent that can write code but can't run the tests, query the services it's modifying, or reach the APIs it depends on can't actually close the loop on its own work, no matter how well its behavior is specified on paper. It can propose changes but never verify them. An agent granted the ability to run tests, query services, and reach live APIs, but operating outside any sandboxed boundary, can apply those same capabilities somewhere the config never intended. A well-structured repository of YAML and Markdown files establishes what an agent is supposed to do and who is accountable for changing that definition. Whether the agent can only do what the file says, and nothing more, is a property of the sandbox it runs inside, not of the file itself.
Sources
- template-repo/.agents.yaml at main · AndrewAltimit/template-repo
- GitHub - open-gitagent/gitagent: A universal git-native AI agent framework. Your agent lives inside a git repo — identity, rules, memory, tools, and skills are all version-controlled files.
- Agent Configuration as Code: Declarative Definitions and Reproducible AI Agent Deployments
- The Rise of Agent Infrastructure as Code: Why Securing AI Agents Starts in the Repository - Cycode
- Agent‑Ready Repo Structure (2026)
- AI Agent Compliance and Governance in 2026: A Practical Guide
- What are the best AI agent observability platforms in 2026?
- Agentic AI Observability: A Practical Guide for 2026 - Coralogix


