Audit Trails for AI Agent Tool Calls and Diffs
Capturing the full decision chain—not just final commands—is essential for AI agent accountability.

A coding agent deleted a production database during a code freeze. That incident is documented, and it is the clearest illustration of what this piece sets out to establish: audit trails for AI coding agents have to record the full decision-and-action chain, not just the final command, because the final command alone cannot answer who authorized it, what the agent believed it was doing, or which delegation path led there.
Agent audit trails versus application logging
Teams that lacked a structured record of that database-deletion incident found themselves unable to answer basic questions after the fact. No one could say who had authorized the delete, what the agent thought it was doing when it ran, or which chain of delegation had handed it the authority to touch production during a freeze. That gap did not come from a missing log line. It came from a mismatch between what application logging was built to capture and what an agent's failure actually looks like.
Application logs record discrete events: a request came in, a status code went out, an error got thrown. That model works because the unit of harm in a traditional system is usually a single bad response or a single unhandled exception. An AI coding agent's unit of harm is a sequence. Agents call tools, hand off work to subagents, and take actions that change real systems, so a record that only captures the last API call misses everything that led up to it: what the agent retrieved, what it concluded from that retrieval, what policy got checked before the call was allowed, and what human, if any, was in the loop. Research on long-horizon agent auditing finds that the dominant failure mode in these systems is a wrong sequence of actions, not a single wrong answer, and each of these, a destructive command that runs in production, a prompt injection that spreads from one agent to another, an unauthorized escalation of privilege, is a chain of decisions that raises questions of causality, authority, and verification that a single log entry was never built to hold.
The tools most teams already run, OpenTelemetry spans, vendor dashboards, framework callback logs, do real work, and they help a cooperative operator trace why a pipeline failed. But they share two assumptions that break down for agent governance. They assume the internal log can be trusted as-is, and they assume one party both produces the telemetry and consumes it when something needs explaining. Neither assumption holds once a dispute involves a vendor, a customer, and a regulator who each hold a different slice of the evidence, or once the operator running the agent is itself a suspect in the incident. An audit trail built for that situation has to start from a different premise: the record must be complete enough to reconstruct intent and authority, and it must be verifiable by someone who has no reason to trust whoever produced it.
The eight fields every agent event must bind together
A single complete agent event is a structured record, not a log line, and it has to bind eight distinct pieces of information into one canonically serialized unit. Research on tamper-evident agent auditing defines these eight fields as intent, policy evaluation, human approval, execution, effects, context provenance, code provenance, and delegation provenance, and the schema requires all eight to appear in every event, not just the ones tied to an incident.
Intent is the reasoning step the agent produced before it acted. Without it, a record shows that a file was deleted, but it says nothing about whether the agent believed that deletion was the correct thing to do. Policy evaluation captures which governance rule the runtime checked before letting the tool call through, and what decision came back; this is separate from the system prompt, since it reflects what the runtime actually enforced rather than what the model was told. Human approval records whether an approval gate fired, who responded to it, and what exactly they signed off on, and this field carries legal weight: the EU AI Act's Article 14, on human oversight, requires exactly this kind of record of human-in-the-loop decisions. Execution is the specific tool or API the agent invoked, along with its arguments, a plain fact distinct from what the agent intended to do. Effects are the external consequences, the file writes, the API calls, the database changes that followed, and they need their own field because a tool call and its downstream effect are separable events that can diverge from each other.
Context provenance tracks which retrieval sources, memory entries, or earlier conversation turns shaped the agent's intent, and it is the field that lets a team trace a prompt injection as it moves from one agent to the next. Code provenance records the agent version, the model version, and the prompt version active at the time, which matters because a given behavior can be a known regression under one version and expected conduct under another. Delegation provenance, finally, records which parent agent or human trigger authorized a child agent's action, and in multi-agent workflows it is the field that resolves who is accountable when responsibility would otherwise spread across several agents with no clear owner. Every event should also carry a unique agent identifier and its delegated permission scope, so that a single record can be verified on its own, without pulling in surrounding context to make sense of it. Dropping any one of these eight fields makes a specific class of incident unreconstructable: without human approval, an unauthorized action is indistinguishable from an approved one; without delegation provenance, a multi-agent incident has no clear author.
The reasoning trace: the hardest field to capture, the most consequential to skip
Of the eight fields, intent, the reasoning trace, is the one that most clearly separates a forensic audit from a glorified access log, and it's also the one field that a gateway sitting between the agent and its tools cannot produce by itself. A gateway sees the final API request. It does not see the chain of deliberation the agent went through to arrive at that request, so a gateway-layer log can confirm that a tool was called but cannot explain why the agent decided to call it.
This matters most acutely in prompt-injection forensics, where the question is not just what happened but where a malicious instruction entered the chain and how it propagated. Structured queries run against the full eight-field schema, including context provenance and intent, deliver precise results on guardrail and delegation lookups; running the equivalent search as unstructured text over the same events gives results close to random. The reasoning trace is what turns a pile of logs into something a structured query can actually answer. Some teams push back on this, arguing that capturing thinking tokens adds latency and storage overhead that isn't justified for routine, low-stakes tasks, and that objection has real engineering weight behind it. But the tasks where reasoning capture matters most, production writes, credential access, multi-agent delegation chains, are exactly the tasks with the largest blast radius, and that is where the cost of capturing the trace is easiest to justify. Capturing the reasoning trace requires instrumenting the agent harness itself, not just the gateway it talks through, and that architectural choice shapes everything that follows in how these systems get built.
Tamper-evidence and hash chaining for legally defensible audit trails
An audit trail that can be edited after the fact without anyone noticing isn't an audit trail at all; it's a claim, and nothing stops a third party from rejecting that claim outright, because a mutable internal log gives them no way to confirm the record wasn't rewritten.
Tamper evidence research in this space builds that confirmation through hash chaining and Merkle batching. Each event's hash gets chained to the hash of the event before it, so an edit, a deletion, a reordering, or a forked version of the record all break the chain in a way that's detectable. The full integrity stack built this way catches all four of those attack classes at a 100% detection rate with zero false positives. Hash chaining gives per-event ordering and tamper evidence within a single trust boundary, and Merkle batching adds something further: a compact inclusion proof, which lets a verifier confirm that one specific event belongs to the record without downloading the whole log to check.
That's sufficient for a single organization checking its own records. It isn't sufficient when a dispute crosses organizational lines, when a vendor, a customer, and a regulator each hold a different slice of the evidence and none of their infrastructure counts as neutral ground to the others. For that situation, periodic on-chain anchoring of epoch roots gives every party a public commitment to check against: anyone holding the disclosed payload and its Merkle proof can verify the record independently, without trusting the operator and without a pre-agreed intermediary standing between them. The on-chain footprint for this is small by design. Each anchor stores only a 32-byte epoch root and a back-pointer to the previous one; no actual event content goes on-chain, so data privacy holds while independent verification still works. A comparison across five anchoring mechanisms, signed digests, transparency logs, timestamping authorities, internal hash chains, and on-chain anchoring, found that on-chain anchoring is the only one of the five that gives a neutral third party a way to verify the record without requiring trust in the operator or in some intermediary chosen in advance. The overhead for all of this is modest: the full integrity stack adds roughly 48 microseconds of median latency per event, plus a small, fixed storage cost per event, numbers low enough that tamper evidence belongs in production agent infrastructure as a baseline property, not as an optional mode a team turns on after something goes wrong. On-chain anchoring specifically answers the cross-organizational dispute case; it isn't a requirement for every team running agents internally, where hash chaining and Merkle proofs inside a single trust boundary may be enough.
Delegation provenance and multi-agent accountability chains
Multi-agent workflows introduce an attribution problem that is not just a matter of better logging; it's the actual mechanism by which accountability disappears. A coordinator agent that invokes three specialist subagents produces one compound action, and no single agent's individual log explains that action on its own.
Delegation provenance, the record of which parent agent or human trigger authorized a given child agent's action, has to exist as its own structured field. It cannot be inferred later from timestamps or from which log entries happen to sit near each other in time, because that kind of inference breaks down exactly when it matters most: during an incident, when the sequence of calls was unusual to begin with. Without delegation provenance captured directly, there's no reliable way to determine after the fact whether a destructive action stayed within its delegated scope or exceeded it. The 2026 Singapore Consensus on Global AI Safety Research Priorities addresses agentic risk through a companion report setting out ten foundational principles, among them auditability, traceable identity, and human oversight, but that report does not single out multi-agent accountability diffusion as the single leading under-regulated risk in agentic systems. That leaves delegation provenance as largely an engineering responsibility rather than a settled regulatory requirement. This is why the eight-field schema treats it as a first-class field: the audit trail has to reconstruct not only what each agent did, but who authorized each agent to act, and whether that authorization stayed inside the scope the original human trigger granted.
The governance implication follows directly. An escalation rule that fires when a tool call exceeds its allowed blast radius needs a counterpart that fires when a delegation chain exceeds the scope of the original approval, and both of those events need to appear in the audit trail as linked records rather than as two unrelated entries that happen to share a timestamp. Prompt injection that spreads across agents shows why the linkage matters. An attacker-controlled instruction buried in one agent's tool output can cause a different, downstream agent to take an action nobody approved, and the only way to catch that in the record is if context provenance and delegation provenance are both captured, because the point where the injection entered and the point where the bad action executed sit in two different agents' logs.
Policy versioning and agent-as-code configuration as audit infrastructure
An audit trail that logs every tool call faithfully but can't say which version of the governing policy was active at the time, who approved that version, or when it went live is missing a piece that matters for any incident tied to a configuration change.
Code provenance, the agent version, model version, and prompt version active at the moment of each event, earns its place as a first-class field for a specific reason: a behavior that was perfectly compliant under one policy version can be a clear violation under the next one, and the audit trail needs to be able to tell those two cases apart. That requires treating policy the way software teams already treat code. Every prompt change, every guardrail update, every policy revision should carry a diff, a named author, an approval signature, and a deploy timestamp, with Git serving as the source of truth and the agent platform pulling versioned configs at deploy time. This is the agents-as-code pattern, applied to governance rather than just to the agent's own execution logic.
The acceptable use policy and the system prompt are not the same artifact, and the distinction matters. The system prompt is what the model itself sees. The acceptable use policy is what the surrounding runtime actually enforces, and it has to be version-controlled in its own right, with a direct link to the audit trail, so that any enforcement decision can be traced back to the exact policy version that produced it. A structured agent definition file that spells out budget caps, tool allow-lists, and escalation hooks works the same way. That file is an audit artifact in its own right: once it's checked into a repository, reviewed, and deployed through a pipeline, the deploy event itself becomes part of the audit trail, and the configuration at that specific commit is the policy every session running against it has to be judged by.
Credential provenance and sandbox isolation as audit-trail prerequisites
An audit trail is only as good as the identity attached to each event, and if agents run on shared or hardcoded credentials, the question of who performed a given action has no reliable answer, no matter how well everything else about the record is built.
Shadow agents make this failure mode concrete. These are agents deployed outside a security team's governance, connecting to production APIs using hardcoded credentials or developer tokens meant for a person, not a service. Because the credential is shared across multiple actors and sessions, any action those agents take is invisible to the security team watching for it and unattributable in whatever audit trail does exist, since the credential itself can't distinguish one session or one actor from another. The fix runs the other direction: credentials minted fresh at the start of each session, scoped tightly to what that session actually needs, and tied to a specific agent identity rather than shared across a fleet of them. That is the condition the rest of this schema depends on. Intent, policy evaluation, execution, delegation, all eight fields rely on the assumption that the identity attached to an event is real and specific, and that assumption only holds when credential issuance and execution isolation are built into the infrastructure from the start, not bolted on after an incident shows where the gaps were.


