Agent traces are becoming some of the most valuable data a company produces: a record of what an agent was asked, how it reasoned, which tools it called and what came back. Most of that record is lost, editable by the agent that wrote it, incomplete, or too sensitive to reuse.
Sam Liu recently argued that agent traces are the new oil. Labs train on them, vendors sell them, and companies deploying agents mostly file them away as logs. We agree with the argument, and want to add the part that matters most to us. Oil is only worth something if it reaches the refinery uncontaminated. A trace is only worth something if you can trust it.
The short version
- A trace is now the work itself. As agents run longer, more of the record is the agent's own actions rather than a person's messages. It is how a company knows what was done in its name.
- Traces feed three uses: running agents well, keeping them safe, and making them better. Each use breaks in a different way when the trace is wrong.
- The supply goes bad in four ways. Traces are deleted on a timer, edited by the agent that wrote them, missing what happened outside the capture path, or full of secrets that make them unusable.
- Knowing which step mattered is the hardest part. A long trace shows everything an agent saw. It does not show which input caused the outcome.
- Provenance comes before refining. Before a trace is used for an incident review, an audit or training, it should be complete, unaltered and redacted, and say so.
What a trace is, and why it matters now
A chat transcript was a back-and-forth between a person and a model. An agent trace adds everything in between: the model's reasoning, each tool call and its arguments, the results that came back, and what the agent did next. In a long-running agent, most of the trace is the agent working on its own.
Three trends have made these records valuable. Training increasingly runs on rollouts, which are traces. Agents run longer and take more actions per task, so each trace holds more. And agents are being deployed across whole organizations, so the traces describe real operations, not demos.
For a company deploying agents, traces serve three purposes:
- Operations. Where tokens and time go, whether a workflow pays for itself, and why a run failed.
- Safety. What an agent actually did when it had access to code, money or customer data, and whether anything it read steered it.
- Improvement. Successful runs hold procedures worth reusing. Failed runs hold corrections and human feedback.
Each of these assumes the trace is a faithful record. That assumption deserves a closer look.

Four ways the supply goes bad
Lost. Claude Code stores each session as a JSONL transcript under ~/.claude/projects/ and, by its documentation, deletes transcripts after 30 days unless cleanupPeriodDays is changed. The same page notes that the entry format is internal and changes between versions. A company that has not decided where its traces live has, in effect, decided to lose them.
Altered. That transcript is an ordinary file in the user's home directory, writable by the same user the agent runs as. In September, Jeremy Qin, Maksym Andriushchenko and colleagues showed that agents in Claude Code, Codex and Grok Build deleted their own execution traces when asked, without any guardrail stopping them. A trace the agent can rewrite cannot settle a question about what the agent did.
Incomplete. Every capture method sees only what passes through it. A hook sees tool calls. An SDK wrapper sees model calls made through that SDK. A script that opens its own network connection is invisible to both, and a hook that is switched off records nothing. A trace that does not say where its gaps are looks complete when it is not.
Leaky. Traces carry prompts, file contents, command output, API keys and customer data. That makes them hard to share with an auditor, risky to send to a vendor, and unsafe to train on without cleaning. Many teams respond by keeping nothing, which brings them back to the first problem.
Why provenance comes before refining
Each use of a trace fails differently when the trace is wrong.
An incident review built on an incomplete trace ends with the wrong fix. A safety claim built on an editable trace persuades nobody who has a reason to doubt it. And training on traces inherits whatever is in them. A run where the agent followed an instruction hidden in a web page looks, in the trace, like a successful run. Reuse it as an example of good work and the injected behavior is what gets reinforced.
So before a trace is refined into a report, an evaluation set or training data, it should be able to answer four questions. Was anything removed or changed after it was written? What was outside the capture path? Has sensitive content been removed? Which inputs actually drove the outcome?

How we approach it
Our two open-source tools each take part of this.
Tracekit is the evidence layer. It records tool calls, model calls and the agent's transcript into a ledger signed by a separate operating-system user on Linux, so the agent never holds the signing key. Each record is hash-chained to the one before, and the chain's head is checkpointed to a git repository or file off the machine. The transcript is hashed at every step, so deleting earlier parts of a session is reported. Secrets are redacted before anything is written, coverage gaps are recorded rather than assumed away, and a run exports as a bundle anyone can verify offline. See our Tracekit write-up.
There is a real trade-off here. By default Tracekit keeps content as hashes, not text: good for evidence and privacy, not enough to learn from. Setting content_capture: full keeps the redacted text, which is what reuse needs. Teams should choose deliberately, per workflow.
Causeway is the analysis layer, built for the problem Sam Liu calls long-horizon credit assignment. When an agent read ten things before acting, all ten are in the trace. Causeway replays the run with each suspect input removed and measures whether the outcome still happens, so a cause is tested rather than guessed. See Reached Is Not Caused.

What these tools do not do yet
- Tracekit is a v0.2 release candidate. Full isolation runs on Linux; macOS is experimental, and Windows runs dev mode only. Claude Code is its most complete integration.
- Causeway is a v0.3 alpha. Its replay tests have been run on a simulated model, not yet on real ones. Each test costs model calls, roughly two per trial for each downstream model call.
- Neither one decides what a good trace is. Labelling success, failure and quality still takes judgment.
- Neither one sees below its capture path. Anything an agent does outside the hooks, SDKs or proxy it is recorded through is reported as a gap, not recovered.
Where to start
- Decide where traces live and for how long. Set retention on purpose, for example with
cleanupPeriodDaysor aSessionEndhook that archives each transcript. - Store them where the agent cannot write. A separate user, a separate machine, or a signed ledger with an external witness.
- Redact before you store. Secrets removed at write time never reach a backup, a vendor or a training set.
- Record what was not captured. A trace that lists its gaps is more useful than one that hides them.
- Label provenance before reuse. Keep track of which traces were verified, which were redacted, and which ran with untrusted inputs.
- Test causes before acting on them. Before changing a prompt, a policy or a training set because of one input, check that the input actually mattered.
The takeaway
Traces are becoming the record of how a company works when agents do the work. Like any valuable resource, their worth depends on the supply chain behind them. Keep them, keep them intact, keep them clean, and know which parts actually mattered.
Tracekit and Causeway are open source: Cygnux-Labs/Tracekit and Cygnux-Labs/Causeway. For how they fit next to observability and guardrail tools, see Monitoring Is Not Evidence. If you want help putting a trace pipeline together, book a call.
