On September 24, researchers showed that AI coding agents in Claude Code, Codex and Grok Build would delete their own execution traces when asked, and nothing stopped them. A log the agent can erase is weak evidence of what the agent did.
The paper, LLM Agents Can Easily Tamper With Their Own Traces by Jeremy Qin, Maksym Andriushchenko and colleagues, recommends that trace logging run through an independent mechanism outside the agent's control.
That is the problem Tracekit was built for. It also raises a question we hear often: how does Tracekit compare with the tools teams already use to monitor agents? This piece sets out the main kinds of tool in use today, what each is good at, and where Tracekit fits. It also says plainly where other tools are the better choice.
The short version
- Most agent monitoring tools are built for debugging, cost and evaluation. They do that well, with far more mature integrations and interfaces than Tracekit.
- Guardrail tools decide whether a tool call may run. They lead on enforcing one policy across several agents.
- Shared audit logs collect activity for compliance reporting. They are only as strong as the place each event came from.
- A small group of tools, Tracekit among them, produce evidence: a record of what the agent did that the agent cannot quietly rewrite, and that a third party can verify.
- These work as layers, not substitutes. Most teams that need evidence also need an observability tool.

Observability platforms
Langfuse, LangSmith, Arize Phoenix, AgentOps, Datadog LLM Observability
These capture traces, tool spans, token counts, latency and evaluations, usually through OpenTelemetry or framework SDKs. Langfuse and Phoenix are open source and can be self-hosted. LangSmith comes from the LangChain team and goes deepest on LangChain and LangGraph. AgentOps focuses on replaying agent sessions. Datadog adds LLM traces to an existing monitoring stack.
Use them to find out why a run failed, where the money went and whether quality is slipping. They are more mature for that work than Tracekit. Langfuse, for example, documents tracing for Claude Code and other coding agents. Our piece When Nothing Errors covers what observability measures and where it stops.
They are weaker as audit evidence, and they are not designed to be. Records sit in an ordinary database, and capture runs through hooks or SDKs in the same environment being watched, where it can be switched off. Langfuse's Claude Code integration works by reading the session transcript after each response. A record built that way is only as trustworthy as that transcript.
Policy and guardrail tools
Arcjet, Preloop, Noma, Runlayer, native Claude Code hooks
These act before a tool call runs. They can allow it, deny it or hold it for a person to approve.
- Arcjet enforces one policy, written in Rego, across Claude Code, GitHub Copilot, Cursor and Codex, through the hooks each agent already fires.
- Preloop is an open-source, self-hostable control plane with an MCP firewall, approval rules and support for most major coding agents.
- Noma and Runlayer add agent discovery, runtime detection of prompt injection and data exfiltration, and access control.
- Claude Code's own hooks can block a call before it runs, with rules you write yourself.
These tools are better than Tracekit at enforcing rules across many agents and at ready-made threat detection. Their record is kept for operations and alerting. Arcjet, for example, records sessions and decisions in its console and exports them to tools like Datadog, Splunk and S3. That is useful, but the record is only as trustworthy as whoever runs the console.
Shared audit logs
WorkOS audit harness, Claude Compliance API
The WorkOS audit harness sends events from Claude Code, Codex, OpenClaw and pi into one audit log. Its documentation is candid about the limit: events sent from a developer's machine "can be fabricated or withheld." Events composed on the server side are stronger.
The Claude Compliance API records Claude activity on Anthropic's side and feeds it into existing security tools. On Claude Enterprise this includes Claude Code sessions.
Both are good choices for compliance reporting across a team. Neither gives you a record that a third party can verify offline.
Tamper-evident records
Agent Receipts, AgentLedger, Tracekit
This group exists to produce evidence.
- Agent Receipts is an open specification with working tools. A separate daemon holds the signing keys, each receipt is signed with Ed25519 and hash-chained, and an MCP proxy can sign every tool call without code changes.
- AgentLedger writes signed receipts and, when policy denies an action, signs and stores the denial before the error is raised. Its design is published as an IETF Internet-Draft.
- Tracekit belongs here (write-up, code). What sets it apart:
Separate signer. On Linux in system mode, the signing service runs as its own operating-system user. The agent never holds the key.
Signed hash chain. Each record carries the previous record's hash, a sequence number and an Ed25519 signature. Editing, deleting or reordering a record breaks verification at the first bad record.
External witness. The head of the chain is checkpointed to a git repository or file off the machine. That catches the case a chain alone cannot: someone rebuilding the whole log.
Transcript hashing. The agent's session transcript is hashed at every step, so deleting or editing earlier parts of the conversation is reported.
Offline bundles. A run exports as one file that anyone with the public key can verify, without trusting us or you.
Tracekit also has a policy gate that blocks known-dangerous calls and holds risky ones for approval by a different user. Those rules are tripwires for honest mistakes and obvious cases, not a replacement for the guardrail tools above.

What Tracekit does not do
- Claude Code is the most complete path. Custom agents work through the Python and TypeScript SDKs, LangChain, LangGraph and MCP through adapters, and other systems through OpenTelemetry ingest, where calls are recorded after the fact rather than gated. Hooks for Codex CLI, Cursor and Gemini CLI are on the main branch: the installer sets them up in dev mode only, they have been tested with each tool's documented hook format but not yet inside the live tools, and reasoning capture is Claude Code only.
- It records only what passes through its integrations. It does not discover agents on its own.
- Its policy rules can be evaded by a determined agent. Hard enforcement belongs in sandboxes and operating-system controls, with Tracekit as the independent witness.
- It is early. Tracekit is a v0.2 release candidate and not yet on PyPI. Full isolation runs on Linux. macOS system mode is experimental and has only been tested with mocks. Windows runs dev mode only, which checks integrity but not isolation, and has not yet been tested on real hardware.
Using them together
Tracekit is a specialized audit tool, not a replacement for Langfuse-style monitoring. Teams that need both run both: an observability backend for day-to-day operations, and a signed, witnessed ledger for the runs whose record has to hold up even if the agent itself was compromised or uncooperative. Tracekit can also receive OpenTelemetry traces and push signed runs to OTLP backends such as Jaeger, so the two layers can share one instrumentation path.

Tracekit is open source under MIT at Cygnux-Labs/Tracekit. For a practical pilot plan, and how it works alongside Causeway, see our business guide to Tracekit and Causeway. If you want help piloting it, book a call.
