Test every agent change against a real run.
AgentTrace records a real run, every tool call and every answer. Replay the next prompt, model or code change against that recording without calling a single real tool, and get a PASS or FAIL you can gate a release on.
- 01get_customer("c-42")exact
- 03get_order("A-1")exact
- 04get_order("B-2")exact
- 07get_delivery_status("A-1")exact
- 09get_delivery_status("B-2")not called
- 11format_reply("Dana")exact
- 000
- real tool calls during replay
- 001
- request to store a whole run
- 011
- finding codes, each with a severity
- 000
- runtime dependencies in the SDK
regressions
Agents break quietly. AgentTrace shows you where.
Every difference between a recording and its replay becomes a finding with a stable code and a severity. These are the defaults, and a ComparisonPolicy can raise, lower or ignore any of them.
07 get_delivery_status('A-1')
09 get_delivery_status('B-2')
11 format_reply('Dana', …)
get_delivery_status('A-1')match
not calledmissing
format_reply('Dana', …)match
get_delivery_status(order_id="B-2") was recorded (seq 9) but never called. The reply was identical, so an output diff would have passed it.
how it works
Record once. Replay every change.
- 01
Record
Wrap the agent and decorate its tools. The SDK buffers the run in memory, snapshotting each value as it is recorded.
3 get_order("A-1") 4 get_order("B-2") parallel 5 answer A-1 50 ms 6 answer B-2 50 ms
- 02
Save
The finished run is stored in one transaction, keyed by its own id. A retried upload conflicts instead of duplicating.
POST /runs/ingest 201 14 events stored POST /runs/ingest retry 409 already stored
- 03
Replay
Your unchanged entry point runs again. Every decorated tool answers from the recording and never executes.
get_order("A-1") exact get_order(" B-2") normalized get_order("a-1") unmatched real tool calls: 0
- 04
Compare
Skipped calls, unexpected calls, a changed status or a changed output shape each fail the run.
FAIL 1 error, 0 warnings error MISSING_TOOL_CALL get_delivery_status( "B-2") never called
replay demo
Four changes. One recording. Four verdicts.
A support agent was recorded once. Four changed versions were then replayed against that recording. Pick one to see how each tool call matched and what the comparison decided.
real tool executions during all 4 replays: 0
change None. This is the exact code that was recorded.
6 exact · 0 unmatched · 0 unused · output same as recording| tool | recorded | new agent | outcome |
|---|---|---|---|
| get_customer | customer_id='c-42' | customer_id='c-42' | exact |
| get_order | order_id='A-1' | order_id='A-1' | exact |
| get_order | order_id='B-2' | order_id='B-2' | exact |
| get_delivery_status | order_id='A-1' | order_id='A-1' | exact |
| get_delivery_status | order_id='B-2' | order_id='B-2' | exact |
| format_reply | customer_name='Dana', lines=[…] | customer_name='Dana', lines=[…] | exact |
PASS 0 errors, 0 warnings
No findings. Calls, order, status and output all matched.
why it matters Same code, same answers. The baseline every other version is judged against.
dashboard
Read a run like a timeline, not a log.
A read-only view of everything the SDK uploads: projects, runs with their verdicts, each run's timeline and its comparison report.
Projects / Demo
support-agent
completedv1.0.014 events as 8 steps · 2 steps ran in parallel (overlapping bars)
- Paired
- A call and its answer share a call_id, so each step is one row, even when parallel calls interleave.
- Parallel
- Steps are drawn across the run's sequence range. Calls in flight together show as overlapping bars.
- Verdicts
- Every replay links to its recording and to its comparison report, grouped by severity.
python sdk
One decorator per tool. Nothing new in your dependency tree.
The SDK runs inside your agent's process, so it uses the standard library only. Recording never raises into your code: tool results and exceptions pass through unchanged, and a failed upload is logged, never thrown.
$ uv add agenttrace-vcror pip install agenttrace-vcr · on PyPI
Python 3.12+. The package is agenttrace-vcr; you import it as agenttrace, and its CLI is agenttrace.
from agenttrace import AgentTracer tracer = AgentTracer() # reads AGENTTRACE_* @tracer.tool async def get_order(order_id: str) -> dict: return {"id": order_id, "status": "shipped"} async def run_agent(input): async with tracer.trace( "support-agent", input=input, agent_version="v1.0.0", ) as trace: order = await get_order("A-1") trace.set_output({"message": f"Your order is {order['status']}."})
from agenttrace import ComparisonPolicy, Recording recording = await Recording.from_api(run_id) result, report = await tracer.replay_and_compare( recording, run_agent, agent_version="v2.0.0", policy=ComparisonPolicy(ignore_paths=["generated_at"]), ) print(report.format()) assert report.passed # FAIL 1 error, 0 warnings, 0 info # error MISSING_TOOL_CALL get_delivery_status( # order_id="B-2") was recorded (seq 9) but never called
suites and ci
Recordings live in your repo. The exit code is the gate.
Export a stored run to JSON, review it for secrets, and commit it next to a suite.toml. run-suite replays every case offline, with no API, no network and no real tools. This repo's own GitHub Actions workflow runs run-suite on every pull request, so a regression turns the PR red.
name = "support" agent = "examples.async_support_agent:run_agent" [[cases]] name = "two-orders" recording = "recordings/two-orders.json"
$ agenttrace run-suite examples/suites/support/suite.toml PASS two-orders suite support: 1 passed, 0 failed $ agenttrace export $RUN_ID -o recordings/refund.json # prints the [[cases]] entry to paste, and refuses to overwrite without --force
exit codes
- 0Every case passed.
- 1At least one case failed. Each failure's errors are listed under it.
- 2The suite itself could not run.
Any CI runner that fails a job on a non-zero exit can use it as it is.
proof
A real pull request, turned red.
One changed line, and CI named the exact tool call the agent stopped making.
Demo: deliberate regression — do not merge #1
closed, never mergeddemo/regression1 file changed, +1 −1
- orders = await asyncio.gather(*(get_order(order_id) for order_id in order_ids))+ orders = await asyncio.gather(*(get_order(order_id) for order_id in order_ids[:1]))The agent now looks up only the first order it was asked about.
- Regression suitefailing
- SDK (Python 3.12)failing
- SDK (Python 3.13)failing
- APIpassing
- Webpassing
quickstart
Your first recording, replayed and judged.
Only need the SDK in your agent? uv add agenttrace-vcr is all — recording and regression suites work with no server. To run the whole stack (API, dashboard, Postgres) you need Python 3.12+, Node 20+ and Docker. The demo agent is scripted, so there is no LLM to pay for.
01Clone, then start Postgres
$ git clone https://github.com/destroxx/agenttrace && cd agenttrace $ cp .env.example .env # set POSTGRES_PASSWORD in .env, e.g. openssl rand -hex 16 $ docker compose up -d # wait until docker compose ps shows (healthy)
02Configure, migrate and run the API
$ cp apps/api/.env.example apps/api/.env # set the same POSTGRES_PASSWORD in apps/api/.env $ python3.12 -m venv .venv && source .venv/bin/activate $ pip install -e "apps/api[dev]" -e "packages/python-sdk[dev]" $ cd apps/api && alembic upgrade head $ python -m scripts.new_admin_key # put its ADMIN_KEY_SHA256 line in .env $ uvicorn app.main:app --reload
03Create a project and its key, then record and replay
# in a second shell, from the repo root $ source .venv/bin/activate $ ADMIN="Authorization: Bearer <the admin key>" $ export AGENTTRACE_PROJECT_ID=$(curl -s -X POST localhost:8000/api/v1/projects \ -H "$ADMIN" -H 'content-type: application/json' -d '{"name":"Demo"}' \ | python3 -c 'import sys,json; print(json.load(sys.stdin)["id"])') $ export AGENTTRACE_API_KEY=$(curl -s -X POST \ localhost:8000/api/v1/projects/$AGENTTRACE_PROJECT_ID/keys \ -H "$ADMIN" -H 'content-type: application/json' -d '{}' \ | python3 -c 'import sys,json; print(json.load(sys.stdin)["key"])') $ python examples/async_support_agent.py $ python examples/replay_demo.py
04Open the dashboard
$ cd apps/web && cp .env.example .env.local $ npm install && npm run dev # then open http://localhost:3000/projects
Full setup in the README, including troubleshooting and a DATABASE_URL alternative.
faq
Questions teams ask first.
Does replay ever call my real tools?
No. Every @tracer.tool is answered from the recording and its body never runs. A call that no recording matches raises UnmatchedToolCall inside the agent, so it cannot reach the real tool either. The demo counts real executions, and the count stays at 0 across all four replays.
How do you handle LLM non-determinism?
The recording fixes what the tools answer, so the only thing left to vary is the agent, which is the thing you are testing. Arguments are matched exactly first, then after a conservative normalisation: strings are trimmed, 2.0 equals 2, and keys set to None are dropped. Case is never folded. Reworded output text is a warning by default, not a failure; turn on semantic comparison and Claude judges whether its meaning changed.
Can recording slow down or break my agent?
Recording never raises into your code. Tool results and exceptions pass through unchanged, the run is buffered in memory and uploaded once when it ends, and a failed upload is logged and swallowed. Stop the API and run your agent: it finishes normally and reports uploaded: False. Without AGENTTRACE_PROJECT_ID the SDK never opens a socket.
What does it take to add to an existing agent?
Decorate the functions that are your tools with @tracer.tool and wrap the entry point in tracer.trace(...). Tools you cannot decorate go through tracer.record_tool_call. It works with any Python 3.12+ agent, sync or async, and the SDK has no runtime dependencies.
How are parallel tool calls kept apart?
Each call and its answer share a call_id. Sequence numbers give the order, but not which answer belongs to which call, so the explicit key is what lets two overlapping get_order calls be paired back up. The dashboard draws them as overlapping bars.
Where does the data live?
In your own PostgreSQL, behind the FastAPI service you run. A finished run is stored in one transaction, so a half-stored trace cannot exist, and a retried upload conflicts instead of creating a duplicate. Regression suites live in your repo as JSON recordings, and run-suite needs neither the API nor the network.
What is not built yet?
Deployment with private data: every write needs an API key that can only touch its own project, but reads are public, so anyone who can reach the API can see its runs. And evaluation: AgentTrace checks that a new agent behaves like its recording, not whether either answer is good.
Every prompt change is a deploy nobody tested.
Record today's runs, and tomorrow's agent has something to be measured against. It takes a decorator on each tool and one context manager around the run.