rec · every tool call and its answer

Test every agent change against a real run.

AgentTrace records a real run, every tool call and every answer. Replay the next prompt, model or code change against that recording without calling a single real tool, and get a PASS or FAIL you can gate a release on.

replayingsupport-agent · v2.0.0
  1. 01get_customer("c-42")exact
  2. 03get_order("A-1")exact
  3. 04get_order("B-2")exact
  4. 07get_delivery_status("A-1")exact
  5. 09get_delivery_status("B-2")not called
  6. 11format_reply("Dana")exact
fail · missing_tool_callget_delivery_status("B-2") recorded at seq 9, never called
every tool call answered from the recording · real tool calls: 0
000
real tool calls during replay
001
request to store a whole run
011
finding codes, each with a severity
000
runtime dependencies in the SDK

regressions

Agents break quietly. AgentTrace shows you where.

Every difference between a recording and its replay becomes a finding with a stable code and a severity. These are the defaults, and a ComparisonPolicy can raise, lower or ignore any of them.

skipped stepreplay_demo.py · v2
recorded

07 get_delivery_status('A-1')

09 get_delivery_status('B-2')

11 format_reply('Dana', …)

replayed

get_delivery_status('A-1')match

not calledmissing

format_reply('Dana', …)match

FAILMISSING_TOOL_CALL · error

get_delivery_status(order_id="B-2") was recorded (seq 9) but never called. The reply was identical, so an output diff would have passed it.

how it works

Record once. Replay every change.

  1. 01

    Record

    Wrap the agent and decorate its tools. The SDK buffers the run in memory, snapshotting each value as it is recorded.

     3 get_order("A-1")
     4 get_order("B-2") parallel
     5 answer A-1 50 ms
     6 answer B-2 50 ms
  2. 02

    Save

    The finished run is stored in one transaction, keyed by its own id. A retried upload conflicts instead of duplicating.

    POST /runs/ingest
    201 14 events stored
    POST /runs/ingest retry
    409 already stored
  3. 03

    Replay

    Your unchanged entry point runs again. Every decorated tool answers from the recording and never executes.

    get_order("A-1")  exact
    get_order(" B-2") normalized
    get_order("a-1")  unmatched
    real tool calls: 0
  4. 04

    Compare

    Skipped calls, unexpected calls, a changed status or a changed output shape each fail the run.

    FAIL 1 error, 0 warnings
    error MISSING_TOOL_CALL
    get_delivery_status(
      "B-2") never called

replay demo

Four changes. One recording. Four verdicts.

A support agent was recorded once. Four changed versions were then replayed against that recording. Pick one to see how each tool call matched and what the comparison decided.

change None. This is the exact code that was recorded.

6 exact · 0 unmatched · 0 unused · output same as recording
toolrecordednew agentoutcome
get_customercustomer_id='c-42'customer_id='c-42'exact
get_orderorder_id='A-1'order_id='A-1'exact
get_orderorder_id='B-2'order_id='B-2'exact
get_delivery_statusorder_id='A-1'order_id='A-1'exact
get_delivery_statusorder_id='B-2'order_id='B-2'exact
format_replycustomer_name='Dana', lines=[…]customer_name='Dana', lines=[…]exact

PASS 0 errors, 0 warnings

No findings. Calls, order, status and output all matched.

why it matters Same code, same answers. The baseline every other version is judged against.

dashboard

Read a run like a timeline, not a log.

A read-only view of everything the SDK uploads: projects, runs with their verdicts, each run's timeline and its comparison report.

localhost:3000/runs/aedff5c5…

Projects / Demo

support-agent

completedv1.0.0

14 events as 8 steps · 2 steps ran in parallel (overlapping bars)

SeqStepDur.Span
0agent_start
1-2get_customer50 ms
3-5get_order50 ms
4-6get_order50 ms
7-8get_delivery_status51 ms
9-10get_delivery_status51 ms
11-12format_reply0 ms
13agent_end
The run page for a real recording of examples/async_support_agent.py
Paired
A call and its answer share a call_id, so each step is one row, even when parallel calls interleave.
Parallel
Steps are drawn across the run's sequence range. Calls in flight together show as overlapping bars.
Verdicts
Every replay links to its recording and to its comparison report, grouped by severity.

python sdk

One decorator per tool. Nothing new in your dependency tree.

The SDK runs inside your agent's process, so it uses the standard library only. Recording never raises into your code: tool results and exceptions pass through unchanged, and a failed upload is logged, never thrown.

$ uv add agenttrace-vcr

or pip install agenttrace-vcr · on PyPI

Python 3.12+. The package is agenttrace-vcr; you import it as agenttrace, and its CLI is agenttrace.

agent.py
from agenttrace import AgentTracer

tracer = AgentTracer()  # reads AGENTTRACE_*

@tracer.tool
async def get_order(order_id: str) -> dict:
    return {"id": order_id, "status": "shipped"}

async def run_agent(input):
    async with tracer.trace(
        "support-agent", input=input, agent_version="v1.0.0",
    ) as trace:
        order = await get_order("A-1")
        trace.set_output({"message": f"Your order is {order['status']}."})
test_agent.py
from agenttrace import ComparisonPolicy, Recording

recording = await Recording.from_api(run_id)
result, report = await tracer.replay_and_compare(
    recording,
    run_agent,
    agent_version="v2.0.0",
    policy=ComparisonPolicy(ignore_paths=["generated_at"]),
)
print(report.format())
assert report.passed

# FAIL  1 error, 0 warnings, 0 info
#   error  MISSING_TOOL_CALL  get_delivery_status(
#          order_id="B-2") was recorded (seq 9) but never called

suites and ci

Recordings live in your repo. The exit code is the gate.

Export a stored run to JSON, review it for secrets, and commit it next to a suite.toml. run-suite replays every case offline, with no API, no network and no real tools. This repo's own GitHub Actions workflow runs run-suite on every pull request, so a regression turns the PR red.

examples/suites/support/suite.toml
name = "support"
agent = "examples.async_support_agent:run_agent"

[[cases]]
name = "two-orders"
recording = "recordings/two-orders.json"
terminal
$ agenttrace run-suite examples/suites/support/suite.toml
PASS  two-orders
suite support: 1 passed, 0 failed

$ agenttrace export $RUN_ID -o recordings/refund.json
# prints the [[cases]] entry to paste, and refuses to overwrite without --force

exit codes

  • 0Every case passed.
  • 1At least one case failed. Each failure's errors are listed under it.
  • 2The suite itself could not run.

Any CI runner that fails a job on a non-zero exit can use it as it is.

proof

A real pull request, turned red.

One changed line, and CI named the exact tool call the agent stopped making.

github.com/destroxx/agenttrace/pull/1

Demo: deliberate regression — do not merge #1

closed, never mergeddemo/regression1 file changed, +1 −1

examples/async_support_agent.py
- orders = await asyncio.gather(*(get_order(order_id) for order_id in order_ids))+ orders = await asyncio.gather(*(get_order(order_id) for order_id in order_ids[:1]))
The agent now looks up only the first order it was asked about.
checks
  • Regression suitefailing
  • SDK (Python 3.12)failing
  • SDK (Python 3.13)failing
  • APIpassing
  • Webpassing
Regression suite · run-suite · FAIL two-orders, 7 errors, includingMISSING_TOOL_CALL get_order(order_id="B-2") was recorded (seq 4) but never called

See the pull request on GitHub

quickstart

Your first recording, replayed and judged.

Only need the SDK in your agent? uv add agenttrace-vcr is all — recording and regression suites work with no server. To run the whole stack (API, dashboard, Postgres) you need Python 3.12+, Node 20+ and Docker. The demo agent is scripted, so there is no LLM to pay for.

  1. 01Clone, then start Postgres

    $ git clone https://github.com/destroxx/agenttrace && cd agenttrace
    $ cp .env.example .env
    # set POSTGRES_PASSWORD in .env, e.g. openssl rand -hex 16
    $ docker compose up -d
    # wait until docker compose ps shows (healthy)
    
  2. 02Configure, migrate and run the API

    $ cp apps/api/.env.example apps/api/.env
    # set the same POSTGRES_PASSWORD in apps/api/.env
    $ python3.12 -m venv .venv && source .venv/bin/activate
    $ pip install -e "apps/api[dev]" -e "packages/python-sdk[dev]"
    $ cd apps/api && alembic upgrade head
    $ python -m scripts.new_admin_key   # put its ADMIN_KEY_SHA256 line in .env
    $ uvicorn app.main:app --reload
    
  3. 03Create a project and its key, then record and replay

    # in a second shell, from the repo root
    $ source .venv/bin/activate
    $ ADMIN="Authorization: Bearer <the admin key>"
    $ export AGENTTRACE_PROJECT_ID=$(curl -s -X POST localhost:8000/api/v1/projects \
      -H "$ADMIN" -H 'content-type: application/json' -d '{"name":"Demo"}' \
      | python3 -c 'import sys,json; print(json.load(sys.stdin)["id"])')
    $ export AGENTTRACE_API_KEY=$(curl -s -X POST \
      localhost:8000/api/v1/projects/$AGENTTRACE_PROJECT_ID/keys \
      -H "$ADMIN" -H 'content-type: application/json' -d '{}' \
      | python3 -c 'import sys,json; print(json.load(sys.stdin)["key"])')
    $ python examples/async_support_agent.py
    $ python examples/replay_demo.py
    
  4. 04Open the dashboard

    $ cd apps/web && cp .env.example .env.local
    $ npm install && npm run dev
    # then open http://localhost:3000/projects
    

Full setup in the README, including troubleshooting and a DATABASE_URL alternative.

faq

Questions teams ask first.

Does replay ever call my real tools?

No. Every @tracer.tool is answered from the recording and its body never runs. A call that no recording matches raises UnmatchedToolCall inside the agent, so it cannot reach the real tool either. The demo counts real executions, and the count stays at 0 across all four replays.

How do you handle LLM non-determinism?

The recording fixes what the tools answer, so the only thing left to vary is the agent, which is the thing you are testing. Arguments are matched exactly first, then after a conservative normalisation: strings are trimmed, 2.0 equals 2, and keys set to None are dropped. Case is never folded. Reworded output text is a warning by default, not a failure; turn on semantic comparison and Claude judges whether its meaning changed.

Can recording slow down or break my agent?

Recording never raises into your code. Tool results and exceptions pass through unchanged, the run is buffered in memory and uploaded once when it ends, and a failed upload is logged and swallowed. Stop the API and run your agent: it finishes normally and reports uploaded: False. Without AGENTTRACE_PROJECT_ID the SDK never opens a socket.

What does it take to add to an existing agent?

Decorate the functions that are your tools with @tracer.tool and wrap the entry point in tracer.trace(...). Tools you cannot decorate go through tracer.record_tool_call. It works with any Python 3.12+ agent, sync or async, and the SDK has no runtime dependencies.

How are parallel tool calls kept apart?

Each call and its answer share a call_id. Sequence numbers give the order, but not which answer belongs to which call, so the explicit key is what lets two overlapping get_order calls be paired back up. The dashboard draws them as overlapping bars.

Where does the data live?

In your own PostgreSQL, behind the FastAPI service you run. A finished run is stored in one transaction, so a half-stored trace cannot exist, and a retried upload conflicts instead of creating a duplicate. Regression suites live in your repo as JSON recordings, and run-suite needs neither the API nor the network.

What is not built yet?

Deployment with private data: every write needs an API key that can only touch its own project, but reads are public, so anyone who can reach the API can see its runs. And evaluation: AgentTrace checks that a new agent behaves like its recording, not whether either answer is good.

untested

Every prompt change is a deploy nobody tested.

Record today's runs, and tomorrow's agent has something to be measured against. It takes a decorator on each tool and one context manager around the run.