Ground Stationbeta
Get early access

Your agents are production systems. Now they have traces.

Ground Station is open-source observability for AI agents: it captures every model call, tool call, and token of an agent’s run, reconstructs it as a trajectory, and tells you why it was slow, what it cost, and what to change.

Replay of an agent trajectory for the task “Fix the flaky checkout tests”: 4 model calls, 7 tool calls, 125,548 tokens, $0.24, 1 minute 49 seconds. Two runs of cargo test account for 94 of the 109 seconds; the first run failed with exit code 101.
trajectories/tr_01JB7QK4XEFix the flaky checkout tests
claude-codeacme/storefront running
Built for
  • Claude Code
  • Codex
  • OpenCode
  • Pi
  • Hermes
  • OpenClaw
  • browser agents
  • custom frameworks
01 agent.running t+00:00 → t+18:42

Every dashboard is green. Your agent has been looping for eighteen minutes.

Infrastructure monitoring was built for requests, services and machines. Agents execute trajectories: long chains of decisions, model calls and tool calls. The unit of work changed. The telemetry didn’t.

What your infrastructure sees healthy
HTTP 2xx
100.0%
CPU
38%
Memory
1.3 GB
p50 latency
214 ms
14:19:02.114POST/v1/messages2007.81s
14:19:10.402POST/v1/messages2006.92s
14:19:17.955POST/v1/messages2008.04s
14:19:26.310POST/v1/messages2007.47s
14:19:34.128POST/v1/messages2007.66s

All systems operational

What the agent actually did potential loop

Refactor payments moduleclaude-code · acme/payments · claude-sonnet-5-5

duration
18m 42s
tokens
3.12M
tool calls
114
est. cost
$2.81
tool calls over time0m → 18m 42s
context window184,112 / 200,000 92%

Infrastructure monitoring sees the machine. Ground Station sees the execution.

02 trajectory the primitive

The trajectory is the trace of autonomous software.

One complete agent execution, from prompt to completion, captured as a single structured object. Every turn, model call, tool call, observation and subagent, in order, with time, tokens and cost attached.

distributed tracingrequesttracespans
agent observabilityprompttrajectory turns model calls tool calls observations subagents
tr_01JB9M2R7QRefactor authentication to use OAuth
claude-code1m 50s210k tok$0.33 ok
03 data → information → knowledge → action

From events to answers.

Capture what happened. Understand why. Know what to change. The same four steps are the Ground Station roadmap, and each one only works because the one before it is exact.

  1. V0 · Data beta

    What happened?

    Every event, captured at the source: model calls, tool calls, file reads and writes, shell commands, token counts, context size, exit codes, cost. Timestamped, structured, complete.

    • agent.started
    • model.completed
    • tool.started
    • file.read
    • shell.executed
    • agent.failed
  2. V1 · Information beta

    What is happening?

    Hundreds of disconnected events become one trajectory. Not a log file: a timeline with durations, tokens and cost that you can read in seconds.

  3. V2 · Knowledge next

    Why is it happening?

    Patterns surface across the trajectory: repeated commands, duplicate reads, runaway context. Ground Station explains where the time and the tokens actually went.

    “This run was slow because the agent re-ran the full test suite 18 times.”
  4. V3 · Action later

    What should change?

    Findings become specific recommendations with an estimated impact, before you change a single prompt, tool or model.

04 fleet.live 1 → 1,000 agents

One agent or one thousand. See what every one is doing right now.

Ground Station is not just a trace viewer for runs that already finished. It is a live operations view across every agent in your organization, so you notice the one that has been stuck for eighteen minutes before it finishes burning the budget.

trajectories · 24h
1,284
success
92.4%
median runtime
3m 41s
p95 runtime
14m 22s
tokens
84.2M
est. cost
$742
tool errors
1,421
select a row to open its trajectory
agent · taskmodeldurationtoolstokenscontextstatus
sonnet-5.503:2114842k38%running
gpt-5-codex00:417231k21%running
sonnet-5.518:431143.12M92%slow · 3.0× typical
sonnet-5.504:1122670k44%running
opus-5.507:52381.20M57%awaiting approval
haiku-4.500:12318k9%running
gpt-5-codex01:5811190k26%running
sonnet-5.502:379104k17%running
05 detections knowledge, applied

Find the loop before it burns another million tokens.

Ground Station watches trajectories for the behavior that costs you: slow runs, tool loops, cost regressions, runaway context. Every finding comes with its evidence and a cause, not just an alert.

findings4 open
highslow trajectoryFix the flaky checkout tests · claude-code

24m 06s against a typical 6m 18s

evidence

p50 · 6m 18sthis run · 24m 06s
0m5m10m15m20m25m

where the 24m 06s went

cargo test 61%model 22%
other tools 4%harness overhead 13%

explanation

The trajectory was slow because the agent re-ran the full workspace test suite after every edit.

primary contributor

cargo test --workspace18 executions · 14m 42s cumulative · 61% of runtime

06 compare v41 ↔ v42

Know when an agent got worse.

Success rate barely moved. Tokens, runtime and cost did. Compare any two versions, models or prompts on the behavior that actually changed, not just on whether the task passed.

Diff behavior across

  • model upgrades
  • prompt changes
  • tool changes
  • agent releases
  • repository changes
  • infrastructure changes
v41→v42acme-coder
n = 1,904 / 1,877 trajectories

changed in v42 system prompt +212 tokens · new tool repo_search · model unchanged

Agent v41 versus v42, per-task metrics
metricv41v42change
tokens / task46.2k83.4k+80%
runtime3m 41s5m 25s+47%
tool calls / task2744+63%
repo searches / task4.111.8+188%
est. cost / task$0.41$0.74+80%
success rate92.1%92.4%+0.3pp

finding v42 performs 2.9× more repository searches before its first edit. Outcomes are unchanged; cost per task is up 80%.

07 relay open-source collector

One daemon. Every agent.

The relay is a small Rust collector that runs on the same machine as your agents. Adapters stream events into it; it normalizes, redacts, buffers and ships them. Your agents never block on telemetry.

  • Claude Codehooks adapter
  • Codexadapter
  • OpenCodeplugin
  • Custom agentsSDK · TS · Python · Rust
  • Any GenAI appOTLP in
relayrust · local up 14d 3h
  1. 1ingestunix socket · OTLP/gRPC
  2. 2normalizeground schema v1
  3. 3redactsecrets · paths · env · regex
  4. 4sampleper agent · per tool
  5. 5bufferdisk-backed · survives restarts
  6. 6batchzstd · 512 events
  7. 7exportretry · backoff · auth
events/s
1,284
redacted
1,904
dropped
0
p99 overhead
0.4ms
  • Ground Stationself-hosted or cloud
  • OTLP exporterany OpenTelemetry backend
  • Local storeSQLite · local-only mode

Open source, down to the schema.

Telemetry from autonomous software is too sensitive for a black box. The collector, adapters, SDKs and event schema are open, so you can read exactly what is captured and extend it for your own agents.

relay
The collector. A single Rust binary that runs next to your agents.
adapters
Claude Code, Codex and OpenCode, with more on the way.
sdk
Instrument any agent loop in TypeScript, Python or Rust.
schema
A versioned, documented event schema aligned with OpenTelemetry GenAI conventions.
quickstart~/src/storefront
$curl -fsSL https://groundstation.sh/install.sh | sh✓ installed ground + relay → ~/.ground/bin$ground login✓ device registered · org acme$ground connect claude-code✓ hooks written → ~/.claude/settings.json✓ relay running · unix:///tmp/ground-relay.sock$claude "fix the flaky checkout tests"→ trajectory tr_01JBA3KX2M → relay  2 model calls · 5 tool calls · 43,140 tokens · live
08 redaction before the network

Redaction happens before anything leaves the machine.

Agent telemetry contains source code, credentials, customer data and proprietary prompts. The relay applies your rules locally, so privacy is a property of the architecture, not a checkbox in a dashboard. Toggle the rules to see what would be uploaded.

your machine

raw eventfrom claude-code adapter

{
  "event": "tool.completed",
  "trajectory": "tr_01JBA3KX2M",
  "tool": "shell",
  "cwd": "/Users/dana/src/acme-billing",
  "command": "psql $DATABASE_URL -f refunds.sql",
  "env": { "DATABASE_URL": "postgres://billing:hunter2@10.0.3.12:5432/prod" },
  "stdout": "maria.lopez@example.com | 4 refunds\njwu@example.org | 2 refunds\n(5 rows)",
  "prompt": "Why are refunds failing for enterprise customers?",
  "user": "dana@acme.dev",
  "duration_ms": 184,
  "exit_code": 0
}

relay rules~/.ground/relay.toml

Also: sampling per agent and tool, local persistence, retention windows.

uploaded event5 rules applied

{
  "event": "tool.completed",
  "trajectory": "tr_01JBA3KX2M",
  "tool": "shell",
  "cwd": "path:7c1e9a42",
  "command": "psql $DATABASE_URL -f refunds.sql",
  "env": { "DATABASE_URL": "[SECRET:postgres_url]" },
  "stdout": "[EMAIL] | 4 refunds\[EMAIL] | 2 refunds\n(5 rows)",
  "prompt": "[EXCLUDED · 11 tokens]",
  "user": "user:3f9a0c1e",
  "duration_ms": 184,
  "exit_code": 0
}
09 otlp.export fits your stack

Speaks OpenTelemetry. Adds agent semantics.

Ground Station doesn’t replace your tracing. Trajectories export as OTLP traces, so the agent run sits inside the same trace as the request that started it, and the 18 minutes stop being one opaque span.

How Ground Station concepts map to OpenTelemetry
ground stationotelname
trajectory→ traceground.trajectory.*
model invocation→ spangen_ai.chat
tool invocation→ spangen_ai.execute_tool
subagent→ child spangen_ai.invoke_agent
events→ span eventsfile.read · shell.exit
measurements→ metricsgen_ai.client.token.usage
trace 4bf92f3577b34da6a3ce929d0e0e4736
  • POST /api/tasksweb
  • tasks.createapi
  • queue.enqueueapi
  • ground.trajectoryRefactor payments module
  • gen_ai.chatsonnet-5.5
  • gen_ai.execute_toolread_file
  • gen_ai.chatsonnet-5.5
  • gen_ai.execute_toolshell · cargo test
  • gen_ai.execute_tooldatabase · psql
  • db.querypostgres
  • gen_ai.invoke_agentsubagent · explore
  • tasks.update_statusapi
  • + 106 more agent spans · 2 span events each

span attributes · gen_ai.execute_tool · cargo test

gen_ai.operation.name
execute_tool
gen_ai.tool.name
shell
ground.trajectory.id
tr_01JB8W4ZQD
ground.tool.command
cargo test --workspace
ground.tool.exit_code
101
ground.context.tokens
184,112

Every agent leaves a trajectory. Make it observable.

Ground Station is in private beta. The collector, adapters and schema are open source. Bring one agent, and see its first trajectory in minutes.

Get early access Star on GitHub

The observability layer for autonomous software

  1. V0Datacapture everythingbeta
  2. V1Informationreadable trajectoriesbetayou are here
  3. V2Knowledgepatterns & explanationsnext
  4. V3Actionrecommendationslater