Your Agent Runs. Do You Know What It Did?
AI agents in production rarely fail dramatically. More often an agent calls the same tool twenty times, burns a day's token budget in a single session, or finishes a task with an answer that looks right and isn't. Without observability, three basic questions — what did the agent do, why did it fail, and what did it cost — only get answered after they have become incidents.
Hugging Face's agent glossary defines the harness as "the execution layer inside the agent: it calls the model, handles its tool calls, decides when to stop." Since every step an agent takes passes through the harness, that is the natural place to instrument it. This article covers what to record, why it has to be there from day one, how to verify results, how to attribute cost, and how to recover when an agent goes wrong.
What to Trace
Hugging Face calls one full agent run from start to finish a rollout, also known as a trajectory or a trace: what the agent saw, what it did, and the outcome at each step. A trace that is useful in production records at least:
| Data | Why it matters |
|---|---|
| The prompt at each step | Confirms the model received the context you think it did |
| The system prompt actually sent | Templates are often assembled dynamically; what sits in the repo is not necessarily what went out |
| Tool calls (name + arguments) | Shows the model's decisions, not just the final outcome |
| Tool results | Many failures start with a tool output that is empty, truncated, or an error |
| Input/output tokens | The basis for cost and for spotting a bloating context |
| Latency per step | Separates a slow model from a slow tool |
The system prompt point is the one teams skip. Scaffolding, in Hugging Face's definition, covers the system prompt, tool descriptions, how the model's responses get parsed, and what it remembers across steps. If any of that is assembled from variables, store the final version that hit the API, not the template.
Observability Belongs in the Harness From Day One
DataNorth's harness engineering guide puts it concisely: the model supplies the reasoning; the harness decides what the model sees, what it can do, what it remembers, and where it runs. Major frameworks already treat tracing as a core component — the OpenAI Agents SDK, for example, ships Tracing as a primitive alongside Agents, Handoffs, Guardrails, Sessions and State.
The reason is that agent errors compound. DataNorth cites research showing per-step accuracy degrades as the number of steps grows, and METR data showing Claude Opus 4.6 sustains 718.8-minute tasks at a 50% success rate but only 69.9 minutes at 80%. The longer the task, the more likely some step goes wrong — and without a trace you will not know which one. Adding logging after the first outage means that first outage stays undiagnosed.
Computational vs Inferential Checks
DataNorth distinguishes two feedback mechanisms: guides, which steer the agent before it acts, and sensors, which observe afterwards so the agent can self-correct. "Tests, linters, type checks and hooks all qualify," it notes. Within sensors it helps to separate two kinds of check:
- Computational — deterministic, cheap and repeatable: unit tests, linters, type checkers, JSON schema validation. Firecrawl describes coding-agent harnesses that run the test suite after each feature and only check it off when the tests pass.
- Inferential — another model acting as a grader (LLM-as-judge) for qualities rules cannot measure, such as relevance or tone. More flexible, but more expensive and not fully deterministic.
Run the computational checks first because they are cheapest, and reserve LLM-as-judge for what genuinely needs judgment. Record both in the trace. One trick from Firecrawl is worth copying: OpenAI wrote linter error messages specifically to teach the fix, so every failure message became context for the next attempt.
Cost Attribution: Per Run, Per Tool, Per User
A budget cap only means something if you know where the tokens go. Tag every model call with a run_id, the tool that triggered the step, and the user_id or tenant. That lets you answer which tool is the most expensive, which users trigger the costliest runs, and whether cost per run rose after a prompt change.
The same data shows where optimization pays off. DataNorth reports that Vercel cut token usage from roughly 102,000 to 61,000, and execution time from 274.8 to 77.4 seconds, by replacing 15 specialized tools with bash. Anthropic reportedly cut one workflow from 150,000 tokens to 2,000 by presenting tools as code modules, and a PreToolUse hook filtered test output from tens of thousands of tokens down to hundreds. Savings like these are only visible when cost is recorded per tool.
Failure Recovery Patterns
Bounded retries. Retry transient failures (timeouts, rate limits) with backoff and a cap on attempts. Do not retry a logic failure with the same prompt — feed the error back as context, as in the linter pattern above.
Checkpoints. Save state after significant steps so a failed run can resume instead of starting over. DataNorth notes that LangGraph provides short-term memory through checkpointers, and an interrupt function that saves state and waits for a human when a step needs judgment. Firecrawl adds that JSON tracking files hold up better than Markdown because models are less likely to corrupt them.
Stopping a looping agent. The harness is what "decides when to stop," so give it hard limits: a maximum step count, a per-run token budget, and detection of repeated tool calls with identical arguments. When a limit trips, save a checkpoint, mark the run as failed, and alert someone — do not let the agent keep burning tokens.
Human in the loop. For sensitive actions, Firecrawl recommends interrupts that pause the agent until a human approves.
Conclusion
Observability for AI agents is not an add-on; it is part of the harness. Record the prompt, the real system prompt, tool calls, tool results, tokens and latency at every step; pair computational checks with LLM-as-judge; attribute cost per run, tool and user; and put retries, checkpoints and stop conditions in place. Then, when an agent misbehaves in production, the answer is already in the trace rather than in a guess.