Cover: See what your AI agent does – turn trace of an AI agent

If You Run an AI Agent, You Need to See What It Does

A chatbot that gives a wrong answer is an annoyance. An agent that sends an email on its own, changes a record or drafts an offer is something else. It acts — and anything that acts has to be able to account for itself.

Most teams putting agents into production discover this at the first incident. The question lands on the table: what exactly happened here? And the honest answer is, surprisingly often, that nobody can say any more.

Why logs are not enough for agents

Classic logging was built for software that executes instructions. One line per event, timestamp, severity level. For a web server that is appropriate.

An agent works differently. Given an input, it decides for itself which tools to use, calls one, judges the result, possibly calls another, and eventually formulates an answer. We call that sequence a turn. What matters is not the individual event but the connection between them: which model was involved, which tools were called with which arguments, what came back — and at what point the decision was made that later turned into the problem.

A chat transcript shows only the surface of that: question in, answer out. Everything in between — which is precisely where the acting happens — stays invisible.

Turn tracing: the sequence as the unit

So we built observability around the turn rather than around the log line. A recorded turn opens up like a protocol: every tool call with its arguments, every result, every model call, in the order it actually happened.

That makes a practical difference. Faced with “why did the agent contact this customer?”, you no longer reconstruct a probable story from scattered log lines. You look at the sequence.

One detail we insisted on: the system does not fill gaps. Where execution history is missing, it says so rather than substituting a plausible-looking reconstruction. Evidence that guesses at the uncertain parts is worse than none, because people believe it.

Completeness you can verify

An agent runs in its own container, often on the customer’s side, and reports its events back to the platform. That raises an uncomfortable question: how do you know nothing was lost on the way — or altered afterwards?

Every event therefore carries a running sequence number and a hash that incorporates its predecessor. On arrival the platform checks both: does the number follow the previous one without a gap, and does the hash match the chain so far? If not, the state is marked broken, with the reason recorded.

This is deliberately stricter than monitoring. Monitoring tells you whether a system is running. A verified event chain tells you whether the record of what it did is complete. For a system that writes into customer data on its own, that is the difference between a log and evidence.

Noticing before someone asks

The same data feeds alerts. Four types cover most of what matters in practice:

  • Agent offline — the agent has stopped reporting. For a system meant to work in the background, this otherwise surfaces only when a result fails to appear.
  • Spike in denied actions — the agent repeatedly attempts something it lacks the rights for. Either its task is scoped wrong, or it is heading somewhere nobody intended.
  • Spike in errors — the failure rate rises against normal operation.
  • Chain integrity broken — the verification above has failed.

Each alert carries a severity from low to critical and goes to the dashboard, to email, or to a webhook. The point is not the notification itself but that the anomaly is derived from the agent’s behaviour rather than from system load.

See it in the interface

This short video walks through what it looks like in practice — the dashboard, turn search, and the detail view of a single run:

Why this is not a side issue

There is an operational reason and a regulatory one.

Operationally: an agent whose behaviour you cannot follow cannot be improved. You can switch it off or keep trusting it, but you cannot fix the specific point where it went wrong — because you do not know where that point is.

Regulatory: the AI Act requires logging and human oversight for systems at the corresponding risk level. “We have the chat transcript” does not carry that weight. A verified, gap-free record of which tools were called with which data comes considerably closer.

Our position on this is inconvenient but we think it is right: if you deploy agents that act independently, you take responsibility for that action. Observability is the instrument that makes the responsibility exercisable in the first place. Without it, all that is left is trust — and for autonomous systems, trust without verifiability is a poor foundation.

Agent observability is part of the HybridAI platform and available for hosted HybridClaw agents.