If You Run an AI Agent, You Need to See What It Does

A chatbot that gives a wrong answer is an annoyance. An agent that sends an email on its own, changes a record or drafts an offer is something else. It acts — and anything that acts has to be able to account for itself.

Most teams putting agents into production discover this at the first incident. The question lands on the table: what exactly happened here? And the honest answer is, surprisingly often, that nobody can say any more.

Why logs are not enough for agents

Classic logging was built for software that executes instructions. One line per event, timestamp, severity level. For a web server that is appropriate.

An agent works differently. Given an input, it decides for itself which tools to use, calls one, judges the result, possibly calls another, and eventually formulates an answer. We call that sequence a turn. What matters is not the individual event but the connection between them: which model was involved, which tools were called with which arguments, what came back — and at what point the decision was made that later turned into the problem.

A chat transcript shows only the surface of that: question in, answer out. Everything in between — which is precisely where the acting happens — stays invisible.

Turn tracing: the sequence as the unit

So we built observability around the turn rather than around the log line. A recorded turn opens up like a protocol: every tool call with its arguments, every result, every model call, in the order it actually happened.

That makes a practical difference. Faced with “why did the agent contact this customer?”, you no longer reconstruct a probable story from scattered log lines. You look at the sequence.

One detail we insisted on: the system does not fill gaps. Where execution history is missing, it says so rather than substituting a plausible-looking reconstruction. Evidence that guesses at the uncertain parts is worse than none, because people believe it.

Completeness you can verify

An agent runs in its own container, often on the customer’s side, and reports its events back to the platform. That raises an uncomfortable question: how do you know nothing was lost on the way — or altered afterwards?

Every event therefore carries a running sequence number and a hash that incorporates its predecessor. On arrival the platform checks both: does the number follow the previous one without a gap, and does the hash match the chain so far? If not, the state is marked broken, with the reason recorded.

This is deliberately stricter than monitoring. Monitoring tells you whether a system is running. A verified event chain tells you whether the record of what it did is complete. For a system that writes into customer data on its own, that is the difference between a log and evidence.

Noticing before someone asks

The same data feeds alerts. Four types cover most of what matters in practice:

  • Agent offline — the agent has stopped reporting. For a system meant to work in the background, this otherwise surfaces only when a result fails to appear.
  • Spike in denied actions — the agent repeatedly attempts something it lacks the rights for. Either its task is scoped wrong, or it is heading somewhere nobody intended.
  • Spike in errors — the failure rate rises against normal operation.
  • Chain integrity broken — the verification above has failed.

Each alert carries a severity from low to critical and goes to the dashboard, to email, or to a webhook. The point is not the notification itself but that the anomaly is derived from the agent’s behaviour rather than from system load.

See it in the interface

This short video walks through what it looks like in practice — the dashboard, turn search, and the detail view of a single run:

Why this is not a side issue

There is an operational reason and a regulatory one.

Operationally: an agent whose behaviour you cannot follow cannot be improved. You can switch it off or keep trusting it, but you cannot fix the specific point where it went wrong — because you do not know where that point is.

Regulatory: the AI Act requires logging and human oversight for systems at the corresponding risk level. “We have the chat transcript” does not carry that weight. A verified, gap-free record of which tools were called with which data comes considerably closer.

Our position on this is inconvenient but we think it is right: if you deploy agents that act independently, you take responsibility for that action. Observability is the instrument that makes the responsibility exercisable in the first place. Without it, all that is left is trust — and for autonomous systems, trust without verifiability is a poor foundation.

Agent observability is part of the HybridAI platform and available for hosted HybridClaw agents.

Your CRM Problem Isn’t Data — It’s Input

Every company running a CRM hears the same complaint from sales: “This system costs me more time than it gives back.” The usual responses are training, mandatory fields, reminder emails, and eventually pressure from management. They rarely work, and the reason is structural.

Why CRM records are incomplete

A field rep has a forty-minute meeting at a customer site. Inside it: a quantity, a reservation about price, the news that the actual decision-maker just changed, a complaint about the last delivery, and a sense of how likely the deal is to close.

How much of that reaches the CRM depends on whether that person — two meetings later, on a train, or at home in the evening — can summon the energy to fill in eleven form fields. Usually one sentence makes it. Sometimes nothing does.

This is not a discipline problem. It is an interface problem. The conversation is spoken language: unstructured, full of context. The CRM expects form fields. Between the two sits a translation job someone has to do by hand — and it falls to the person whose time is the most expensive in the company.

Most of the AI conversation over the past few years has focused on the opposite direction: better analysis of the data that is already there. Dashboards, forecasts, lead scoring. None of that is wrong, but it treats the symptom. If half of what happens in the market never enters the system, even the best model on top of it is analysing the gaps.

Why realtime voice is different from chat

Voice control for business software has existed for twenty years and was almost always disappointing. What changed is not speech recognition — that was already good enough — but what happens between recognition and execution.

A realtime voice model does not listen and transcribe. It understands a spoken paragraph, identifies several distinct facts inside it, routes them to different target structures, and asks when something is missing. So

“Just left Müller GmbH. They want 200 units at the Q3 price, decision by end of month — I’d put it at 70 percent.”

becomes three separate records: a visit report, an offer draft, and a deal assessment. Each lands where it belongs, with the right fields filled in.

Timing is what makes it work. This happens in the car park, two minutes after the meeting, while the details are still sharp — not in the evening, when the memory has worn down and the motivation is gone.

The second difference from chat is that it is hands-free. A rep driving between appointments cannot type, but can talk. And that window — between two meetings — is exactly when people are most willing to deal with something that otherwise gets postponed indefinitely.

Reading and writing belong together

Dictation alone would only be half a solution. It gets interesting when the same voice also answers questions.

“What do I need to know about this customer?” on the drive over — and the answer covers open opportunities, recent orders, notes from the last visit, and the relevant emails from recent weeks. Same data as on the desktop, through an interface that works in a car.

Technically these are two different jobs. Questions about revenue, quantities and pipeline have to become exact queries over structured data — nothing may be estimated, numbers are numbers. Questions about background, conversation history or product detail have to run semantically over unstructured documents: meeting notes, emails, spec sheets, contracts.

We deliberately split this into two paths — text-to-SQL for the figures, semantic search for the story behind them — and let a third model combine and explain the results. The arithmetic is never done by the language model, always by code. An offer whose total an LLM “estimated” would be worthless.

What this looks like in practice: Sales Companion

That is what we built the Sales Companion for: an iPhone app with CarPlay support that sits on the existing CRM as a voice interface.

Sales Companion app: asking customer data by voice on iPhone
The Sales Companion voice view: ask a question instead of hunting for a form.

Worth stressing: it is not a new CRM. The data stays in Salesforce, SAP, Dynamics or Zoho, where it already lives. A layer goes on top, reaching those systems through adapters. No migration, no parallel system, no retraining for the teams who keep working at their desks in the tool they know.

Day to day, that means:

  • Briefing before the meeting — what matters about this customer, summarised on the way there.
  • Visit report after it — dictated rather than typed, with customer, contacts, topics and next steps.
  • Offer drafts — spoken quantities and prices become a draft, calculated exactly.
  • Analysis on demand — “Which customers ordered less this quarter than last?” in seconds, without waiting for a BI report.

One part surprised us with how well it lands: the same voice that knows your customers is also good for practice. A roleplay where the AI plays the counterpart, raises objections and pushes back — followed by a score per skill and concrete advice on what to work on. For new sales hires that beats any handbook.

Privacy is not a footnote

If you send customer conversations through a language model, you need to know where that data goes. Our infrastructure runs on European servers, GDPR and AI Act compliant. Personal data can be masked before a prompt reaches a model, and role-based access controls which teams can reach which data sources at all.

This is not a feature added later. For a system that records visit reports about real people, it is the precondition for being allowed to deploy it.

See for yourself

On our Agentic CRM page you can watch the demo videos and try the voice interface directly — the public demo runs on sample sales data, no sign-up required.

The more interesting question was never whether the technology works. It is whether your sales team will use it. Our experience so far: when the input takes two minutes in the car park instead of twenty minutes in the evening, they do.