Every AI agent dashboard leads with the same proud number: conversations handled. It is the emptiest metric in the building. A conversation handled badly is worse than no conversation, and a thousand of them is a thousand small withdrawals from your brand. Judging an agent requires numbers that carry consequences, and five of them do most of the work.
Containment rate, read with suspicion
Containment is the share of interactions the agent completes without human help, and it is the headline metric of the industry. It is useful, and it lies. High containment can mean the agent is genuinely resolving things, or it can mean customers are giving up inside the conversation, which counts as contained because no human was involved.
The honest reading pairs containment with what the contained conversations produced. A contained booking call ends in a booking. A contained question ends with the answer delivered and the customer gone happily. Containment without an outcome attached is just the absence of escalation, and absence is not success.
Escalation quality, the metric almost nobody tracks
Some conversations should reach a human, and the agent's job in those is to hand off well. Escalation quality asks: when the agent passed a conversation over, did it pass at the right moment, to the right person, with the context intact? The measurable proxies are time-to-human after the escalation trigger, and how often the human had to re-ask questions the agent already asked.
This number matters because it prices the agent's judgment. An agent that escalates too late frustrates the exact customers who most needed a person; one that escalates too early is an expensive phone tree. The transcripts of escalated conversations, reviewed monthly, are where this gets tuned, and that review is a standing part of how we run client automations.
Booking yield, where the agent meets the P&L
For a service business, the agent's revenue story runs through one funnel: conversations with booking intent, bookings made, bookings held. Yield is the ratio through those stages, and it is the number that lets an agent stand in a lineup with your other lead sources and justify its existence in dollars rather than vibes.
Watch the held stage hardest. An agent can inflate bookings by making the calendar too easy to claim, and the no-show rate will quietly tell you whether those bookings were commitments or clicks.
Abandonment, the silence inside the numbers
Abandonment is the share of customers who start with the agent and simply stop: mid-conversation, mid-booking, mid-question. Every abandonment is a customer who tried you and left without a resolution or an escalation, invisible in both containment and booking counts.
The diagnostic power is in where the drop happens. Abandonment clustered at a specific question usually means the question is asked too early or reads as invasive. Clustered after a long agent reply, the replies are too long. Spread evenly, the channel may be attracting traffic with no real intent. The location of the leak is the repair order.
Correction rate, the trust thermometer
Correction rate is how often a human had to fix something the agent did: a wrong booking detail, a misfiled record, a bad answer caught after the fact. It is the metric that governs how much autonomy the agent should have next quarter. A falling correction rate earns expanded scope; a rising one triggers investigation before anything else grows. Tracked honestly, it is also your early-warning system for drift after knowledge-base changes, which is why it belongs on the monthly report rather than in an incident post-mortem.
The benchmark trap, and what to compare against instead
The first question owners ask about each number is what a good one looks like, and the honest answer is: compared to whom? Published benchmarks for agent metrics blend industries, channels, and definitions so thoroughly that they mostly measure who chose to publish. A dental office's containment rate and an emergency plumber's are different animals; the plumber's calls should escalate more, because more of them genuinely need a human decision, and a vendor bragging about ninety-percent containment on that line is describing a hazard, not an achievement.
The comparisons that actually inform are internal and longitudinal. This month against last month, after the knowledge-base update. The after-hours line against the daytime line, which isolates what staffing coverage is worth. The agent's booking yield against what the same channel produced before automation, which is the only apples-to-apples judgment of whether the deployment paid. And where a business runs multiple locations or lines, the spread between them, because an outlier location is a finding either way: something to fix or something to copy.
The discipline this implies is unglamorous: freeze the definitions, keep the history, and resist re-basing the numbers when a change makes last quarter look bad. A metric series that survives honest continuity for a year becomes the most valuable measurement asset the deployment has, worth more than any industry report, because it describes the only business whose decisions you are making.
The report an owner should actually receive
One page, monthly: the five numbers, their trend, three transcript excerpts worth reading, and the single change recommended next. That is the whole discipline. It takes an hour to produce from a well-instrumented system and it converts the agent from a black box into a managed employee with a performance record.
This report is a standard deliverable in our managed engagements, and our results reporting rolls these numbers into the same view as the rest of the funnel. If your current agent reports conversations handled and nothing else, get in touch and we will instrument it properly, or tell you honestly that it is not worth instrumenting.