How do you measure whether a voice agent works?

Containment rate, transfer rate, task completion, and what actually happens after the call end up mattering more than call volume or a generic satisfaction score.

An admin approves every new account by hand. Nothing is created until then. We reply by email; no newsletter, no sequence.

app.salescrew.io/inbox
The unified reply inbox with classified threads

The short answer

  • Containment rate, the share of calls the agent resolves without transferring to a human, is the headline metric most vendors lead with, but it is only meaningful alongside task completion, since a contained call that did not actually solve the caller's problem is a false positive.
  • Transfer rate is not automatically bad. What matters is whether transfers happen for legitimate reasons, an explicit request, a genuinely out-of-scope topic, low confidence, rather than the agent failing on things it should be able to handle.
  • Task completion, whether the call achieved its actual purpose (booked, answered, resolved), is the metric closest to business value, and it requires checking outcomes, not only conversation flow.
  • What happens after the call, whether a booking sticks, whether a promised callback actually happens, whether the caller's issue stays resolved, is the part most dashboards do not show but is often the real test of whether the agent is working.

Why volume and star ratings are not enough

A dashboard showing call volume and a satisfaction score feels like measurement, but neither number says much about whether the agent is actually doing its job well. Volume just says calls are coming in and getting answered, which was already true before the AI agent existed if a human or voicemail was picking up. A satisfaction score, especially one gathered from a short post-call survey, tends to reflect how pleasant the voice sounded more than whether the caller's actual problem got solved.

The more useful metrics require looking past the call itself and into whether it accomplished something. Containment rate, how often the agent handles the call without a human, is a reasonable starting point, but read alone it rewards an agent that ends calls quickly even when it failed to actually help, since ending the call without a transfer still counts as "contained." Pairing containment with task completion, did the caller actually get what they called for, catches that gap.

What to actually track

MetricWhat it tells youWhat it misses alone
Containment rateHow often the agent avoids a human transferWhether the contained call was actually resolved
Transfer rateHow often the agent hands off to a humanWhether the transfer was for a legitimate reason
Task completionWhether the call's actual purpose was achievedWhat happens after the call ends
Post-call outcomeWhether a booking sticks, a promise is keptRequires tracking beyond the call itself

How to actually build this measurement

Start by defining what a successful call looks like for each call type your agent handles, a booking confirmed, a question actually answered, a transfer completed cleanly, rather than using one generic success definition for every call. Then periodically review a sample of calls, especially transfers and calls the agent marked as contained, to check whether the automated outcome matches what a human reviewing the transcript would call a success.

Disclosure: SalesCrew is our product, and voice agents are on our roadmap and not shipped today, so this page describes general measurement practice rather than a SalesCrew reporting feature available now. The framework here, containment paired with task completion, applies to evaluating any vendor's voice agent.

A high containment rate can hide a low-quality agent

An agent that avoids transferring but also fails to actually resolve the caller's issue looks good on containment and bad on the outcome that matters. Always pair containment with task completion, never read it alone.

Questions

What is containment rate?
The share of calls the agent handles fully on its own without transferring to a human. A high containment rate is only good if the calls being contained were actually resolved correctly, which is why it needs to be read alongside task completion, not alone.
Is call volume a useful metric?
Volume tells you the agent is answering calls, not whether those calls went well. It is worth tracking as context but is not evidence of quality by itself.
How do you know if a transfer happened for a good reason?
Review a sample of transferred calls periodically and check whether the transfer trigger, an explicit request, a topic outside scope, low confidence, matches what actually happened on the call, rather than assuming every transfer was necessary.