---
title: "5 Metrics That Prove Whether Your AI Assistant Is Working"
url: "https://referent.app/blog/measure-ai-assistant-roi-metrics"
description: "Logins and message counts tell you nothing. These five metrics, from action depth to interruption deflection, tell you whether your AI assistant is producing real ROI, how to baseline each one, and what healthy looks like."
---

Guides7 min read

# 5 Metrics That Prove Whether Your AI Assistant Is Working

Rodrigo Mahamud

·

September 1, 2026

![5 Metrics That Prove Whether Your AI Assistant Is Working](/images/blog/measure-ai-assistant-roi-metrics.svg)

The metrics that prove an AI assistant is working are the ones tied to completed work: actions per user, time to resolution, workflows finished end to end, expert interruptions avoided, and hours recovered. Vanity numbers like seat activations and message counts can all rise while the assistant produces nothing. This guide gives you five metrics that cannot be gamed, with baselines and healthy ranges for each.

This article is for operations leaders and executive sponsors who need to defend, expand, or kill an AI assistant investment with evidence. You’ll learn:

*   Why the default usage metrics mislead
*   The five metrics that track real value
*   How to baseline each one before rollout
*   The review cadence that keeps the number honest

## Table of Contents

*   [Why Usage Metrics Mislead](#why-usage-metrics-mislead)
*   [Metric 1: Action Depth per Active User](#metric-1-action-depth-per-active-user)
*   [Metric 2: Time to Answer, Time to Action](#metric-2-time-to-answer-time-to-action)
*   [Metric 3: End-to-End Workflow Completion](#metric-3-end-to-end-workflow-completion)
*   [Metric 4: Interruption Deflection](#metric-4-interruption-deflection)
*   [Metric 5: Hours Recovered, Priced](#metric-5-hours-recovered-priced)
*   [The Review Cadence](#the-review-cadence)
*   [Frequently Asked Questions](#frequently-asked-questions)

## Why Usage Metrics Mislead

Every AI dashboard leads with the same numbers: seats activated, messages sent, daily active users. They are easy to collect and they all share the same flaw: they measure attention, not outcomes. A confused team generates high message counts. A failing rollout looks identical to a thriving one right up until renewal.

This is how organizations end up in the [adoption gap](/blog/ai-adoption-gap-why-impact-lags): usage everywhere, impact nowhere, and no instrument on the dashboard capable of telling the difference. The five metrics below are chosen for one property: each one only moves when real work gets done.

> **Key Takeaway:** If a metric can rise while zero work gets completed, it is not an ROI metric. It is a engagement statistic wearing a suit.

## Metric 1: Action Depth per Active User

**What it is:** completed actions (reports compiled, meetings scheduled, records updated, briefs delivered) per active user per week. Not messages. Actions.

**Why it matters:** it separates teams that chat with the assistant from teams that work through it. Depth is where the chatbot-versus-agent distinction from [AI Agents vs Chatbots](/blog/from-chatbots-to-ai-agents-that-act) becomes measurable.

**Baseline and healthy range:** baseline is zero by definition. In week one expect low single digits per user. A healthy connected deployment trends toward 10 to 20 actions per user per week by month two. Flat depth with high message volume means people ask, distrust the answer, and do the work manually: investigate retrieval quality.

## Metric 2: Time to Answer, Time to Action

**What it is:** elapsed time from question to verified answer, and from request to completed action, for your five most common cross-system requests. Measure the manual path before rollout with a stopwatch. Literally.

**Why it matters:** this is the metric executives feel. “Where is the Meridian contract?” going from 15 minutes of hunting (or a 4-hour wait for someone to reply) to 10 seconds with a citation is the product working, quantified.

**Baseline and healthy range:** typical manual baselines run 10 to 45 minutes for cross-system questions, per the app-switching toll documented in [Asana’s Anatomy of Work](https://asana.com/resources/anatomy-of-work). Healthy assistant performance is seconds for answers and under five minutes for confirmed multi-step actions.

## Metric 3: End-to-End Workflow Completion

**What it is:** for each workflow you have delegated (the weekly report, meeting prep, lead routing), the share of instances the assistant completes without a human stepping in to rescue it, excluding designed confirmation stops.

**Why it matters:** workflow completion is where individual convenience becomes organizational capacity. It is also your quality alarm: a completion rate that sags after an integration change tells you something broke before your users do.

**Baseline and healthy range:** measure per workflow. Above 90 percent, expand the workflow’s scope or remove a confirmation step that has become a rubber stamp. Below 70 percent, the workflow was under-specified: tighten it as a [skill](/blog/ai-skills-tribal-knowledge-playbooks) rather than blaming the tool.

## Metric 4: Interruption Deflection

**What it is:** questions answered by the assistant that previously interrupted a specific person: the ops lead who knows where everything is, the senior engineer, the one person who understands the pricing sheet.

**Why it matters:** interruptions are the most expensive tax on senior time. UC Irvine research puts the cost of regaining focus after an interruption at [over 23 minutes](https://www.ics.uci.edu/~gmark/chi08-mark.pdf). Every deflected question returns that time to your scarcest people.

**Baseline and healthy range:** have your three most-interrupted people tally interruptions for one week before rollout (a sticky note works). Re-tally monthly. Deflecting half is common within two months once the assistant is connected to real context; your experts will report the difference unprompted, which is its own signal.

## Metric 5: Hours Recovered, Priced

**What it is:** the roll-up: (baseline time minus current time) × instances per week, summed across measured workflows and deflected interruptions, multiplied by loaded hourly cost.

**Why it matters:** this is the number that survives a budget meeting. Time-to-answer improvements are felt; hours-priced is defensible. It is also the number that decides expansion: when one team shows 30 recovered hours weekly, the second team’s business case writes itself.

**Baseline and healthy range:** conservative teams count only measured workflows and deflected interruptions, ignoring diffuse gains. Even so, a 20-person team typically shows 25 to 50 recovered hours weekly by month three. Price honestly, count conservatively, and the number is bulletproof.

## The Review Cadence

Metrics without cadence decay into dashboard wallpaper:

1.  **Before rollout:** baseline metrics 2, 4, and 5 manually. One week of measurement is enough. Skipping this step is the single most common measurement mistake, and it is unrecoverable.
2.  **Weekly for the first month:** review depth and completion with the pilot team. Fix retrieval and workflow specs while attention is high.
3.  **Monthly thereafter:** the five metrics on one page, owned by one named person, per the governance model in the [Agent Development Lifecycle](/blog/agent-development-lifecycle-enterprise).
4.  **Quarterly:** re-price metric 5, decide expand or fix, and re-baseline any workflow whose definition changed.

## Frequently Asked Questions

### We already rolled out without baselines. Is it too late?

You lose the clean before-and-after, but not the program. Baseline now, measure the delta of your next workflow properly, and use time-to-answer comparisons against the manual path, which can be reconstructed with a stopwatch any time.

### Should we survey users about satisfaction?

As a complement, yes: a two-question pulse (“did the assistant save you time this week? where did it fail?”) catches qualitative signals the metrics miss. As the primary measure, no. Satisfaction is a lagging echo of the five metrics, not a substitute.

### What if the metrics say it is not working?

Then you learned it cheaply, which is the point of measuring. Diagnose in order: retrieval quality (are answers right?), context coverage (are the right systems connected?), workflow specification (is the task well-defined as a skill?), and placement (is the assistant [where the team already works](/blog/ai-assistant-slack-whatsapp-telegram)?). Most failures live in those four, and all four are fixable.

### Who should own these metrics?

One named person with authority to change the deployment: typically the operations lead who owns the pilot. Committee ownership is how five metrics become zero metrics.

## Measure Like You Mean It

An AI assistant is an investment, and investments get instruments. Five numbers, baselined honestly and reviewed on cadence, will tell you what anecdotes never can: whether the thing works, where it does not, and what the next dollar should fund.

**Want metrics from day one?** Referent logs every action it takes, so depth, completion, and time-to-action come from the audit trail rather than surveys. [Book a 15-minute demo](/demo) and we will help you design the baseline week.

* * *

_Sources: [Asana — Anatomy of Work Global Index](https://asana.com/resources/anatomy-of-work) · [Gloria Mark et al., UC Irvine — The Cost of Interrupted Work](https://www.ics.uci.edu/~gmark/chi08-mark.pdf) · Related: [The Agent Development Lifecycle](/blog/agent-development-lifecycle-enterprise)_

Share this article

[](https://www.linkedin.com/shareArticle?mini=true&url=https%3A%2F%2Freferent.app%2Fblog%2Fmeasure-ai-assistant-roi-metrics&title=5%20Metrics%20That%20Prove%20Whether%20Your%20AI%20Assistant%20Is%20Working)

## Be the first to hear about Referent news.