Measuring response time and agent workload
Most inbox dashboards flatter you. Here is how to read yours honestly.
The useful inbox measures are first response time, resolution time, open backlog and workload per person. Everything else is either derived from those or a vanity number.
The harder question is not which to watch but whether yours are counting what you think they are.
The four worth watching
| Measure | Tells you | Watch out for |
|---|---|---|
| First response time | How long a customer waits for any human answer | Bot replies counting as responses |
| Resolution time | How long the whole thing took | Conversations never marked resolved |
| Open backlog | What is outstanding right now | Only meaningful if status is used consistently |
| Conversations per agent | Whether the load is shared | Says nothing about how hard each one was |
Every one of those depends on something upstream being done properly — status, assignment, or the distinction between a bot reply and a human one.
The counting rule that decides everything
A reply that came from an automation is not a human response, and if your reporting treats it as one, your first response time will look excellent while customers wait exactly as long as before.
This is the single most common way inbox dashboards mislead. The bot answers in two seconds, the metric records two seconds, and the customer who needed a person waits an hour. Make sure your reporting separates the two, and read a few real conversations to sanity-check it.
What "resolved" has to mean
Resolution time is only meaningful if resolved means resolved. If half your conversations are never marked, your average is computed from a biased sample — the easy ones — and it will look better than reality.
Two habits fix this. Automate the obvious transitions, and reopen a conversation when the customer replies to it. The conventions are in open, pending, resolved.
Workload is not performance
Conversations per agent is genuinely useful for spotting an unbalanced rota and genuinely misleading as a performance measure. Someone handling 10 complaints has had a harder day than someone handling 40 stock questions.
If you want to compare people, compare like with like — the same category of conversation, over a long enough period to smooth out the days when the difficult ones happened to land on one person.
The measures worth ignoring
- Total messages sent. Rewards verbosity.
- Conversations handled per hour, in isolation. Rewards rushing.
- Bot deflection rate, unless you also know how many of those customers came back unhappy.
- Anything with no denominator. "500 conversations" is not a number until you know out of how many, over what period.
None of those are wrong to look at. They are wrong to target, because a team told to move a number will move it.
Satisfaction, measured properly
A rating question at the end of a conversation is the one measure that comes from the customer rather than from your own logs. Asked as buttons after a resolution, it is cheap and it is the only thing on this page that tells you whether the fast resolution was actually any good.
Keep it to one question, ask it only after the conversation has genuinely ended, and do not chase people who ignore it. The setup is in customer satisfaction scores from a flow.
A monthly review that takes twenty minutes
- First response time, humans only, compared with last month.
- Open backlog at the same time of day each month.
- Anything pending for more than a week, read individually.
- Workload spread across the team.
- Five real conversations, read end to end.
Step five is the one that finds things the other four cannot. Numbers tell you something changed; conversations tell you what.
Where cost shows up
Slow first responses have a price as well as a cost in goodwill. Replies inside the 24-hour customer service window are free and unlimited under Meta's pricing documentation, and the window opens when the customer messages you, per Meta's sending messages documentation.
Let it close and your next message needs an approved template. A team that consistently answers the same day is quietly saving money as well as customers.
Frequently asked
What is a good first response time?
Faster than you manage now. Absolute benchmarks vary so much by category that chasing somebody else's number is not useful; your own trend is.
Should I measure the bot separately?
Yes. It answers a different kind of question, and mixing the two hides both. The automation side is in measuring whether your automation works.
Why does one agent always look slower?
Frequently because they get the hard conversations. Check the mix before you draw a conclusion.
What if we have no assignment at all?
Then most of this cannot be measured, and that is the first thing to fix.
Getting a baseline before you change anything
The most common mistake is improving something before knowing what it was. Spend a month collecting numbers you have not acted on, and everything afterwards has something to be compared against.
- First response time, humans only, recorded weekly.
- Open backlog at a fixed time each day.
- Conversations per agent per week.
- The share of conversations the bot handled end to end.
- How many were reopened after being resolved.
The fifth is the quality check on the fourth. A high automation share with a high reopen rate is not automation working; it is customers coming back because the answer did not land.
How the numbers mislead, specifically
| The number looks | Because | What to check |
|---|---|---|
| Response time improved sharply | A new bot reply is being counted as a response | Whether the metric separates bot from human |
| Resolution time fell | Only easy conversations are being marked resolved | The share of conversations ever resolved |
| Backlog dropped | Somebody bulk-resolved old chats | The resolve events, not the total |
| One agent handles far more | They are getting the short conversations | The mix by category |
Every row here is something that happens routinely, and none of them involve anybody behaving badly. They are ordinary consequences of measuring a system while changing it.
What to do with what you find
Two rules that keep this useful. Change one thing at a time, so the next month's numbers mean something. And treat every metric as a prompt to read conversations rather than as a verdict — the number tells you where to look, and the conversation tells you what is actually happening.
A team that reads 5 real conversations a month learns more than one that watches a dashboard daily, and the two together are better than either.
What to report upwards, and what to keep internal
| Measure | Good for a summary | Better kept for the team |
|---|---|---|
| First response time, trend | Yes | |
| Open backlog | Yes | |
| Satisfaction score | Yes | |
| Conversations per agent | Yes — it invites the wrong comparison | |
| Individual response times | Yes — useful for coaching, corrosive as a league table |
The split is not about secrecy. It is that a number shown to people who cannot see the context behind it will be read as a verdict, and the two rows in the right column are exactly the ones whose context matters most.
The one number most stores should watch
Open conversations older than a day. It is a single figure, it needs no statistical care, and it correlates with almost everything else that matters — if it is near zero your team is keeping up, and if it is growing nothing else on the dashboard is worth reading first.
Check it at the same time each morning. That is the whole practice, and it is more useful than most reporting projects.
The short version
Watch first response time for humans only, resolution time on a status system you trust, the open backlog, and how the work is spread. Take a baseline before you change anything, change one thing at a time, and read real conversations alongside the numbers.
Turning a measurement into a change
Numbers only help if something happens afterwards. Four patterns that reliably produce a change worth making:
| What you see | What it usually means | What to do |
|---|---|---|
| First response time worse on specific days | A rota gap | Move the overlap, not the people |
| The same question dominating the backlog | A missing flow | Build it — it is the cheapest capacity you can add |
| High reopen rate after resolution | Answers that do not land | Read those conversations before changing anything |
| One person consistently at capacity | Routing, not effort | Check the assignment mode and the topic routing |
Each of those is a change to a system rather than to a person, which is generally where the improvement is. A team asked to be faster gets faster for a fortnight; a flow that answers the most common question keeps working.
What not to do with the data
Do not publish individual response times as a ranking. It produces exactly the behaviour you would expect — short replies, premature resolves, and a reluctance to take the difficult conversations — and it degrades every other number on the page in the process.
Where to start if you measure nothing today
- Check open conversations older than a day, every morning, for a fortnight.
- Write the number down.
- Read the oldest one each time.
That is a complete measurement practice, it costs five minutes a day, and it will tell you more about your service operation in two weeks than a dashboard nobody has agreed the definitions for.