💬 Shared Team Inbox

Measuring response time and agent workload

Most inbox dashboards flatter you. Here is how to read yours honestly.

The useful inbox measures are first response time, resolution time, open backlog and workload per person. Everything else is either derived from those or a vanity number.

The harder question is not which to watch but whether yours are counting what you think they are.

The four worth watching

MeasureTells youWatch out for
First response timeHow long a customer waits for any human answerBot replies counting as responses
Resolution timeHow long the whole thing tookConversations never marked resolved
Open backlogWhat is outstanding right nowOnly meaningful if status is used consistently
Conversations per agentWhether the load is sharedSays nothing about how hard each one was

Every one of those depends on something upstream being done properly — status, assignment, or the distinction between a bot reply and a human one.

The counting rule that decides everything

A reply that came from an automation is not a human response, and if your reporting treats it as one, your first response time will look excellent while customers wait exactly as long as before.

This is the single most common way inbox dashboards mislead. The bot answers in two seconds, the metric records two seconds, and the customer who needed a person waits an hour. Make sure your reporting separates the two, and read a few real conversations to sanity-check it.

What "resolved" has to mean

Resolution time is only meaningful if resolved means resolved. If half your conversations are never marked, your average is computed from a biased sample — the easy ones — and it will look better than reality.

Two habits fix this. Automate the obvious transitions, and reopen a conversation when the customer replies to it. The conventions are in open, pending, resolved.

Workload is not performance

Conversations per agent is genuinely useful for spotting an unbalanced rota and genuinely misleading as a performance measure. Someone handling 10 complaints has had a harder day than someone handling 40 stock questions.

If you want to compare people, compare like with like — the same category of conversation, over a long enough period to smooth out the days when the difficult ones happened to land on one person.

The measures worth ignoring

None of those are wrong to look at. They are wrong to target, because a team told to move a number will move it.

Satisfaction, measured properly

A rating question at the end of a conversation is the one measure that comes from the customer rather than from your own logs. Asked as buttons after a resolution, it is cheap and it is the only thing on this page that tells you whether the fast resolution was actually any good.

Keep it to one question, ask it only after the conversation has genuinely ended, and do not chase people who ignore it. The setup is in customer satisfaction scores from a flow.

A monthly review that takes twenty minutes

  1. First response time, humans only, compared with last month.
  2. Open backlog at the same time of day each month.
  3. Anything pending for more than a week, read individually.
  4. Workload spread across the team.
  5. Five real conversations, read end to end.

Step five is the one that finds things the other four cannot. Numbers tell you something changed; conversations tell you what.

Where cost shows up

Slow first responses have a price as well as a cost in goodwill. Replies inside the 24-hour customer service window are free and unlimited under Meta's pricing documentation, and the window opens when the customer messages you, per Meta's sending messages documentation.

Let it close and your next message needs an approved template. A team that consistently answers the same day is quietly saving money as well as customers.

Frequently asked

What is a good first response time?

Faster than you manage now. Absolute benchmarks vary so much by category that chasing somebody else's number is not useful; your own trend is.

Should I measure the bot separately?

Yes. It answers a different kind of question, and mixing the two hides both. The automation side is in measuring whether your automation works.

Why does one agent always look slower?

Frequently because they get the hard conversations. Check the mix before you draw a conclusion.

What if we have no assignment at all?

Then most of this cannot be measured, and that is the first thing to fix.

Getting a baseline before you change anything

The most common mistake is improving something before knowing what it was. Spend a month collecting numbers you have not acted on, and everything afterwards has something to be compared against.

  1. First response time, humans only, recorded weekly.
  2. Open backlog at a fixed time each day.
  3. Conversations per agent per week.
  4. The share of conversations the bot handled end to end.
  5. How many were reopened after being resolved.

The fifth is the quality check on the fourth. A high automation share with a high reopen rate is not automation working; it is customers coming back because the answer did not land.

How the numbers mislead, specifically

The number looksBecauseWhat to check
Response time improved sharplyA new bot reply is being counted as a responseWhether the metric separates bot from human
Resolution time fellOnly easy conversations are being marked resolvedThe share of conversations ever resolved
Backlog droppedSomebody bulk-resolved old chatsThe resolve events, not the total
One agent handles far moreThey are getting the short conversationsThe mix by category

Every row here is something that happens routinely, and none of them involve anybody behaving badly. They are ordinary consequences of measuring a system while changing it.

What to do with what you find

Two rules that keep this useful. Change one thing at a time, so the next month's numbers mean something. And treat every metric as a prompt to read conversations rather than as a verdict — the number tells you where to look, and the conversation tells you what is actually happening.

A team that reads 5 real conversations a month learns more than one that watches a dashboard daily, and the two together are better than either.

What to report upwards, and what to keep internal

MeasureGood for a summaryBetter kept for the team
First response time, trendYes
Open backlogYes
Satisfaction scoreYes
Conversations per agentYes — it invites the wrong comparison
Individual response timesYes — useful for coaching, corrosive as a league table

The split is not about secrecy. It is that a number shown to people who cannot see the context behind it will be read as a verdict, and the two rows in the right column are exactly the ones whose context matters most.

The one number most stores should watch

Open conversations older than a day. It is a single figure, it needs no statistical care, and it correlates with almost everything else that matters — if it is near zero your team is keeping up, and if it is growing nothing else on the dashboard is worth reading first.

Check it at the same time each morning. That is the whole practice, and it is more useful than most reporting projects.

The short version

Watch first response time for humans only, resolution time on a status system you trust, the open backlog, and how the work is spread. Take a baseline before you change anything, change one thing at a time, and read real conversations alongside the numbers.

Turning a measurement into a change

Numbers only help if something happens afterwards. Four patterns that reliably produce a change worth making:

What you seeWhat it usually meansWhat to do
First response time worse on specific daysA rota gapMove the overlap, not the people
The same question dominating the backlogA missing flowBuild it — it is the cheapest capacity you can add
High reopen rate after resolutionAnswers that do not landRead those conversations before changing anything
One person consistently at capacityRouting, not effortCheck the assignment mode and the topic routing

Each of those is a change to a system rather than to a person, which is generally where the improvement is. A team asked to be faster gets faster for a fortnight; a flow that answers the most common question keeps working.

What not to do with the data

Do not publish individual response times as a ranking. It produces exactly the behaviour you would expect — short replies, premature resolves, and a reluctance to take the difficult conversations — and it degrades every other number on the page in the process.

Where to start if you measure nothing today

  1. Check open conversations older than a day, every morning, for a fortnight.
  2. Write the number down.
  3. Read the oldest one each time.

That is a complete measurement practice, it costs five minutes a day, and it will tell you more about your service operation in two weeks than a dashboard nobody has agreed the definitions for.

Need help? Message us