Measuring whether your WhatsApp automation actually works
A channel can report a lot of recovered carts while claiming orders that would have arrived anyway.
Measuring an automation means asking whether it changed an outcome you care about — support load, conversion, or repeat purchase — rather than counting how many messages it sent. Message volume is the easiest number to produce and the least informative one available.
This is the honest difficulty with WhatsApp as a channel. A customer asks a question, disappears for two days, then buys through your website. Email would show you a click. WhatsApp shows you a conversation and leaves the causal question open.
The four numbers worth tracking
| Number | What it tells you | Where it comes from |
|---|---|---|
| Conversations resolved without a human | Whether the bot is actually absorbing load | Flows that completed without an Assign Chat step |
| Time to first human reply | Whether service improved | Analytics |
| Reply rate on transactional messages | Whether people still read you | Analytics |
| Cost per outcome | Whether it pays | Message charges ÷ orders confirmed, carts recovered, etc. |
The third one is the leading indicator nobody watches. When reply rates on order messages start falling, you are being tuned out — which shows up as a quality rating downgrade weeks later.
The attribution trap
Any recovery flow will report impressive numbers, because a share of the people it messages would have come back regardless. A tool that claims every subsequent purchase is measuring your customers' intentions, not its own effect.
The honest way to find out is a holdout: for a period, deliberately exclude a random slice of eligible customers from the flow and compare. If the flow is working, the messaged group converts meaningfully better. If it is not, you have learned something expensive for free.
Be careful with published benchmarks. Widely-repeated figures like "98% open rates" or specific cart-recovery percentages generally trace to vendor marketing rather than to a study with a method. We do not publish them here for that reason, and you should treat them the same way when comparing tools.
What the platform can and cannot tell you
- Delivered and read are available per message, though read counts undercount because recipients can disable read receipts.
- Link clicks are a raw count. There is no click-through rate, and you cannot build an audience of people who clicked.
- Automation activity shows every run and where it stopped — the best debugging tool you have.
- Revenue is not attributed automatically. WhatsApp traffic arriving at your website generally lands in "direct" unless you tag your links.
That last point is fixable and worth twenty minutes: add UTM parameters to the links your flows send, so your store analytics can separate WhatsApp-driven sessions from everything else.
Judging a flow by its job
Different flows deserve different questions.
- Order confirmation: what share of orders get confirmed, and how many cancellations arrive before you paid to ship? That second number is the whole value in cash-on-delivery markets.
- "Where is my order?": what share of order-status questions finished without a human? Compare against your inbox volume before you launched it.
- Abandoned checkout: holdout test, or you are guessing.
- Review requests: reviews received per hundred messages sent.
- Ad lead qualification: cost per qualified lead, not per conversation started.
The cost side
Cost per outcome needs both halves of your bill: the subscription and Meta's per-message charges. Since messages sent inside an open conversation window are free under Meta's pricing rules, a well-designed flow can have a near-zero marginal cost — which makes "cost per outcome" a genuinely useful comparison between flows rather than a formality.
It also shows you where money is actually going. Usually one marketing flow accounts for most of the message spend while the utility flows run for nothing.
A monthly review that takes twenty minutes
- Check quality rating. Green?
- Check each automation's activity view. Any flow with zero runs is broken or unnecessary — both worth knowing.
- Read the unanswered branch of your AI step. Every question there is a gap.
- Check opt-out volume against list growth.
- Pick the flow with the worst cost per outcome and either fix or retire it.
Point two catches more problems than anything else. A keyword flow that stopped matching because customers changed how they phrase things does not error — it simply stops running, silently, and keyword lists are the usual culprit.
The uncomfortable question
For each flow, ask what would happen if you switched it off for two weeks. If you cannot name the number that would move, you are not measuring it — and a flow nobody can defend is a flow adding to your frequency budget for nothing.
Frequently asked questions
How do I know if WhatsApp is actually driving sales?
Run a holdout: exclude a random slice of eligible customers from a flow for a period and compare conversion. Without that, a recovery flow claims orders that would have arrived anyway.
Can I see click-through rates on WhatsApp links?
No. Link tracking gives a raw click count only — there is no click-through rate and you cannot build an audience of people who clicked. Add UTM parameters so your store analytics can attribute the sessions.
Why does read count look lower than delivered?
Because recipients can disable read receipts, so read is structurally undercounted. Treat delivered as the reliable number.
What is the best early warning that an automation is failing?
A falling reply rate on transactional messages, and any automation showing zero runs in its activity view — usually a keyword list that no longer matches how customers phrase things.