Support metrics go wrong in two directions. Some teams track nothing and run on vibes until a bad month arrives with no explanation. Others copy an enterprise dashboard — twenty KPIs, none owned by anyone — and stop looking after a fortnight. For a team of two to ten agents handling WhatsApp, Instagram, Telegram and email from a shared inbox, five metrics are enough to answer the three questions that matter: are we fast, are we good, and are we drowning. This is the overview of all five; where a metric deserves a deep dive of its own, we link to it.
Why These Five and Not Twenty
Each metric earns its place by covering a failure mode the others miss. First response time catches slowness the customer feels immediately. Resolution time catches conversations that start fast and then drag. CSAT catches interactions that were quick but bad. Contact rate catches problems upstream of support — the tickets you should never have received. Backlog catches the operational debt that predicts next week's numbers. Drop any one and a whole class of problem becomes invisible; add more and the weekly review stops happening. A useful discipline: every metric on your scorecard must have a named owner and a known action that moves it. If nobody can say what they would do differently if the number worsened, it is decoration, not measurement.
1. First Response Time (FRT)
FRT is the elapsed time between a customer's first message in a new conversation and the first human reply. Two computation rules keep it honest. Use the median, not the mean — one conversation that sat overnight will wreck an average and tell you nothing. And compute it against business hours if you publish them: a WhatsApp message at 2am answered at 8:05am is a five-minute business-hours response, not a six-hour failure. Keep a separate eye on the 90th percentile, which is where customers actually feel pain.
Auto-replies do not count as a first response. An acknowledgement message has value — it sets expectations — but the clock should keep running until a human (or a bot genuinely resolving the issue) engages. Expectations differ sharply by channel: messaging users expect an answer in minutes, email users in hours. Channel-by-channel targets, percentile maths and how to bring FRT down without hiring are covered in depth in our first response time benchmarks for 2026.
As broad orientation for a small team during business hours: under 15 minutes on WhatsApp and live chat is strong, under an hour on Instagram and Telegram DMs is acceptable, and under four hours on email keeps you comfortably ahead of most inboxes.
2. Resolution Time
Resolution time is the elapsed time from the customer's first message to the conversation being genuinely closed. Track two variants. Full resolution time (median, business hours) shows how long problems take end to end. First contact resolution (FCR) — the percentage of conversations resolved without the customer needing to return — shows how often you fix things properly the first time. A team can have excellent FRT and terrible resolution time; that pattern almost always means fast acknowledgements followed by slow internal answers, and it points at knowledge gaps or missing escalation paths rather than lazy agents.
Watch reopens. If a conversation marked resolved is reopened by the customer within 48–72 hours, count it against FCR. Reopen rates creeping above roughly one in ten resolved conversations usually mean agents are closing threads to make queues look clean — a metric-gaming problem, not a customer problem. Segment resolution time by topic label: refunds legitimately take longer than password resets, and a single blended number hides the topic that is actually broken.
3. CSAT
CSAT asks the customer directly: how was this interaction? The standard mechanics — a one-question, five-point survey sent in-thread shortly after resolution — take an afternoon to set up in a shared inbox, and the score is the percentage of responses rating 4 or 5. For small teams on messaging channels, 75–85% is the common healthy range, with sustained scores above 90% genuinely strong.
The number is only the start. Response rates, survey timing inside WhatsApp's 24-hour window, segmentation by agent, channel and topic, minimum sample sizes, and the closed-loop follow-up that every 1-star rating deserves are a discipline of their own — we cover the complete system in how to collect, analyse and act on CSAT scores. For the weekly scorecard, two numbers suffice: the rolling 30-day score, and the count of low ratings that have not yet had a human follow-up. The second should be zero.
4. Contact Rate
Contact rate is the metric most small teams have never computed, and the one founders should care about most. It is support volume normalised by business activity: conversations per 100 orders for an e-commerce operation, or conversations per 100 active customers per month for a subscription product. Raw ticket counts rise when the business grows; contact rate rises only when something is generating avoidable work.
Its power is diagnostic. Segment contacts by topic and ask, for each of the top five topics: should this contact exist? "Where is my order?" messages signal a tracking-communication gap, not a support problem. "How do I…" questions signal missing or unfindable self-service content. Payment-failure messages signal a checkout bug. Cutting contact rate is mostly deflection work — better order notifications, an answer library, AI auto-replies for the questions that repeat daily — and the playbook is laid out in our guide to AI ticket deflection for small support teams. A falling contact rate with stable CSAT is the single clearest sign a support operation is maturing: the team handles growth without headcount by deleting work instead of absorbing it. Review it monthly rather than weekly — the denominator needs a full cycle of orders or billing to be meaningful, and week-to-week wobble in this metric is almost always noise.
5. Backlog and Oldest Open Conversation
Backlog is the count of open, unresolved conversations at a point in time — snapshot it at the same hour daily. On its own the count is ambiguous (twenty open chats mid-morning is normal traffic), so pair it with age: the age of the oldest unanswered conversation, and the number of conversations open beyond 24 and 72 hours. Volume spikes are weather; rising age is rot.
Backlog is the leading indicator in the set. FRT and CSAT report last week; a backlog trending upward for four consecutive days predicts next week's misses before they happen. The fixes are operational: a morning triage sweep, routing rules so nothing sits unassigned, and labels that separate "waiting on customer" from "waiting on us" so the queue reflects real debt. Unassigned conversations are the biggest silent killer in shared inboxes — if that is your failure mode, start with automatic chat assignment and the backlog age problem usually halves on its own.
The Weekly Scorecard
Put the five on one page and review them at the same time every week — twenty minutes, whole team present. Typical healthy ranges for a 2–10 person messaging-first team:
| Metric | How to compute | Healthy range (small team) | First move if it slips |
|---|---|---|---|
| First response time | Median, business hours, per channel | <15 min WhatsApp; <1 h DMs; <4 h email | Check unassigned queue and coverage gaps |
| Resolution time / FCR | Median close time; % resolved without return | Same-day close common; FCR 70–80% | Segment by topic; fix the slowest topic |
| CSAT | % of responses rating 4–5, rolling 30 days | 75–85% | Read verbatims; follow up every low score |
| Contact rate | Conversations per 100 orders or active users | Flat or falling month over month | Deflect the top repeating topic |
| Backlog / oldest open | Daily snapshot; age of oldest unanswered | Nothing unanswered >24 business hours | Morning sweep; auto-assignment rules |
Run the review to a fixed agenda so it survives busy weeks: two minutes per metric against last week and the four-week trend, five minutes on the single worst number, and a closing decision — one change to make this week, with an owner. Skip the temptation to discuss everything; the meeting's job is to produce one action, not a retrospective. Write the action in the same document as the scorecard so next week's meeting opens by checking whether it happened and whether it moved the number. Over a quarter, that log of change-and-effect becomes the most honest record you have of what actually improves your operation — and it is what makes the difference between a team that measures and a team that improves.
When several metrics slip at once, fix them in this order: backlog first (it is upstream of everything), then FRT, then resolution, then CSAT, then contact rate. Speed problems are staffing and routing problems and respond to changes within days; satisfaction and deflection are slower levers.
Metrics You Can Safely Ignore for Now
Knowing what not to track is half the discipline. Four metrics that small teams borrow from enterprise dashboards and then regret:
- Average handle time (AHT). On messaging channels, conversations are asynchronous by nature — a WhatsApp thread that spans forty minutes of wall-clock time may contain ninety seconds of agent work. Optimising AHT on chat pushes agents to rush, and rushing shows up in CSAT within weeks.
- Tickets closed per agent. The fastest way to close many tickets is to close them badly. Reopens and repeat contacts absorb the saving. Per-agent volume belongs in workload balancing, not in a quality scorecard.
- NPS. A relationship metric driven by product, pricing and brand as much as by support. Useful at company level eventually; noise at support-team level now.
- Channel-level CSAT differences of a point or two. Different channels attract different populations; small gaps are demographics, not performance. Investigate gaps of five points or more, sustained.
The pattern behind all four: they measure activity or sentiment you cannot directly act on. The five on the scorecard each have a lever attached — that is the admission test.
Setting Targets When You Have No History
A team instrumenting for the first time faces a chicken-and-egg problem: targets should come from data, but there is no data yet. The sequence that works: run two weeks with no targets at all, purely collecting. Take the medians from those weeks as your provisional targets — not the benchmarks in this article, your own numbers, however unflattering. Improve towards the benchmark ranges over the following quarter by fixing the biggest gap first. Teams that adopt external benchmarks on day one set themselves up to miss everything simultaneously, conclude that measurement is demoralising, and stop. Teams that chase their own last-month numbers improve steadily and hit the published ranges within a quarter or two without the morale damage.
Instrumenting This in a Shared Inbox
None of this works if the data lives in four separate apps — a personal WhatsApp on someone's phone, an Instagram login shared in a spreadsheet, a Gmail account. The precondition for measurement is that every channel flows into one place with timestamps, assignees and labels attached. In OmniDesk, the analytics dashboard breaks response time, resolution rate and agent workload down by inbox, channel and time period, so the scorecard above is a report you open rather than a spreadsheet you maintain. Labels drive the topic cuts, routing rules keep the unassigned queue at zero, and the same data feeds the SLA targets you will eventually want to formalise — when you are ready for that step, our guide to building SLA rules your team will actually follow turns these benchmarks into commitments.
Frequently Asked Questions
Which single metric should a brand-new team start with?
First response time. It is the number customers feel most directly, it is fully within the team's control, and instrumenting it forces the operational hygiene — one inbox, assignment, business hours — that every other metric depends on. Add CSAT in month two and the rest as volume grows.
Should we use averages or medians?
Medians for anything time-based. Support time data is heavily skewed — a handful of overnight or week-long conversations drag an average far from typical experience. Use the median for the headline and the 90th percentile to see the worst cases; if those two diverge sharply, your problem is consistency rather than speed.
Do bot and auto-reply conversations count in these metrics?
Count a bot response as first response only if the bot actually resolved the issue — an auto-acknowledgement does not stop the FRT clock. Fully bot-resolved conversations should appear in resolution and CSAT metrics as their own segment, so automation quality is visible rather than blended into human performance.
Is this overkill for a two-person team?
Track three: FRT, CSAT and backlog age. Contact rate starts mattering once volume grows faster than the team; formal resolution-time tracking matters once conversations regularly span days. The weekly-review habit matters more than the metric count — two people looking at three numbers beats ten people ignoring twenty.
How do these metrics relate to SLAs?
Metrics describe what happens; an SLA is a promise about what should happen, with an alert when it is about to be broken. Run the scorecard for a month or two first so targets come from your own data, then formalise the response-time promises as SLA rules with escalation paths.