Two clinics in the same network, same specialty, same three-person front desk.
One of them handles roughly double the combined patient contact volume of the other, with faster answers. Nobody knows which one. Ask the regional manager, and you'll get the site that has been asking for help. That answers a different question.
Front desk work is measurable in theory and measured almost nowhere. Phone systems count calls. Case queues count cases. Texts, portal messages, and reminder replies land somewhere else, or nowhere. No report puts one number on what a team handled last Tuesday.
People decisions then get made on the worst available evidence. Who complained most recently, and how well.
Our position: this is the same organization that benchmarks reimbursement rates against peers by specialty and region. It knows how to use comparative data. It applies that discipline everywhere except the one domain where the evidence is about people. Coaching without data rewards visibility over output, every time.
Below: what a full workload picture includes, how to read a response-time number fairly, four checks to run before you conclude anything, and what two weeks of real data tends to show.
Every patient contact does get counted somewhere. That's the frustrating part.
Your phone system produces call volume by extension. The case queue reports on cases opened and closed. Reminder replies sit inside the reminder tool.
Each report is accurate on its own. None can be added to the others, because they count different things over different periods, with different ideas of what resolved means.
A manager who wants one number for one site builds it by hand from three exports. It goes stale in a week. So nobody builds it.
Complaint volume is a poor stand-in for workload. It is also the one in use.
The location asking for a fourth person has a manager comfortable escalating. The location absorbing more volume has one who prides herself on handling it.
Six months on, the second team has lost two people. Exit conversations mention being stretched. That's the first real data the network gets, and it arrives as turnover.
Front-office turnover costs a network real money, and it lands on a different budget line than the hire that would have prevented it.
A manager notices response times sag at one site between 2 and 4 PM.
Without volume data, that reads as a discipline problem, and the coaching conversation goes there. Volume data often tells another story. That site runs a 1:30 provider block that drops 40 check-outs and their questions into a two-hour window. The team is queued rather than slow.
Same observation, opposite conclusion. The only difference is whether anyone could see the curve underneath.
Missing data would matter less if what filled the gap were random.
It never is. Impressions favor proximity, so the site the manager visits most gets understood best. They favor a good talker, so the manager who frames well gets believed. They favor recency, so last week's crisis outweighs six months of steady overload.
Nobody involved is acting in bad faith. All of it still produces staffing decisions a network would refuse if the subject were claims.
One inbox produces one number.
Texting, reminders, recall replies, and campaign responses all flow through one platform, tied to the athenahealth record. That makes the workload picture complete rather than partial. Messages handled, median first-response time, resolution rate, and volume curves by hour come out per location and per queue.
Nobody fills in a timesheet. The measuring is a byproduct of the work, which is what keeps it honest.
A raw ranking invites a wrong conclusion. Four checks turn it into a diagnosis:
|
What you see |
What it might mean |
Check before concluding |
|---|---|---|
|
Slow median, high volume |
Understaffed, not underperforming |
Volume per FTE against network median |
|
Slow median, normal volume |
Coaching or routing issue |
Hourly curve and queue mix |
|
Fast median, low resolution |
Answering fast, solving little |
Reopen rate and thread length |
|
Sudden slowdown, one week |
Campaign day or absence |
Campaign calendar and coverage log |
Run all four and the number stops reading as an accusation. Staff coaching data across a network works only if the people being coached accept the comparison. These four checks are what earn that.
Two sites can handle identical message counts and carry very different loads.
A clinic whose queue is 70% scheduling threads is doing fast, closable work. A clinic fielding billing questions and prior-auth follow-ups is doing slow work that often needs a second person. Comparing their medians without segmenting by queue produces a ranking that punishes the harder job.
Segment first, then rank. Scheduling against scheduling, billing against billing, and the comparison holds up when the manager at the bottom asks how it was calculated. That question gets asked, and a benchmark that can't answer it stops being used within a month.
Staffing cases usually fail on evidence, not on merit.
"We're swamped" competes against every other location saying the same thing. "We handle 1.8 times the median contacts per FTE, with a median response 40% slower" is a different conversation. It cuts both ways. The same comparison sometimes shows a site is staffed fine and has a routing problem.
Message volume per location, held against headcount, is the number that settles these arguments. Most networks have never seen it, which is why staffing debates in this domain run on conviction instead.
Team Performance Views report what each front desk actually handled. Messages handled, median first-response time, resolution rate, and volume curves by hour, broken out per location and per queue.
Views stay team-level and trend-oriented by design. The purpose is workload fairness and coaching, so the reporting shows a team's shape over time rather than tracking individuals through their day.
Context travels with the numbers — campaign days, volume spikes, and coverage gaps appear alongside the response times they explain.
Most front desks end up welcoming it, in our experience, because it's the first instrument that proves how much work they do. Data lives on a HIPAA-compliant, SOC 2 Type II certified platform, and the conversations behind the metrics stay access-controlled.
Most networks expect the ranking to confirm what they already believe. It usually doesn't, in one particular way.
The site everyone assumed was struggling often sits mid-pack on volume per FTE, slow for a structural reason. A queue nobody owns. A provider block that clusters. A coverage gap at lunch. Meanwhile, a site nobody discussed turns out to be carrying the heaviest load.
Two weeks is enough to see it, because response time and volume are high-frequency numbers. Attendance metrics need a quarter before a trend can be trusted.
Staffing and coaching decisions inherit the credibility of the evidence behind them.
That credibility is worth more than the decision itself. A manager who can show why one site got the hire can also explain it to the site that didn't, using a comparison that the team can check for itself.
Confirmation follow-through above 75% becomes a fair baseline for every location, based on our internal data. You can now see whether a site misses it, and whether volume explains the miss.
Turnover is the costliest way to find a workload problem.
A team carrying 1.8 times the median gives off signals long before anyone resigns. Response times drift. Resolution rates dip. The hourly curve flattens into a standing backlog. All three are visible weekly on a scoreboard and invisible entirely without one.
Recognition matters here as much as rescue does. The strong location whose methods were invisible becomes the one other managers call. That's how one team's habits turn into everyone's baseline.
Front desk teams are the only group in a benchmark-native network still judged mostly on impressions. They notice, and the ones who notice most are usually the ones carrying the most.
Before the next staffing decision, get one week of real per-location volume data and compare it against contacts per FTE. It reverses the decision often enough to be worth the week.
Schedule a consultation and we'll show what two weeks of benchmark data reveals about a network your size, including the sites a ranking usually surprises people about.