EMR Integration

Benchmarking Front Desk Performance Across athenahealth Locations

Written by Jo Galvez | Sep 7, 2026, 7:00:01 PM
💡Benchmarking front desk performance across athenahealth locations extends a network's benchmark culture to the teams it has never covered. The measure is comparable per-location data on message volume, response times, and resolution rates, weighted against the workload each team actually carries.

Most networks count only part of that today: the phone system counts calls, the case queue counts cases, and texts get counted nowhere, so no one holds a total.

Curogram measures all of it as conversations move through one inbox, with no time studies and no self-reporting. The villain is the unfair comparison — the location that complains loudest gets the extra hire while the team drowning in silence gets a performance conversation.

That's an evidence standard the same network would reject in any other domain.


Two clinics in the same network, same specialty, same three-person front desk.

One of them handles roughly double the combined patient contact volume of the other, with faster answers. Nobody knows which one. Ask the regional manager, and you'll get the site that has been asking for help. That answers a different question.

Front desk work is measurable in theory and measured almost nowhere. Phone systems count calls. Case queues count cases. Texts, portal messages, and reminder replies land somewhere else, or nowhere. No report puts one number on what a team handled last Tuesday.

People decisions then get made on the worst available evidence. Who complained most recently, and how well.

Our position: this is the same organization that benchmarks reimbursement rates against peers by specialty and region. It knows how to use comparative data. It applies that discipline everywhere except the one domain where the evidence is about people. Coaching without data rewards visibility over output, every time.

Below: what a full workload picture includes, how to read a response-time number fairly, four checks to run before you conclude anything, and what two weeks of real data tends to show.

The Villain: The Unfair Comparison

Three Systems, No Total

Every patient contact does get counted somewhere. That's the frustrating part.

Your phone system produces call volume by extension. The case queue reports on cases opened and closed. Reminder replies sit inside the reminder tool.

Each report is accurate on its own. None can be added to the others, because they count different things over different periods, with different ideas of what resolved means.

A manager who wants one number for one site builds it by hand from three exports. It goes stale in a week. So nobody builds it.

The Team That Absorbs It

Complaint volume is a poor stand-in for workload. It is also the one in use.

The location asking for a fourth person has a manager comfortable escalating. The location absorbing more volume has one who prides herself on handling it.

Six months on, the second team has lost two people. Exit conversations mention being stretched. That's the first real data the network gets, and it arrives as turnover.

Front-office turnover costs a network real money, and it lands on a different budget line than the hire that would have prevented it.

"Slow Afternoons" and What They Usually Are

A manager notices response times sag at one site between 2 and 4 PM.

Without volume data, that reads as a discipline problem, and the coaching conversation goes there. Volume data often tells another story. That site runs a 1:30 provider block that drops 40 check-outs and their questions into a two-hour window. The team is queued rather than slow.

Same observation, opposite conclusion. The only difference is whether anyone could see the curve underneath.

Where Impressions Bend

Missing data would matter less if what filled the gap were random.

It never is. Impressions favor proximity, so the site the manager visits most gets understood best. They favor a good talker, so the manager who frames well gets believed. They favor recency, so last week's crisis outweighs six months of steady overload.

Nobody involved is acting in bad faith. All of it still produces staffing decisions a network would refuse if the subject were claims.

The Guide: The Network Scoreboard

Counting the Whole Workload, Not the Phone Slice

One inbox produces one number.

Texting, reminders, recall replies, and campaign responses all flow through one platform, tied to the athenahealth record. That makes the workload picture complete rather than partial. Messages handled, median first-response time, resolution rate, and volume curves by hour come out per location and per queue.

Nobody fills in a timesheet. The measuring is a byproduct of the work, which is what keeps it honest.

Reading a Response-Time Number Fairly

A raw ranking invites a wrong conclusion. Four checks turn it into a diagnosis:

What you see

What it might mean

Check before concluding

Slow median, high volume

Understaffed, not underperforming

Volume per FTE against network median

Slow median, normal volume

Coaching or routing issue

Hourly curve and queue mix

Fast median, low resolution

Answering fast, solving little

Reopen rate and thread length

Sudden slowdown, one week

Campaign day or absence

Campaign calendar and coverage log

 

Run all four and the number stops reading as an accusation. Staff coaching data across a network works only if the people being coached accept the comparison. These four checks are what earn that.

Not All Message Volume Weighs the Same

Two sites can handle identical message counts and carry very different loads.

A clinic whose queue is 70% scheduling threads is doing fast, closable work. A clinic fielding billing questions and prior-auth follow-ups is doing slow work that often needs a second person. Comparing their medians without segmenting by queue produces a ranking that punishes the harder job.

Segment first, then rank. Scheduling against scheduling, billing against billing, and the comparison holds up when the manager at the bottom asks how it was calculated. That question gets asked, and a benchmark that can't answer it stops being used within a month.

Volume Per FTE Is the Argument Staffing Requests Lack

Staffing cases usually fail on evidence, not on merit.

"We're swamped" competes against every other location saying the same thing. "We handle 1.8 times the median contacts per FTE, with a median response 40% slower" is a different conversation. It cuts both ways. The same comparison sometimes shows a site is staffed fine and has a routing problem.

Message volume per location, held against headcount, is the number that settles these arguments. Most networks have never seen it, which is why staffing debates in this domain run on conviction instead.

Curogram Highlight: Team Performance Views

Team Performance Views report what each front desk actually handled. Messages handled, median first-response time, resolution rate, and volume curves by hour, broken out per location and per queue.

Views stay team-level and trend-oriented by design. The purpose is workload fairness and coaching, so the reporting shows a team's shape over time rather than tracking individuals through their day.

Context travels with the numbers — campaign days, volume spikes, and coverage gaps appear alongside the response times they explain.

Most front desks end up welcoming it, in our experience, because it's the first instrument that proves how much work they do. Data lives on a HIPAA-compliant, SOC 2 Type II certified platform, and the conversations behind the metrics stay access-controlled.

The Success: Coaching Where the Data Points

What Two Weeks of Baseline Usually Shows

Most networks expect the ranking to confirm what they already believe. It usually doesn't, in one particular way.

The site everyone assumed was struggling often sits mid-pack on volume per FTE, slow for a structural reason. A queue nobody owns. A provider block that clusters. A coverage gap at lunch. Meanwhile, a site nobody discussed turns out to be carrying the heaviest load.

Two weeks is enough to see it, because response time and volume are high-frequency numbers. Attendance metrics need a quarter before a trend can be trusted.

From "Who Seems Busy" to "Who Carries What"

Staffing and coaching decisions inherit the credibility of the evidence behind them.

That credibility is worth more than the decision itself. A manager who can show why one site got the hire can also explain it to the site that didn't, using a comparison that the team can check for itself.

Confirmation follow-through above 75% becomes a fair baseline for every location, based on our internal data. You can now see whether a site misses it, and whether volume explains the miss.

The Overloaded Team Gets Help Before the Exit Interview

Turnover is the costliest way to find a workload problem.

A team carrying 1.8 times the median gives off signals long before anyone resigns. Response times drift. Resolution rates dip. The hourly curve flattens into a standing backlog. All three are visible weekly on a scoreboard and invisible entirely without one.

Recognition matters here as much as rescue does. The strong location whose methods were invisible becomes the one other managers call. That's how one team's habits turn into everyone's baseline.

Conclusion: Be Fair By Being Informed

Front desk teams are the only group in a benchmark-native network still judged mostly on impressions. They notice, and the ones who notice most are usually the ones carrying the most.

Before the next staffing decision, get one week of real per-location volume data and compare it against contacts per FTE. It reverses the decision often enough to be worth the week.

Schedule a consultation and we'll show what two weeks of benchmark data reveals about a network your size, including the sites a ranking usually surprises people about.

 

Frequently Asked Questions