Every company selling prediction in the perioperative space eventually shows you the same slide. A line chart with two lines that hug each other, and a number underneath. 94% accurate. 96% accurate. Sometimes a MAPE figure, which most people nod at and nobody asks to see the math on.
It’s a fine number. It’s just not the number that decides whether next Tuesday goes well.
We’ve spent a lot of time inside OR and anesthesia departments, and we’ve never once heard a scheduler say the forecast was wrong. What we hear is that it didn’t tell them anything they could act on, or it told them too late to matter, or it told them they were short and left them to figure out the rest with a phone and a spreadsheet.
If you’re getting ready to evaluate predictive staffing tools, accuracy is the easiest thing to compare and the least useful thing to compare on. Here’s what we’d look at instead.
Accurate about what, exactly?
Start with the thing being predicted. Most tools in this category forecast case volume, block utilization, or room minutes. That’s real work and it’s genuinely hard. It’s also one step removed from the decision you actually have to make.
Knowing you’ll run 42 cases two Tuesdays from now doesn’t tell you whether you need another CRNA. Knowing you’ll run 310 room minutes in ortho doesn’t tell you whether your anesthesiologist count works. The forecast has to survive a translation before anyone can use it, and that translation is where the difficulty lives.
Volume forecasting was built for the OR’s operational view. Which rooms open, which blocks get released, how the day flows. Anesthesia staffing sits on top of the same data and doesn’t answer to the same math.
Coverage models are where the forecast becomes real
Here’s the part that gets waved past in most demos.
Anesthesia coverage isn’t cases divided by a ratio. You’ve got a care team model where one physician supervises two, three, or four rooms depending on acuity, location, and what your state and payer rules allow. You’ve got MD-only rooms. You’ve got sites that are physically far enough apart that supervision doesn’t stretch. You’ve got credentialing that says this person can do hearts and that person can’t. You’ve got call, late rooms, first-start requirements, break and relief coverage, and a post-call day that takes someone off the board entirely.
You also have local rules that exist nowhere in writing. The Thursday room that always runs long. The surgeon whose posted times are optimistic by 40%. The site that needs its own dedicated person no matter how light the day looks.
A tool that forecasts volume accurately and then applies a flat ratio to it will be confidently wrong in a specific way. It tells you that you’re fine on days you aren’t, because it doesn’t know that two of your four rooms can’t be supervised together. The accuracy number stays high and the recommendation is still useless.
So the question isn’t “how accurate is the forecast.” It’s “does this understand how we cover rooms, and can it tell me what to change.” A forecast that says you’re two short on the 14th and here’s the shift that closes it beats a more accurate forecast that stops at a number.
Doing this by hand is possible, and it never stops
Plenty of groups already do a version of this. Someone senior, usually a scheduler who has been there a decade or an anesthesiologist with administrative time, works the grid. Pulls the schedule, cross-references call and time off, and forms an opinion about where the trouble is.
That opinion is often good. The person is often right. The problem is that this isn’t a monthly exercise, it’s a daily one. The grid moves constantly. Add-ons, releases, call-outs, a surgeon who blocked four rooms and filled two. Every one of those changes the answer, so the analysis has to be redone, and it gets redone in someone’s head in between everything else they’re doing.
Add that up across a department and it’s a real cost, spread thin enough that nobody counts it. What you get for it is a view that’s already slightly stale, can’t be checked by anyone else, and walks out the door when that person does. So the comparison isn’t software versus nothing. It’s software versus a smart person re-deriving the same answer every morning from data that changed overnight.
The benchmark isn’t perfection, it’s your block target
This is the one we’d argue hardest for.
Most departments staff to block. It’s the obvious starting point and it isn’t a dumb one, it’s just wrong most of the time. Blocks get released late. Cases get added. Posted times drift. The grid you staffed to last week is not the day you run. Manage to block target and you’ll be off on most days, some of them badly.
So the bar a forecast has to clear isn’t perfection. It’s block target. A number that’s meaningfully better than block, far enough out that you can still act on it, changes how the department runs.
And that’s the second half of the question. Better than block is only worth something if it arrives while you still have options.
Two weeks out is a scheduling question, two days out is a phone call
We start forecasting around 60 days out, and the long view has real value. Time off approvals, locums lead times, spotting a heavy stretch before it gets baked in. But 60 days is planning. The window where the forecast is sharp enough to trust and there’s still room to move is roughly 7 to 14 days, and that’s where the work gets done.
At two weeks out, a gap is a scheduling question. You can move a shift, adjust a call assignment, ask someone if they’d rather work the 14th than the 16th. Most of those conversations are neutral, and some are genuinely welcome. People have preferences. A swap that lands well is a favor, not an imposition.
At two days out, that same gap is a phone call. Somebody has plans. You’re asking them to break them. They might say yes, and plenty of people will, but you spend something every time you ask. Goodwill, incentive pay, or both. You also have fewer levers by then. Time off is approved, people have made arrangements, and the option that’s left is usually the most expensive one.
Those asks don’t get spread evenly either. Under pressure you call the people who say yes, which is rational on the day and corrosive over a year. The same three or four reliable people end up absorbing most of the extra work. Lead time doesn’t fix that by itself, but it widens the pool of people who could plausibly say yes, and a wider pool is most of the fix.
So when you evaluate predictive staffing tools, don’t take the headline accuracy figure at face value. Ask two things. What’s the accuracy at 7 to 14 days, and what is it being compared against? If a tool is 90% accurate and your block target is already 90% accurate, you bought a dashboard. The gap between the forecast and the naive baseline is the entire product.
Knowing you’re short is not the same as being covered
Say the forecast works. It understands your coverage model, it runs far enough ahead, and it tells you the week of the 14th is light by two people.
Now what?
Somebody still has to figure out who could work. That means knowing who’s already scheduled, who’s credentialed for the location and case mix, who’s near an hours or fatigue limit, who’s on post-call, and who has historically said yes to this kind of ask. Then somebody has to make contact, track responses, and update the schedule when someone accepts.
That work is not a footnote. In most departments it’s the majority of the effort. The forecast is maybe 20% of the job, and it’s the 20% that vendors talk about because it’s the part that demos well.
If a tool hands you a gap and walks away, you’ve automated the easy part. Ask what happens after the gap appears. Does it produce a list of eligible people or just a number? Can you send the ask from inside the system? Does it track who was contacted and what they said? The more of that loop that lives in one place, the less of it lives in somebody’s inbox.
What we’d actually ask a vendor
If we were on the buying side, this is the list. It applies to us as much as anyone else.
- What are you forecasting? Cases, minutes, or staffed positions? If it’s the first two, who does the translation to people?
- Show me accuracy at 7 to 14 days, and show me how much better it is than staffing to block target.
- Does the model know our coverage ratios, our sites, our credentialing rules, and our call structure? Or does it apply a generic ratio?
- What does the output look like? A number, or a specific recommended change to a specific day?
- What happens after a gap is identified? Walk me through the next four clicks.
- Who does it suggest we ask, and on what basis?
- Where does the data come from, and what breaks when our block grid changes?
- When the forecast is wrong, how do we correct it? Is there a way to tune it against what actually happened?
- How long until this is live and producing something we’d act on?
That last one matters more than people expect. A tool that takes nine months to configure has already cost you two scheduling cycles.
The scoreboard that matters
Accuracy is table stakes. Every serious tool in this space predicts volume reasonably well, and the differences between them are smaller than the marketing suggests.
The real questions are how much better it is than what you’re doing today, how early you find out, whether it understands how your department covers rooms, and how much of the follow-through it handles. Those are harder to put on a chart, which is exactly why they don’t show up on one.
The number we’d track a year in isn’t forecast accuracy. It’s how many times somebody had to make the two-days-out phone call, and how many of those turned into a premium shift. That number going down is what the whole thing is for.
This article also appears on the ORlogic Substack.
See what your OR forecast looks like 60 days out
Book a working session with our team and we'll walk through your staffing picture using your own data.
Book a time