Contact Us

Forecasting and Statistical Modeling

Accurate at the Horizon Where the
Decision Is Actually Made

Forecasting and statistical models built around a specific planning decision, benchmarked against the naive baseline, delivered as a range rather than a point, and monitored for the decay that follows every model into production.
CaliberFocus builds demand, capacity, workforce, financial and risk models for health systems, hospitals, physician groups and ambulatory organizations. We start from the decision and its lead time, because a model that is highly accurate one week out is worthless if the staffing commitment is made six weeks out. And we publish the comparison against a simple baseline every time, because a meaningful proportion of forecasting projects never beat one.
If a model cannot beat last year same week at the horizon you plan on, it should not be deployed. We will tell you when that is the case.
The Challenge

Most provider planning runs on last year plus a percentage

Budgets are built from prior year actuals with a growth assumption. Staffing is planned from a historical average. Capacity decisions are made from a trend line drawn by eye. This is not incompetence, it is a rational response to forecasts that arrived late, could not be explained, or turned out to be no better than the assumption they replaced.

The failure is usually not the mathematics. It is that the model was built to a horizon nobody plans on, delivered a single number with implied precision it did not have, could not explain itself to the manager expected to act on it, and quietly stopped working eight months later without anyone noticing.

Accuracy optimized at the wrong horizon

A staffing model measured one week out tells you nothing about a six-week commitment.

Point estimates where a range is needed

Capacity and staffing are often planned against the upper end of plausible demand, not the average.

Never benchmarked against something simple

Seasonal naive methods are surprisingly strong in healthcare. A meaningful share of models do not beat them.

History with structural breaks in it

Pandemic years, service openings, provider departures, EHR conversions and payer changes create regime shifts.

Models that cannot explain themselves

A department manager will not staff to a number they do not understand. Explainability is an adoption requirement.

Multiple versions of the future, built on multiple versions of the past

Finance, workforce, operations and service lines often plan from different history windows, assumptions and definitions.

Decay that nobody is watching

Referral patterns move, providers leave and payer rules change while the model keeps producing confident output.
A forecast with no decision attached is a number, not a forecast.
Before modeling begins we establish who acts on the output, what they decide, when the commitment is made and what would change if the number were different. If nobody can answer those four questions, the correct recommendation is not to build the model.
Our Approach

Baseline first. Horizon second. model third

The order matters. Establishing what a simple method achieves at the required horizon sets the bar the model has to clear and frequently determines whether the project should proceed at all. Only then is it worth deciding what to build.

Step 1

Name the decision and the lead time

Who acts, what they commit to, how far ahead the commitment is made, and what would change at a different number.

You get a required forecast horizon set by the decision.

Step 2

Establish the naive baseline

Measure what last year same period, a moving average and a seasonal naive method achieve at that exact horizon.
You get the bar any model must clear.

Step 3

Audit the history

Identify structural breaks, regime changes, data quality shifts and one-off events, and decide how each is handled.
You get a training window that reflects the world the forecast applies to.

Step 4

Engineer features

Build predictors from the governed platform. Seasonality, calendar, capacity, referral pipeline, and external signals where they genuinely help.

You get a reproducible feature set with lineage.

Step 5

Start simple and escalate only if it pays

Statistical methods first. More complex approaches only where they measurably improve accuracy at the required horizon.
You get the simplest model that clears the bar.

Step 6

Backtest at the decision horizon

Walk-forward validation out of time, evaluated at the lead time that matters rather than at whatever horizon flatters the model.
You get an honest estimate of the accuracy the planner will experience.

Step 7

Quantify uncertainty

Prediction intervals with stated confidence, calibrated so an eighty percent interval contains the outcome about eighty percent of the time.
You get a range a planner can size capacity against.

step 8

Deploy into the decision

Deliver into the planning meeting, the scheduling process or the operational system on the cadence the decision runs on.
You get output before the commitment, not after it.

Step 9

Monitor, revalidate, retire

Track accuracy at horizon, calibration and feature drift, revalidate on a schedule, and retire the model when it stops earning its place.
You get a model with a lifecycle rather than a launch.
Property Last Year Plus a Percentage Spreadsheet Trend Vendor Black Box Governed Forecasting
Accurate at the decision horizon Untested Untested Claimed Measured and published
Compared to a naive baseline It is the baseline No Rarely Always
Uncertainty quantified No No Sometimes Calibrated intervals
Explainable to the operator Yes Yes No Yes, required
Handles structural breaks Poorly Poorly Unknown Explicitly
Monitored for decay Not applicable No Vendor dependent Yes, with revalidation
Retired when it stops working Not applicable No No Yes, by design
Sometimes the honest answer is that the baseline wins.
We have finished assessments by recommending that a client keep planning on a seasonal average and spend the money on data quality or process instead. That is a real outcome and it saves considerably more than a marginal model would have earned. A partner who has never delivered that recommendation has probably not been measuring properly.
Capabilities

The Model Is a Quarter of the Work

Fitting a model is the fastest part of any forecasting engagement. Framing the decision, cleaning a history that contains several different operating regimes, validating honestly at the right horizon, and getting the output into a planning process people already run take most of the time and determine whether any of it is used.

Frame and Prepare

Decision and Horizon Framing

Establishing the decision, the decision maker, the commitment lead time and the value of being right, which together define what the model has to achieve to be worth building.

Baseline Establishment

Measuring what simple methods achieve at the required horizon before any modeling starts, so improvement is demonstrated rather than assumed.

Historical Audit and Regime Detection

Identifying structural breaks, one-off events, capacity changes and data quality shifts in the history, and deciding explicitly how each is treated in training.

Feature Engineering and External Signals

Predictors built from the governed platform, plus calendar, seasonality, demographic and other external signals where they measurably improve accuracy rather than because they are available.

Model and Validate

Time Series Forecasting

Volume, census, arrivals, cash and utilization forecasting using statistical and machine learning methods selected on measured performance at the required horizon rather than on novelty.

Hierarchical and Reconciled Forecasting

Department, service line, facility and enterprise forecasts that reconcile to each other. Without reconciliation, the sum of the department plans disagrees with the enterprise plan and finance stops using both.

Classification and Risk Models

Propensity and risk estimation for no-show, denial likelihood, readmission, deterioration and similar outcomes, with calibration treated as a first class requirement alongside discrimination.

Uncertainty Quantification

Calibrated prediction intervals and scenario ranges, so planners can size against the upper end of plausible demand rather than the average.

Scenario and Sensitivity Analysis

Modeling of capacity, staffing, volume and service line changes with assumptions stated openly, and presented as a range of outcomes rather than as a prediction.

Model Selection on More Than Accuracy

Candidate models compared on stability, interpretability, maintainability, data dependency, runtime and monitoring burden alongside accuracy. The winning model is the one that best supports the decision over three years, not the one that scored highest in week six.

Deploy and Sustain

Decision Workflow Integration

Output delivered into the planning meeting, scheduling system, staffing tool or operational queue on the cadence the decision runs on, rather than into a dashboard someone has to remember to open.

Explanation and Adoption

Drivers of each forecast surfaced in operational language, because a manager who cannot see why the number moved will revert to their own judgment and the model will quietly stop being used.

Performance Monitoring and Revalidation

Accuracy at horizon, calibration, feature and population drift tracked continuously, with a defined revalidation cadence and a retirement path.

Fairness and Clinical Safety Assessment

Subgroup performance evaluation for any model that ranks, prioritizes or targets patients, conducted before deployment and repeated on the revalidation cycle.

What CaliberFocus does, and does not do.
We do not sell a forecasting product and we do not have a model looking for a use case. We frame the decision, establish the bar, build the simplest thing that clears it, prove it out of time at the horizon that matters, and put it where the decision is made. We also build the data foundation underneath it, which means we cannot blame the data for a model that does not work.
The Domains

The horizon column is the one that decides feasibility

Every organization asks for the same forecasts. What separates the ones that get used is whether the model is accurate at the lead time the commitment is actually made.
Domain What Is Forecast Decision Horizon Who Acts
Patient demand Visit and referral volume by specialty, site and payer Weeks to quarters Access, capacity planning, recruitment
Inpatient census Occupancy, admissions, discharges, length of stay distribution Hours to days for flow, weeks for staffing Bed management, nursing leadership
Emergency arrivals Arrival volume by hour, day and acuity Days to weeks ED staffing and scheduling
Perioperative demand Case volume, block utilization, case duration Weeks to months Block allocation, staffing, capital planning
Staffing requirement Worked hours needed by unit, shift and skill mix Four to eight weeks, set by the schedule cycle Nursing and workforce leadership
No-show and cancellation Probability by patient, appointment type and lead time Days to weeks Scheduling, overbooking policy, outreach
Cash and revenue Collections timing, net revenue, payment lag by payer Months to a year Finance, treasury, budgeting
Denial likelihood Probability of denial at or before claim submission Pre-submission, in workflow Revenue cycle, coding, prior authorization
Supply and pharmacy demand Consumption by item, procedure and location Weeks to months Supply chain, pharmacy, purchasing
Workforce attrition Turnover and vacancy risk by unit and role One to two quarters HR, nursing leadership, recruitment
Readmission and deterioration risk Patient level risk scores In-encounter to 30 days Care management and clinical teams
Population and risk adjustment Cost, utilization and risk trajectory for attributed populations Quarters to a year Value based care, actuarial, contracting
Validation

Point-in-time correctness is where most healthcare models silently fail

The most common serious defect we find in provider models is not a modeling error. It is training on the corrected record. A lab result was amended, a diagnosis was added at coding, an encounter was reclassified after discharge, a charge posted three weeks late. If the model trains on the final version of history, it learns from information that did not exist at the moment it will be asked to predict. Backtests look excellent and production performance collapses.

Point-in-time feature construction

Every feature reconstructed as it was known at the prediction moment, not as the record was eventually corrected. This is unglamorous work and it is the difference between a model that works and one that only backtests well.

Out-of-time walk-forward validation

Trained on the past, tested on a future the model has not seen, rolled forward repeatedly. Random splits are not valid for time dependent healthcare data and they routinely overstate accuracy.

Evaluated at the decision horizon

Performance reported at the lead time the planner commits on, alongside the naive baseline at that same horizon.

Calibration as well as discrimination

For risk models, whether a stated seventy percent probability actually occurs about seventy percent of the time. A well ranked but poorly calibrated model produces confident and wrong resource allocation.

Hierarchical reconciliation

Unit, department, facility and enterprise forecasts made to sum consistently, so operational plans and the financial plan are not built on different numbers

Subgroup performance

For any model affecting patients, accuracy and calibration examined across relevant populations before deployment, not after a concern is raised.
Bias is a separate failure from error
A model that consistently under-forecasts demand by four percent has an acceptable average error and produces chronic understaffing every single period. Directional bias is tracked separately from magnitude of error, and it is tracked by segment, because a single enterprise accuracy figure routinely hides poor performance in the exact units where the forecast is needed. Error is reported by horizon, by facility, by service line and by the operating segment that will act on it.

PHASE 1

Beats the naive baseline

At the decision horizon.

PHASE 2

Uncertainty is calibrated

The range performs as stated.

PHASE 3

Drivers are
explainable

To the person who will act.

PHASE 4

Owner and
revalidation date

Named before production.

PHASE 4

Completed a
shadow cycle

Compared with the existing planning method against real outcomes.
Governance

Every model is decaying from the day it ships

Healthcare models degrade faster than in most industries because the underlying system keeps changing. A service line opens, three physicians leave, a payer changes its authorization rules, an EHR build alters how something is captured. None of that announces itself to the model, which continues producing confident output derived from a world that has moved.
Signal What It Detects Typical Response
Accuracy at horizon The model is losing skill against actuals at the lead time that matters Investigate, then retrain or retire
Skill versus baseline The naive method has caught up or overtaken the model Retire. A model that no longer beats the baseline is a maintenance cost
Calibration drift Stated probabilities no longer match observed rates Recalibrate before anyone sizes capacity on the interval again
Feature and population drift Inputs or the population have shifted from the training distribution Assess whether a regime change has occurred and retrain on the new regime
Subgroup performance Accuracy diverging across patient or operational groups Escalate to clinical governance for any patient-facing model
Intervention effect The forecast changed behavior, which changed the outcome it is trained on Model the intervention explicitly rather than letting it contaminate training

Decide What Bad Looks Like Before It Happens

Acceptable performance, warning threshold, failure threshold, escalation owner and required response are agreed before deployment and written into the model register.

Retraining Is a Controlled Process, Not a Reflex

A trigger produces a candidate model, which is validated out of time, compared with production, approved, versioned and kept rollback-ready.

The Feedback Problem Nobody Mentions

A successful forecast changes behavior, which changes the outcome it later trains on. Staffing, capacity and no-show models can contaminate their own history unless intervention effects are handled deliberately.

A Model Register

Every model in production with its purpose, owner, risk tier, training window, performance at deployment, current performance, revalidation date and retirement criteria. This is the artifact regulators and insurers will eventually ask for.

Four Governance Tiers

Planning and scenario; operational forecasting; financial and enterprise; clinical prediction—with increasing requirements for validation, ownership, monitoring and oversight.

Overrides Recorded and Analyzed

Who overrode, why, what value replaced the forecast and whether it improved the outcome. A repeated pattern is treated as a missing feature, not user error.

Retirement as a Normal Outcome

Models that no longer beat the baseline, no longer drive a decision, or are no longer used are switched off. A register that only grows is not being governed.
A model should never quietly move from planning support into clinical decision making because somebody found another use for the score.
Change of use is a change of tier, and it goes back through approval.
Architecture

Reproducibility is the architecture requirement

Six months after deployment someone will ask why the model produced a particular number on a particular day. If you cannot reconstruct the inputs, the feature values, the model version and the code that ran, you cannot answer, and for a patient facing model that is a governance failure rather than an inconvenience.

Built on the governed platform

Models consume certified data products rather than bespoke extracts. A model with its own private pipeline diverges from the reporting everyone else uses, and then the forecast and the actual disagree for reasons nobody can explain.

Consistent features between training and serving:

The same feature logic used in both, so a model is not trained on one definition and scored on another. Training and serving skew is a common and near-invisible failure.

Versioned everything

Data snapshot, feature definitions, model version, hyperparameters and code recorded together, so any historical prediction can be reproduced exactly.

Separate environments with a promotion path

Development, validation and production separated, with controlled promotion and a rollback that can be executed quickly.

Serve where the decision is made

Into the staffing tool, the scheduling system, the work queue or the planning pack. A forecast in a portal that a planner must remember to open is a forecast that gets used twice.
Service What It Provides to This Page
Healthcare Data Platform Engineering Enterprise ingestion, identity resolution, governance and reusable data products
EHR Data Warehousing and Lakehouse The EHR-centered analytical estate and the historical depth models train on
Operational and Financial Analytics Governed metric definitions and the semantic layer, so the forecast target means the same thing as the actual it is compared against
Forecasting and Statistical Modeling The forward-looking layer: what is likely, how uncertain, what is driving it, and what to commit to before it happens

If finance and operations define volume differently, the model must not silently pick one.
The forecast target uses the governed definition and its owner, or the forecast and the actual will disagree for reasons nobody can explain.

Security and compliance

PHI minimization in training

Models trained on the minimum necessary, using de-identified or limited data sets where the use case permits.

No client PHI used to train third-party models

Enforced contractually and technically, with model hosting and data residency defined and contractually fixed.

Access control on predictions

A patient-level risk score is PHI and is protected as such. Predictions inherit the access controls of the data they derive from.

Transparency where required

Predictive decision support in clinical settings carries disclosure and source-attribute expectations; model characteristics are documented accordingly.

Complete audit trail

Prediction, inputs, model version and consumer logged so any output can be traced and explained.

Trust

Clinical AI requires clinical-grade governance

Ambient documentation involves some of the most sensitive information in healthcare: the conversation between a patient and a clinician. That places it under consent law, privacy law, records retention policy and patient trust obligations that most enterprise AI never touches. These decisions belong to your privacy, legal, compliance and clinical leadership, and we bring them the analysis to make them.

Consent and patient trust

Data protection and retention

Model governance

Responsible AI in clinical use

Governance is not added after the ambient solution is deployed. It determines how the solution is designed in the first place.
Outcomes

Skill against the baseline, and what it changed

Model accuracy alone proves nothing. A model can be accurate, deployed and completely without effect if nobody changed a decision because of it. We report both halves.
Category What We Measure Why It Matters
Skill Accuracy at the decision horizon versus the naive baseline, tracked over time The only honest measure of whether the model is worth its maintenance
Bias Directional error by segment, separately from magnitude of error A small consistent under-forecast produces chronic understaffing while average accuracy looks fine
Calibration Whether stated intervals and probabilities match observed outcomes Determines whether a planner can size against the range
Decision lead time How much earlier a constraint or variance becomes visible, and whether that is before the commitment A correct forecast delivered after the deadline has no value
Decision impact Plans adjusted because of the forecast, and the decisions that changed A model nobody acts on has no value regardless of its accuracy
Operational result Agency and overtime spend, unfilled slots, boarding hours, stockouts, cash forecast variance Where the money actually appears
Adoption and overrides Planner use, override rate, override reasons, and whether overrides improved on the model A consistent override pattern is a missing feature, not user error
Model health Models within tolerance, revalidations completed on time, models retired Whether the portfolio is governed or merely accumulating
Portfolio discipline Models proposed versus deployed, and the proportion stopped at the baseline gate A programme that deploys everything it starts is not measuring honestly
Expect a proportion of candidate models not to survive the baseline comparison.
Agree with the sponsor in advance that a well-evidenced no counts as a successful outcome. Otherwise the engagement gets judged on how many models it shipped rather than whether any should have shipped.

Application innovation backed by deep engineering..

cf difference
Measurable Results

50% reduction in technical debt for enterprise clients

True Partnership Model

Dedicated teams integrated with your workflow

Rapid Innovation Velocity

Ship features 3X faster with our DevSecOps pipeline

Enterprise-Grade Security

SOC 2 compliant engineering practices

Turn healthcare data into forward-looking decisions

We will establish what a simple method already achieves at your decision horizon, assess whether your history can support better, and tell you what a model would realistically add. If the answer is not much, you will have that in weeks and can spend the budget somewhere it will return more.

Start with the clinical workflow, not the ambient AI platform.

Bring us a specialty or clinical setting where clinicians are spending too much time creating notes. We will assess where ambient documentation fits, what must remain clinician controlled, how it should integrate with your EHR, and how to measure whether it is actually reducing burden.

One conversation with people who have run these deployments, and a written readiness view you can use with or without us.

Security & Compliance

caliberfocus certification

Ready to transform your business? Contact us today.

Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.