Healthcare Data Platform Engineering
One Patient. One Definition
One Number.
A healthcare data platform that resolves identity across your systems, certifies what each metric means, and retires the legacy reports it replaces, so leadership stops arguing about whose number is right.
CaliberFocus builds and modernizes data platforms for health systems, hospitals, physician groups and ambulatory organizations. We work domain by domain against real decisions rather than attempting a complete enterprise model up front, resolve patient and provider identity before anything is built on top of it, and put a certified definition behind every published metric. Every AI initiative on your roadmap depends on this foundation, which is why it is usually the constraint rather than the model.
If two teams answer the same question differently, that is not a reporting problem. It is a definitions problem, and no new platform fixes it by itself.
The Challenge
The meeting is about which number is right.
Finance reports one encounter volume. Operations reports another. Quality has a third. Everyone can defend theirs, all three are traceable to a real system, and the meeting that was supposed to be about performance becomes a meeting about methodology. That happens in almost every provider organization, and it is rarely caused by bad data.
It is caused by the same patient existing under five identities, by four teams computing the same measure against different inclusion rules, and by a reporting estate that grew for a decade without anyone being allowed to retire anything. Meanwhile every acquisition brings another EHR instance, every department buys a system with its own database, and payer, HIE and lab feeds arrive in formats that predate most of the people working with them.
AI has made this urgent rather than merely expensive. Ambient documentation needs write-back. Copilots need governed content. Predictive models need data quality good enough to revalidate on a schedule. The foundation is now the constraint on the AI roadmap, not the models.
It also raises the cost of a defect. A human analyst notices that a field looks wrong. An automated system consumes that field thousands of times before anyone notices. Data quality and provenance stop being reporting concerns and become operational controls..Â
Identity resolved differently in every system
The same patient appears multiple times, and the merge history is inconsistent across the estate. Everything downstream inherits that error.
Definitions that were never written down
Business logic living inside reports
Attributed patient, active provider, completed referral, net revenue and readmission get rebuilt independently inside individual reports. The result is technically correct dashboards that disagree with one another.
Analysts assembling rather than analyzing
Hundreds of reports, unknown consumption, and no mandate to switch anything off. Modernization adds a layer instead of replacing one.
Modernization that recreates the old model
Lift and shift moves the existing warehouse into a new platform and inherits every piece of its debt, at higher cost.
Cloud spend with no owner
Consumption pricing makes cost an engineering decision. Without cost visibility built in from the start, the platform succeeds and then gets challenged in budget.
A lakehouse does not resolve a disagreement about what an encounter is.
Technology modernization and semantic governance are two different programs, and organizations routinely fund the first and skip the second. The result is faster access to numbers people still do not trust. We treat the definition as part of the deliverable, not as documentation to be written afterward.
How it works
Build to the question. Not to completeness.
The most common failure mode in healthcare data programs is scope. An eighteen month effort to model the enterprise delivers nothing usable until month fourteen, by which point the sponsor has changed and the requirements have moved. We work domain by domain, each one driven by a decision somebody is waiting on, and each one shipping something that replaces a thing you are running today.
Step 1
Start from the decision
Identify the question, who asks it, how often, and what would change if the answer were trusted. No domain starts without a named consumer.
Step 2
Profile the sources
Step 3
Resolve identity
Establish resolution for the entities that connect healthcare data: patient, encounter, provider, location, organization, payer, plan, claim
Step 4
Model the domain
Design the domain model against the questions in scope, using an established healthcare model where one fits rather than inventing structure.
Step 5
Ingest and standardize
Build the pipelines and normalize terminology, code sets and reference data into a consistent representation.
Step 6
Define the metrics
Write the definition for every published measure with the business owner. Inclusion, exclusion, grain, timing and known limitations.
Step 7
Certify and publish
Publish as a data product with a named steward, a certification tier, documented lineage and a stated refresh commitment.
Step 8
Retire what it replaces
Identify the legacy reports and extracts the new product supersedes, migrate consumers and switch the old ones off.
Step 9
Operate
Quality monitoring, freshness alerting, lineage, access review and cost management as running processes with owners.
| Capability | EHR reporting layer | Legacy EDW | Lift and shift to cloud | Governed platform |
|---|---|---|---|---|
| Data beyond the EHR | Limited | Yes | Yes | Yes |
| Identity resolved across sources | Within the EHR | Partially | Inherited as-is | Explicit, and measured |
| Certified metric definitions | Vendor defined | Rarely documented | Carried over | Required to publish |
| Supports AI and ML workloads | No | Poorly | Sometimes | Yes, by design |
| Legacy estate retired | Not applicable | No | No | Yes, per domain |
| Cost visibility | Bundled | Fixed and opaque | Consumption, ungoverned | Consumption, attributed |
| Time to answer a new question | Fast if it fits, otherwise no | Weeks to months | Unchanged | Days within a live domain |
Every new data product retires an old one.
Running the legacy report alongside the new one is how organizations end up paying for both indefinitely and trusting neither. Retirement is part of each domain deliverable, agreed with the consumer before the work starts, and it is the step most programs quietly drop when the timeline compresses.
Capabilities
Healthcare data is not just enterprise data with different column names
A general data engineering team can build pipelines. What takes disproportionately longer to learn is why a lab result arrives three times with different statuses, why a claim looks nothing like the encounter it came from, and why merging two patient records incorrectly is a safety event rather than a data quality ticket.
Ingest and Integrate
Clinical and Operational Source Connectivity
EHR reporting databases and APIs, departmental clinical systems, laboratory and imaging, scheduling, ERP, HR, CRM and patient engagement platforms, on the ingestion pattern each source actually supports.
Healthcare Standards Ingestion
FHIR including bulk export, HL7 v2 messaging, X12 claims and remittance, CDA and CCDA documents, and flat file and database extracts where that is what exists.
Currency and Version Control
Streaming and Batch Pipelines
Near real time where a decision depends on it, batch where it does not. Latency is an expensive property and it is worth being deliberate about which domains genuinely need it.
Unstructured and Document Content
Notes, faxed documents, correspondence and reports brought into the platform as governed content, so text based analysis and retrieval have a source of record.
Model and Standardize
Identity Resolution
Patient, provider, location, facility hierarchy and payer identity, built against your existing master data capability where you have one and stood up where you do not.
Terminology and Code Set Management
Mapping and ongoing maintenance across SNOMED CT, LOINC, RxNorm, ICD-10, CPT, HCPCS and local code sets, treated as a maintained asset rather than a one-time exercise.
Domain Modeling
Modeled for the questions being asked, using established healthcare models where they fit the purpose rather than inventing structure that only your team will understand.
Semantic Layer and Metric Definitions
One place where a measure is defined, versioned and owned, consumed by every downstream tool. This is what makes the number the same wherever it is read.
Operate and Govern
Data Quality Monitoring
Automated tests on completeness, freshness, volume, referential integrity and distribution shift, alerting the steward rather than surfacing to the consumer as a wrong number.
Lineage, Catalog and Certification
End to end lineage from source to published metric, a catalog people actually use, and a certification tier so consumers know what is production grade and what is exploratory.
Access, Segmentation and Audit
Row and column level controls, sensitive category segmentation, de-identified and limited data set provisioning, and complete access logging.
Cost Engineering
Consumption modeled, attributed to domain and consumer, alerted on, and designed for. Cost is an architecture decision on modern platforms, not a finance report.
We are not reselling a platform and we are not vendor aligned. We assess what your EHR data layer can and cannot do before recommending anything beside it, build on the cloud platform you have already committed to, and design so that the model and platform layer can change without rebuilding governance and integration around it. We also build the AI systems that consume this foundation.
The Domains
Sequence by dependency, not by enthusiasm
Domains are not independent. Identity underpins everything. Encounter underpins most clinical and financial analysis. Building the interesting domain before the ones it depends on is the most common reason a data program stalls in its second year. The table below is our usual sequencing position, which your systems and priorities will adjust.
| Domain | Primary sources | What it unlocks | Sequence |
|---|---|---|---|
| Patient identity | EHR, registration, HIE, payer files, master data platform | Everything. No domain below is reliable without it. | First, always |
| Encounter and utilization | EHR, scheduling, ADT, claims | Volume, throughput, capacity, access and the denominator for most measures | First |
| Provider and network | Credentialing, HR, scheduling, directory, claims | Productivity, panel, referral patterns, network integrity | First |
| Revenue cycle | Practice management, clearinghouse, remittance, contracts | Days in AR, denial performance, yield, cost to collect | Early, high return |
| Clinical results and orders | EHR, lab, imaging, pharmacy | Quality measures, care gaps, clinical analytics, AI features | Second |
| Access and scheduling | Scheduling, contact center, digital front door | Third next available, no show, slot utilization, leakage | Second |
| Cost and finance | ERP, general ledger, cost accounting, payroll | Service line margin, cost per case, budget variance | Second, needs encounter |
| Quality and measures | EHR, claims, registries, abstraction | Regulatory reporting, value based contracts, improvement work | Third, needs clinical |
| Claims and population | Payer files, HIE, ADT feeds, risk platforms | Total cost of care, leakage, risk adjustment, population health | Third, needs identity |
| Workforce | HR, scheduling, time and attendance, credentialing | Staffing models, turnover, agency spend, productivity | When workforce is a priority |
| Supply chain | ERP, item master, purchasing, clinical documentation | Preference card variance, cost per procedure, standardization | When margin work demands it |
| Patient experience | Survey vendors, CRM, digital, complaints | Experience by service line, provider and access pathway | When linked to encounter |
The Value Is in the Chain, Not in Any One Domain
Referral → appointment → encounter → claim → payment
Where patients fall out of the funnel, and what referral performance is worth in realized revenue.
Clinical documentation → coding → claim → denial
Which documentation and coding patterns are producing avoidable denials, traced to the source behavior.
Scheduling → staffing → capacity → access
Patient → encounter → quality → utilization
Payer → authorization → claim → denial → payment
Architecture
Architect so the platform layer can Change
The cloud data platform market has re-sorted itself twice in five years and will do so again. The parts of your architecture that should be durable are identity resolution, terminology, metric definitions, lineage and access control. The parts that will change are storage, compute and the query engine. Designs that entangle the two are how organizations end up unable to move.
| Standard | Typical use in the platform |
|---|---|
| FHIR, including bulk export | Structured clinical extraction, USCDI aligned data elements, application integration |
| HL7 v2 | ADT, orders, results and scheduling events from systems that will not move to FHIR |
| X12 | Claims, remittance, eligibility and authorization transactions |
| CDA and CCDA | Document based exchange, external records and transitions of care |
| Terminology standards | SNOMED CT, LOINC, RxNorm, ICD-10, CPT and HCPCS normalization across sources |
| National exchange frameworks | HIE, TEFCA and QHIN participation as an external data source and an obligation |
Separate the durable from the replaceable
Identity, terminology, definitions, lineage and access control remain portable assets.
Layered, not monolithic
Source → standardized → curated data products → semantic models → serving layers.
The EHR data layer has a role
Use it where it fits operational reporting, but not as a strategy for cross-source, longitudinal or AI workloads.
Latency is a cost decision
Real time is right for a handful of domains and expensive everywhere else.
Built for people and machines
The same governed foundation serves analysts, applications, ML models, copilots and agents.
Cost designed in, not reviewed later
Partitioning, materialization, refresh cadence and workload isolation are architecture decisions.
We build on Microsoft Fabric, Azure, Databricks, Snowflake, AWS and Google Cloud. Platform selection follows your existing enterprise commitment, skills base and workload profile, not our preference. Where you have already committed, we build there and design for portability rather than reopening the decision.
Trust
An incorrect merge is a patient safety event
Most data quality frameworks treat every defect as a ticket with a severity. In healthcare, one class of defect is categorically different. Linking two patients who are not the same person puts one patient record inside another, and the clinical consequences of that are not a reporting problem. Master data quality is governed here as a safety control, not as a data hygiene metric.
| Metric | What it means | How it is treated |
|---|---|---|
| Duplicate rate | One person existing as multiple records | A quality target, worked down systematically |
| Overlay rate | Two different people merged into one record | A safety metric with a zero tolerance target and mandatory root cause review |
| Unresolved match rate | Records the matching process could not confidently link | A stewardship queue with a service level, never auto-resolved |
| Cross-source linkage | Share of external records successfully linked to a known patient | Determines what population and claims analysis can be trusted |
Completeness and conformity
Required elements present and valid against the expected code set, monitored per source rather than in aggregate.
Timeliness and freshness
A stated refresh commitment per data product, with alerting when it is missed and visible staleness for consumers
Consistency across sources
The same fact agreeing across systems, with reconciliation rules where it does not and a documented winner.
Distribution monitoring
Certification tiers, so consumers know what they have
| Tier | What it means | Appropriate use |
|---|---|---|
| Certified | Defined, owned, monitored, lineage documented, refresh committed | Board reporting, regulatory submission, contractual measures, AI features |
| Provisional | Built and reviewed, definition agreed, monitoring incomplete | Operational analysis and improvement work, with the caveat visible |
| Exploratory | Available, not governed, no commitment made | Analyst investigation only. Never published or presented externally |
Every certified data product carries last refresh, source systems, named owner, definition, quality rules and failures, known limitations, lineage and certification tier where it is consumed.
Stewardship sits in the domain
The finance steward owns financial definitions and the clinical steward owns clinical ones. A central team that owns every definition becomes the bottleneck and then gets bypassed.
A definitions registry with versions
Every published measure has a written definition, an owner, a version history and a change process. When a definition changes, everyone consuming it is notified.
Change control on upstream sources
An EHR upgrade, a vendor release or a new build in a source system can change what your data means. Source change is a governance input, not just an IT one.
Impact analysis before change
Lineage answers what breaks if we change this, before the change is made rather than after a downstream report goes wrong.
Access review as a running process
Entitlements reviewed on a cycle, not set once at go-live and inherited forever by people who changed roles years ago
Compliance
A data platform concentrates risk by design
The purpose of the platform is to bring together data that was previously separated by system boundaries. That is the value and it is also the risk. Controls that were adequate when information was fragmented across a dozen systems are not adequate once it is joined, and the review should be done before the joining, not after.
Access and minimum necessary
- Row and column level security enforced in the platform, so entitlement is a property of the data rather than of each report
- Role based access aligned to enterprise identity, with entitlement reviewed on a defined cycle
- De-identified and limited data set provisioning as a standard path, so the default request does not have to be for identified data
- Complete access logging with the ability to answer who accessed which patient records, when and why
Sensitive category handling
- Segmentation for substance use disorder records, behavioral health, HIV, genetic and reproductive health information, and records relating to minors
- Restrictive authorizations and patient requested restrictions carried through into the platform rather than lost at ingestion
- Employee and VIP patient records handled under break-glass controls with proactive access monitoring
- Segmentation designed at ingestion. Retrofitting it after a domain is built is expensive and frequently incomplete
Environment and AI access
- Development, test and production separated, with controlled rules on how production healthcare data may be used outside production
- For every AI system consuming the platform: what data it can retrieve, whose identity and permissions apply, whether prompts and outputs are retained, whether anything may be used for training, how interactions are logged, and how an output traces back to source data
- Machine and service identities governed on the same basis as human ones, with entitlement reviewed on the same cycle
- Security follows the data. Identity, provenance and permitted use stay intact from EHR to platform to semantic model to dashboard to API to AI application, not only inside the platform boundary
Secondary use governance
- A defined path distinguishing operations, quality improvement, research and any commercial or partnership use, with different approval routes
- Research use routed through your existing review process, with de-identification method documented and defensible
- An explicit organizational position on data sharing and monetization arrangements, agreed before a partner asks rather than during a negotiation
- Contractual and consent constraints from source data carried forward, since not all data you hold can be used for everything you hold it for
Provider organizations are increasingly approached about data partnerships, research collaborations and commercial arrangements. This belongs in the platform design conversation, not in a later negotiation.
Outcomes
Measure trust and time, not terabytes
Platform programs are frequently reported on inputs: sources connected, tables built, volume ingested. None of those tell you whether anyone trusts the output or got an answer faster. We baseline before the first domain and report on what changed for the people asking the questions.
| Category | What we measure | Why it matters |
|---|---|---|
| Time to answer | Time to answer a new question within a live domain, and analyst time assembling versus analyzing | TThe most direct measure of whether the platform is working |
| Trust | Reconciliation escalations, competing versions of the same measure in circulation, certified metric coverage | The which-number-is-right problem, made countable |
| Identity quality | Duplicate rate, overlay rate, unresolved match queue age, cross-source linkage rate | Overlay is reported as a safety metric with a zero target |
| Estate reduction | Legacy reports and extracts retired, duplicate pipelines removed, systems decommissioned | Whether you replaced something or just added a layer |
| Reliability | Freshness commitments met, quality test pass rate, incidents and time to detection | Consumers stop checking whether the data loaded |
| Cost | Platform cost by domain and consumer, cost per workload, trend against volume growth | Consumption platforms succeed technically and then get challenged in budget |
| AI enablement | Domains meeting the readiness gate for planned AI use cases | This foundation is usually the constraint on the AI roadmap |
That is where identity, standards, patterns and governance are established rather than reused. Expect the first to feel slow and the third to feel fast.
Build your healthcare data foundation
Not a data strategy request. A specific measure that gets reported differently by different teams, one that takes three weeks to produce every month, or the reason every new analytics project needs another six months of data engineering before it starts. We will trace it to source, show you why the versions differ, and tell you what it would take to make it a single certified number. That exercise almost always exposes the identity, definition and ownership issues that the wider platform work has to address, and it does it in weeks rather than in a discovery phase.
Start with the clinical workflow, not the ambient AI platform.
Bring us a specialty or clinical setting where clinicians are spending too much time creating notes. We will assess where ambient documentation fits, what must remain clinician controlled, how it should integrate with your EHR, and how to measure whether it is actually reducing burden.
- AI Agents and Workflow Automation
- Voice and Conversational AI
- Document AI and Intelligent Processing
- Generative AI and Enterprise Copilots
- AI Strategy and Governance
- HCC and Risk Adjustment Analytics
Security & Compliance
