Document AI and Intelligent Processing
Read Every Document. File It in the
Right Place. Every Time.
The Challenge
The information arrived. Nobody can find it.
A referral packet lands in the fax queue at four in the afternoon. It contains a cover sheet, a referral form, three progress notes, an insurance card image and a lab report, all in one transmission. Someone has to separate them, identify each one, work out which patient they belong to, extract the detail that triggers the next step, file each document to the right place in the chart, and start the referral. That is roughly ten minutes of work, and it happens hundreds or thousands of times a week.
The variability is the problem, not the volume
Manual classification
Manual extraction
Manual matching and validation
Manual routing
Manual exception handling
A result filed to the wrong chart, or never filed at all, is a clinical risk with a documented history of harm. It should be governed with the seriousness that implies, and measured against a target of zero rather than reported as an accuracy percentage.
How it works
From inbound transmission to correctly filed record
Step 1
Ingest
fax, secure email, SFTP, scanner, portal upload, health information exchange or API, into a single processing pipeline.
Step 2
Split
Step 3
Classify
Step 4
Extract
Relevant fields are located and read wherever they appear, including handwriting, checkboxes, tables and stamps, with a confidence score on each value.
Step 5
Match
Step 6
Validate
Step 7
Route and file
Step 8
Learn
| Capability | Manual indexing | Template OCR | Document AI |
|---|---|---|---|
| Handles a layout never seen before | Yes | No | Yes |
| Splits multi-document bundles | Yes | Rarely | Yes |
| Reads handwriting and checkboxes | Yes | Poorly | Yes |
| Confidence score on every field | No | No | Yes |
| Validates against source systems | Sometimes | No | Yes |
| Adding a new document type | Training a person | A new template build | Configuration and examples |
| Cost and speed at volume | High, slow | Low, brittle | Low, resilient |
A document that is ninety five percent accurate sounds excellent until the five percent is the medical record number. We score and threshold every field separately, and set the tolerance on identifier and clinical fields far tighter than on a fax cover date. Document level confidence hides exactly the errors that matter most.
Capabilities
Extraction is the middle of the job, not the job
Most document AI evaluations test extraction accuracy on clean samples and stop there. In production, the work that determines whether anything is saved sits either side of extraction: separating the bundle correctly at the front, and matching, validating and filing correctly at the back.
Ingest and Understand
Multi-Channel Ingestion
Bundle Splitting
Separates multi document transmissions at the correct boundaries by reading content rather than counting pages or looking for separator sheets. This is the step that silently corrupts everything downstream when it goes wrong.
Document Classification
Extraction on Real-World Quality
Resolve and File
Patient and Encounter Matching
Validation Against Source Systems
Cross-Document and Completeness Validation
Document Summarization
Structured Output and Coded Mapping
Chart Filing and Workflow Trigger
Control and Improve
Per-Field Confidence and Thresholding
Provenance and Source Traceability
Exception Queue and Reviewer Console
Accuracy Monitoring by Type and Sender
What CaliberFocus does, and does not do?
We are not a capture or content management platform vendor, and that is deliberate. We assess your real document mix against your systems, select and configure the right processing stack, build the matching, validation and filing logic that determines whether documents land correctly, establish the confidence thresholds and governance with your health information management leaders, and own accuracy through the first document types.
Where it applies?
Start with the document type You receive most and touch twice
| Workflow | What the system handles | What always routes to a person | What moves |
|---|---|---|---|
| Referral intake | Splits the packet, classifies each document, extracts patient, payer, ordering provider and clinical detail, checks completeness and creates the referral. | Clinical triage, urgency and appropriateness decisions. | Referral leakage, time to appointment |
| Results, reports and clinical correspondence | Classifies, extracts key values, matches to the ordering encounter, files to the correct chart location and notifies the ordering clinician. | Abnormal or critical values, and any ambiguous patient match. | Time from receipt to clinician visibility |
| Orders and requisitions | Reads the order, validates against the record, extracts diagnosis and authorization detail and queues for acknowledgment. | Clinical validation and anything requiring clarification with the ordering provider. | Order entry time, scheduling delay |
| Prior authorization determinations | Classifies approvals, denials, pends and requests for information, extracts reference numbers and criteria, and routes by outcome. | Medical necessity interpretation, peer to peer and appeal strategy. | Determination turnaround, delayed procedures |
| Payer correspondence and remittance | Extracts from paper and non standard remittance and correspondence, reconciles against the claim and routes variances. | Unresolved variances and cash application exceptions | Posting lag, unapplied cash |
| Release of information and records requests | Classifies the request, validates authorization scope and expiry, identifies the records in scope and assembles the response for review. | Every disclosure decision, and any request involving sensitive categories. | Request turnaround, regulatory response times |
| Patient forms and intake documents | Reads completed forms including handwriting and checkboxes, validates against the record and updates registration. | Conflicting information and anything requiring verification. | Registration rework, check in time |
| Credentialing and provider documents | Classifies licenses, certifications, insurance and attestations, extracts expiry dates and updates provider records. | Committee decisions and primary source verification judgment. | Time to credential, revenue lost to lapses |
| Fax and shared inbox triage | Classifies everything arriving in the shared fax and scan inbox, identifies the intended workflow or recipient and routes it without a person opening each item. | Unrecognized documents, uncertain routing and anything flagged urgent. | Inbox backlog, time to route |
| Chart abstraction and audit support | Locates, retrieves and pre-abstracts documents for audits, quality reporting and payer requests. | Clinical validation of every abstracted finding. | Abstraction hours, audit response time |
| Administrative and enterprise documents | Invoices, contracts, vendor documents and HR forms processed on the same pipeline as clinical content. | Commercial decisions and policy exceptions. | Processing cost, cycle time |
Substance use disorder records, behavioral health, HIV and other specially protected results, genetic testing, records relating to minors, and anything arriving with a restrictive authorization are routed to a qualified reviewer before filing. These categories carry disclosure restrictions that a filing rule cannot safely interpret, and a correct extraction filed to the wrong access level is still a disclosure event.
Integration
Extraction without filing is a spreadsheet nobody asked for
Plenty of tools will read a document and give you structured data back. The saving only arrives when that data reaches the chart, the queue and the workflow without a person moving it. Filing correctly is a deeper integration problem than extraction, and it is where these programs are won or lost.
Correct chart destination
The image and the data together
Downstream workflow triggering
Filing is not the end. The referral is created, the ordering clinician is notified, the denial is routed, the request is queued.
Sender feedback loop
One audit trail
Graceful degradation
EHR and practice management
Document management and content platforms
Storage, retention, versioning and retrieval within your existing repository
Fax and secure transmission platforms
Scanning and capture infrastructure
Front desk and back office scanning brought into the same pipeline
Integration engine
Clearinghouse and revenue cycle systems
FHIR and vendor APIs
Structured write back of results, orders and observations where supported
Data and analytics platforms
Keep the systems of record in control
PHASE 1
Document mix assessment
2 to 3 weeks
PHASE 2
Build and calibrate
4 to 6 weeks
PHASE 3
Shadow and pilot
3 to 4 weeks
PHASE 4
Production and scale
Control
Automation earns its autonomy one field at a time
Straight through processing is not a setting you switch on at go live. It is a rate that rises as measured accuracy justifies it, per document type and per field. A vendor who quotes you a straight through rate before seeing your documents is quoting you their average, not your outcome.
| Level | What the system does | What a person does | Typical stage |
|---|---|---|---|
| Classify and route | Identifies the document and sends it to the right queue. Extracts nothing that is relied upon. | Keys and files everything | First weeks, and any new document type |
| Extract and present | Presents extracted values alongside the source image for confirmation. | Confirms or corrects every field, then files. | Once classification accuracy is proven |
| File with exception review | Files automatically where every field clears its threshold. Routes documents with any low confidence field. | Works only the flagged fields, not whole documents. | The working state for most document types |
| Straight through | Files and triggers the workflow with sampled audit only. | Audits samples and investigates alerts. | High volume, low variance types with proven accuracy |
| Filing on an unresolved patient match | Not offered. | Not applicable | Never |
Thresholds set per field,
by risk
Patient match has a
hard floor
Source image beside
every value
Reviewers see the extracted value and the region of the page it came from together, which turns a correction from a search into a glance.
Field level review, not document level
A reviewer confirms the two uncertain fields rather than re-reading the whole document. This is where most of the labor saving actually comes from.
Second review on critical fields
Defined high risk fields can require independent confirmation before filing, in line with your existing HIM controls.
Misfile detection and correction
A defined route for clinicians and staff to report a misfiled document, with correction, root cause analysis and threshold adjustment, not just a re-file.
Corrections drive the model
Every override is captured with its reason and feeds accuracy monitoring, threshold tuning and configuration by document type and by sender.
A system that files everything and flags nothing is not more accurate. It is less honest. We would rather send you a document with one uncertain field than file it confidently in the wrong chart, and we design the review experience so that handling the exception takes seconds.
Trust
Every document is PHI until proven otherwise
Data
protection
- BAA executed before any access to protected health information
- Encryption in transit and at rest for images, transcripts, extracted data and logs, with key management under your control where required
- Minimum necessary access. Each pipeline stage receives only the data its function requires
- Retention and deletion defined separately for source images, extracted data and processing logs, aligned to your legal medical record definition
- Data residency and processing location defined and contractually fixed
Sensitive & misdirected documents
- Specially protected categories, including substance use disorder, behavioral health, HIV, genetic and minor related records, route to a qualified reviewer before filing
- Restrictive authorizations and disclosure limitations are read and honored rather than ignored at filing
- A defined procedure for documents received in error, including quarantine, non-filing and notification to the sender
- Minimum necessary collection, processing and access for the approved documentation workflow
- Role based access to recordings, transcripts, drafts, configuration and administrative functions
- Data residency and processing location defined and contractually fixed
Model
governance
- No client documents, images or PHI used to train foundation models, enforced contractually and technically
- Model version and configuration documented per document type, with change control on anything affecting extraction or classification
- Accuracy evaluated against a gold standard set built from your real documents, including degraded and handwritten samples rather than clean ones
- Performance reported by document type and by sender, never as a single aggregate accuracy figure
- Accuracy requirements set by the consequence of getting the field wrong, not one automation threshold applied to every field and document
- Drift alerts and defined rollback when a sender format change degrades accuracy
Accountability &
audit
- Every extracted value traceable to its source region on the page, and every filing action attributable and time stamped
- Complete audit trail linking document, extraction, review, correction and final filing destination
- Misfile rate reported as a safety metric with a zero tolerance target, not folded into an accuracy percentage
- A named accountable owner for the pipeline, with health information management leadership sign off on thresholds and sensitive category handling
Outcomes
Report accuracy by document type, or do not report it
| Category | What we measure | Why it matters |
|---|---|---|
| Safety | Misfile rate and time to detection, by document type | Reported as a safety metric with a zero tolerance target |
| Accuracy | Field level accuracy by document type and by sender, with identifier fields reported separately | Aggregate accuracy hides the errors that matter |
| Coverage | Straight through rate by document type, and the categorized reason every other document needed a person | The second half is what drives the fix backlog |
| Speed | Receipt to filed, receipt to clinician visibility, receipt to downstream action, backlog age | Clinical information nobody can see is not filed |
| Capacity | Documents per FTE, staff hours returned, touches per document, overtime and outsourcing spend | The operating business case |
| Downstream | Referral leakage, time to appointment, authorization turnaround, posting lag, audit response time | Where document processing actually shows up in the P and L |
| Workflow completion | Whether the complete downstream workflow finishes faster, not only whether the document was processed faster | The test of whether automation reached the work or stopped at the document |
| Quality of intake | Incomplete and unreadable documents by sender | Some of the biggest wins are fixing the sender, not the system |
Automate the right types well and leave the messy ones in review rather than forcing a misfile problem.
Identify your first document automation opportunity
Start with the clinical workflow, not the ambient AI platform.
Bring us a specialty or clinical setting where clinicians are spending too much time creating notes. We will assess where ambient documentation fits, what must remain clinician controlled, how it should integrate with your EHR, and how to measure whether it is actually reducing burden.
- AI Agents and Workflow Automation
- Voice and Conversational AI
- Document AI and Intelligent Processing
- Generative AI and Enterprise Copilots
- AI Strategy and Governance
- HCC and Risk Adjustment Analytics
Security & Compliance
