Contact Us

Product Scalability and Performance

The Thing Everyone Blames
Is Usually Not the Thing

Performance and scalability engineering for healthcare and revenue cycle products, starting with measurement rather than with the component the team has already decided is at fault.
Every engineering team has a theory about why the product is slow. The database. The reporting layer. That one legacy service. Those theories are held confidently, they are frequently wrong, and acting on them costs weeks of work that changes nothing a customer can feel. A slow screen is often queue contention from a batch job. A slow report is often a lock held by something else entirely. The measurement usually takes days and it almost always redirects the effort.
Optimizing the wrong component is not a small mistake. It is six weeks, a release, and a customer who is still complaining.
The Challenge

Healthcare load is bursty, and averages hide all of it

Performance deterioration is rarely sudden. A nightly job takes forty minutes instead of twenty, then ninety. An API that responds in three hundred milliseconds occasionally takes three seconds. A claim file that finished before the workday now runs into business hours.
Healthcare products do not experience even load. Billing runs cluster at month end. Clinic systems peak in the morning. Payer submissions follow a cycle. Eligibility checks spike when a practice opens. A product performing acceptably on average can be unusable during the four hours a customer most needs it.

Diagnosis Happens by Opinion

The prevailing theory is repeated confidently in incident reviews and often has never been tested against measurement.

Averages Conceal the Customer Experience

The median is fine. Customers who churn are living in the tail during peak volume.

Products Are Tested at Development Scale

It works with ten thousand records. Healthcare customers arrive with ten years of history.

Performance Becomes Somebody's Problem After a Complaint

Without monitoring designed for it, the first signal is already an unhappy customer.

One Workload Starves Another

Reporting competes with transactions, batch competes with interactive, and integration competes with both.

Nobody Knows the Capacity Limit

The breaking point is discovered by a customer rather than in a test.

Look at your ninety-ninth percentile during your busiest hour.

Not the average, and not a quiet Tuesday. The slowest one percent of requests during the period when customers are working hardest is the experience that generates complaints, escalations and churn.
Our Approach

Measure, locate, prove, then optimize

The sequence protects against the most expensive failure mode in performance work: spending effort on something that was never the constraint. Nothing gets optimized until it has been demonstrated to be the problem.

Step 1

Establish What Actually Hurts

Which workflows, customers and times, and what business consequence?

Step 2

Instrument Properly

Create a baseline before changing anything.

Step 3

Profile Realistically

Real data volumes, concurrency, workload mix and peak hour.

Step 4

Locate the Bottleneck

Application, database, query, lock, queue, network, integration or infrastructure—from evidence.

Step 5

Test the Hypothesis Cheaply

A small experiment costs a day; a wrong optimization costs a sprint.

Step 6

Optimize and Re-measure

Relieving one bottleneck reveals the next

Step 7

Separate Competing Workloads

Often a larger improvement than optimizing either workload individually.

Step 8

Establish Safe Capacity

Build in headroom for growth, variation, failure and seasonal peaks.

Step 9

Model the Next Growth Stage

Current load, 2×, 5×, largest expected customer and peak plus failure.

Step 10

Build the Monitoring

Catch the next problem before a customer reports it.

Test at the data volume your largest customer will reach, not the one they start with.

Healthcare customers accumulate. Query plans change, indexes stop helping and memory behaviour shifts. Realistic data-volume testing finds failures before your largest customer does.
Capabilities

Find it, fix it, and know before the customer does

Performance engineering is mostly diagnostic work. The optimization is frequently straightforward once the constraint is genuinely known—and almost always wasted when it is not.

Measure and Diagnose

Performance Assessment

Current behaviour across application, database, integration and infrastructure at peak, with the tail reported rather than the mean.

Bottleneck Analysis

Profiling, tracing and query analysis to locate the actual constraint.

Load, Stress, Spike, Soak & Recovery Testing

Expected demand, breaking point, sudden traffic, sustained load and recovery after saturation or dependency failure.

Data Volume Testing

Behaviour at the scale the largest customer will reach.

Optimize

Query & Database Optimization

Query plans, indexing, locking, contention and schema decisions.

Application & API Optimization

Algorithms, caching, payload design, chattiness and concurrency handling.

Workload Separation

Give reporting, batch, integration and interactive work appropriate resources.

Asynchronous & Batch Redesign

Move work out of the request path when a user does not need to wait for it.

Sustain

Capacity Planning & Modelling

Model growth before it becomes an incident.

Performance Monitoring Design

Percentile-based, workload-aware and tenant-aware monitoring.

Performance Regression Testing

Automated checks in the delivery pipeline.

Scale-Cost Engineering

Improve throughput without proportionate infrastructure growth.

What CaliberFocus does, and does not do?

We do not start by adding infrastructure. We do not optimize a component until measurement shows it is the constraint. We will say when separating workloads is the fastest improvement, and when more infrastructure is the honest short-term answer while a structural fix is planned.
Where It Applies

Every healthcare product has a peak somebody scheduled

Healthcare load is driven by calendars, working hours and submission cycles rather than random demand. That is useful because a predictable peak can be planned for.
Product Type What Drives the Load The Peak That Breaks It
RCM Platforms Claim submission, remittance posting, denial work Month end and overnight batch colliding with a customer working late.
Eligibility & Verification Real-time checks before appointments Clinic opening hours across time zones.
Coding & Documentation Processing clinical content per encounter End of clinic day when a full day of documentation arrives in two hours.
Patient Engagement Consumer-driven traffic Campaigns, statement mailings or reminder sends.
Clinical Applications Continuous clinical workflow Shift change and morning rounds.
Payer Platforms Administrative and enterprise workload Open enrollment, renewal and regulatory submission windows.
Analytics Products Query load over accumulating data Reporting periods when every customer runs month-end analysis.
Do Not Optimize Your Side of an Eight-Second Transaction From 300ms to 200ms
If the EHR or payer API takes eight seconds, optimize where the customer actually waits—often through caching, asynchronous processing, prefetching, parallelism or a different expectation.

The Slow Screen Is Rarely the Slow Code

The screen may be waiting on a database, waiting on a lock, held by a batch job for another customer. Trace the whole request path and ask what else was happening.
The Method

Six places the constraint actually sits

They produce similar symptoms and require entirely different remedies. Establish which is the constraint before optimizing anything.
Constraint What It Looks Like How to Confirm It
Database Query Slow response worsening with data volume Query plans, execution statistics and realistic data scale.
Lock & Contention Intermittent slowness affecting unrelated features Lock waits and blocking analysis during the actual period.
Queue & Backlog Work completes eventually with growing delay Queue depth and age over time.
Application Code Consistent slowness proportional to input size Profiling under load.
External Dependency Latency the product does not control End-to-end tracing and separate dependency measurement.
Resource Saturation Everything degrades together at a threshold Utilization against capacity, including the resource nobody watches.
The workload model matters more than the tooling.
It passed the load test” is meaningless alone. At what concurrency? Against what transaction mix? At what data volume? For how long? With which dependencies? What happened to cost while it passed?

Report Percentiles, Not Averages

P95 and P99 during peak describe the experience customers report.

Reproduce Before You Fix

If the problem cannot be reproduced, the optimization cannot be shown to have worked.

Change One Thing at a Time

Preserve attribution and avoid creating new regressions.

Re-measure Every Change

A performance programme is a sequence of bottlenecks.

Ask What Else Was Running

Interference between workloads is a common healthcare failure mode.

Protect the Fix

Turn the improvement into a regression test.

Scaling up is the answer that always works and always costs.

Adding infrastructure can relieve almost any performance problem without understanding it. Sometimes that is the right short-term choice. It should have an end date and a documented underlying problem—not become a reflex that grows with the customer base.
Integration

The layer with the problem and the layer with the symptom are different

Performance problems present at the top of the stack and originate lower down. Tracing that chain end to end is the work.

Application & Service

Algorithms, concurrency, caching and service chattiness under load.

Database

Query plans, indexing, locking, contention and connection behaviour.

API

Payload size, round trips, realistic pagination and rate limiting.

Integration

Partner latency, retry behaviour, queue depth and external throughput.

Data & Analytics

Reporting and analytical load competing with transactional work.

Infrastructure & Cloud

Sizing, storage, network and unmonitored resource saturation.

Performance principles

Trace across boundaries. Instrument the business transaction. Keep analytical load off the transactional path. Measure external dependencies separately. The slowest layer owns the customer experience.
Trust

In healthcare products, slow and broken are the same thing

A clinician who cannot get a screen to load moves to paper. A biller watching a spinner during month end calls support. An eligibility check taking ninety seconds is functionally unavailable to a front desk with a patient waiting.

Monitoring

Percentile latency by workflow, tenant and time; business transactions; queue depth and age; degradation alerts; latency, traffic, errors and saturation together.

Reliability & Resilience

Customer-experienced commitments, graceful degradation, backpressure and circuit breaking.

Testing & Regression

Realistic load, data-volume testing, automated regression and performance budgets for critical workflows.

Engineering Governance

Ownership, targets, capacity assumptions, escalation paths, forward capacity review and evidence-based decision records.

Set a performance budget and let it fail the build.

Choose the workflows that matter most, agree the latency they must stay within at realistic load, and make exceeding it break the pipeline. That catches degradation when it is still cheap to fix.
Outcomes

Faster at peak, larger volumes, cost that does not follow

Performance work reported on average response time understates what was achieved. Measure what customers experience and what the business pays.
Category What We Measure Why It Matters
Peak Percentile Latency P95 and P99 during the busiest period, by workflow The experience generating complaints, escalations and churn.
Volume Headroom Transaction and data volume before degradation What can be promised to prospects and how long before the issue returns.
Cost per Transaction Infrastructure cost against throughput and its trend Distinguishes engineering improvement from simply buying hardware.
Workload Isolation Incidents where one workload or tenant degraded another Common source of unexplained complaints.
Detection Problems found by monitoring versus customers Shows whether the team is ahead of the problem.
Regression Regressions caught in pipeline versus production Shows whether improvement is durable.

Honest expectation setting

Not every problem requires architectural change. Sometimes the highest-value improvement is an index, a query or separating reporting from transactions. Not every slow transaction is yours to fix. Measurement may contradict the prevailing theory, and the first fix will often reveal a second constraint. Agreeing in advance that measurement governs is what makes the engagement useful.

Handle more volume, improve performance and scale without increasing cost at the same rate

It may be a claim batch that no longer finishes overnight, a report that takes thirty seconds, an eligibility request that occasionally times out, an ERA process that cannot clear Monday backlog, or a customer whose volume affects everybody else. Those are better starting points than asking whether the platform needs more servers. We will instrument the workflows that matter, profile them under realistic volume and concurrency at your actual peak, locate the constraint from evidence, and test the fix before committing engineering time to it. Part of the output is usually that the prevailing theory was wrong, which is the most valuable finding available and rarely the expected one.

Start with the clinical workflow, not the ambient AI platform.

Bring us a specialty or clinical setting where clinicians are spending too much time creating notes. We will assess where ambient documentation fits, what must remain clinician controlled, how it should integrate with your EHR, and how to measure whether it is actually reducing burden.

One conversation with people who have run these deployments, and a written readiness view you can use with or without us.

Security & Compliance

caliberfocus certification

Ready to transform your business? Contact us today.

Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.