Define the population before paying for healthcare AI outcomes

Healthcare leaders should agree on the population and measurement rules before tying digital health payments to outcomes. A result among patients who complete a program is not the same as the result across everyone the program was intended to serve.

PHTI’s purchasing survey reports that 68% of employers and 55% of health plans surveyed use performance-based contracts. It also finds that 47% of purchasers say fewer than a quarter of their eligible members enroll in digital health solutions.

It is important to make the population definition part of the evidence discussion from the beginning. For a payer, that definition affects whether an intervention can improve results across its membership. For a provider, it helps clarify which patients receive actionable support and which remain outside the workflow.

Start with the population

A buyer should be able to follow patients from eligibility through enrollment, continued participation, and outcome measurement. At each stage, the report should show how many people remain, why others leave or drop out, and whether their clinical information is still available. These groups answer distinct questions about reach, participation, and observed results.

Consider a hypothetical six-month program with 1,000 eligible patients. Two hundred enroll, 160 have the required follow-up measurements, and 80 meet a defined improvement threshold. Documented improvement is 50% among patients with follow-up data, 40% among all enrollees, and 8% across the eligible population. Each percentage describes the same program using a different denominator.


The 40 enrollees without follow-up measurements have unknown outcomes. The 8% figure describes documented improvement across the eligible population; it does not estimate the intervention’s causal effect. A useful report preserves both distinctions and explains what additional information would change the interpretation.

As a buyer, it is important to look at vendors who do not just focus on the narrowest possible definition of the metric, but are also able to take responsibility to increase the “enrolled” population to maximize engagement and in turn, the efficacy of the program.

Show how data changes care

For an AI-enabled intervention, the next question is what happens after the system produces an insight. A team needs to know whether the information arrived in time, reached the appropriate clinician, and led to a completed action. An alert count provides limited insight into those steps.

Take post-discharge support as an example. The evaluation could examine the completeness and timing of discharge information, the proportion of appropriate follow-up tasks completed, and a separately defined patient outcome. Those measures let an organization examine whether a problem arose in the data, in the clinical workflow, or later in the patient’s care.

Care-process measures and clinical outcomes should remain visible separately. Better task completion can demonstrate an operational change, but establishing improvement in patient health requires an appropriate outcome measure and evaluation. Connecting the two helps leaders decide what to fix without assuming that every workflow gain has already produced a clinical benefit.

Economic evaluation needs a further step. If automation saves staff time, report how that capacity is used and whether costs or service capacity change. When estimating financial value, account for the program’s fees, integration, training, human review, and additional care activity. Clinical improvement and net financial savings are distinct questions.

Agree on the comparison before the result

Buyers and vendors should document the eligibility rules, measurement window, comparison approach, and handling of missing information before reviewing results. A comparison with patients who never enrolled may need an assessment of why those patients did not participate, and a simple YoY comparison with last year’s spending also needs to consider other changes in care and membership.

A randomized rollout can help evaluate the effect of offering an intervention. Otherwise, an observational comparison should explain how the groups differ and what uncertainty remains. The design should match the claim: evidence of uptake, evidence of a care-process change, and evidence of an incremental outcome each require an appropriate analysis.

Reproducibility also depends on data access. The parties should agree on a common eligibility roster, permitted exclusions, the relevant clinical and claims records, and the delay before those records are sufficiently complete. Someone other than the team producing the headline result should be able to reconstruct the key measures from the agreed data.

These measurement choices affect implementation as well as contracting. They help buyers recognize when to improve outreach, repair a data connection, change a care-team workflow, or reconsider a financial assumption. They also give vendors a clearer way to demonstrate value and explain the limits of the evidence.

Before signing an outcome-linked contract, I would ask the vendor to show the result for the measured group alongside enrollment, missing follow-up data, and the intended population. Then I would ask what comparison supports the claimed improvement. Those answers make it easier to decide whether the program deserves broader deployment.

Comments

Popular posts from this blog

Why Walmart's Clinic Shutdown Was the Best Thing for its Healthcare Ambitions

The Great Migration: Why Health Systems Are Abandoning the Hospital Bed

Healthcare's attempts to "meet patients where they are"