Skip to content
Buying health AI? Here's a four-part framework

Buying health AI? Here's a four-part framework

A field guide for the people buying healthcare AI and the people building it
7 min read
In this thoughtful piece, Included Health's Akin Oyalowo and Nupur Srivastava offer a practical framework for evaluating healthcare AI, whether you're a health plan, employer, provider organization, or digital health company.

"You're moving fast, and honestly, it makes me nervous."

A benefits leader at one of the country's largest public healthcare purchasers told us this recently, after a few days of presentations about how AI would reshape care. He wasn't an AI skeptic; he was sold on the destination. He just wanted more detail on how to get there: the timing, the guardrails, and some assurance that implementing AI wouldn't break the services his members rely on today.

As leaders at a company with an AI solution in the market, we've heard some version of this question dozens of times. His concern speaks to a common misconception among healthcare buyers. The main risk with healthcare AI isn't the speed of adoption. It's the direction.

Consumers are increasingly turning to general-purpose AI for everyday health concerns, while 75% of U.S. health systems are already running at least one AI application. But what are we racing toward? Consumers are following the path of least resistance. Providers, payers, and the operators supplying them are gravitating toward use cases where the financial pull is strongest and the ROI is fastest, not necessarily where patients and purchasers stand to gain the most.

Will this create efficiency? How quickly will this pay off? These aren't the most important questions. The question that matters most is: What's the net impact of this AI within the payment model, workflow, and population where it's deployed?

We've been on both sides of this question. One of us (Nupur) is an engineer and executive who has spent the last decade building and shipping clinical AI. The other (Akin) is a gastroenterologist who spent years at the bedside before advising health systems and carriers on strategy, and now sits across the table from the employers and health plans deciding where to allocate budget.

This vantage point has taught us that everyone at the table, buyer or builder, needs to get sharper on how we evaluate AI solutions. As a buyer, you need a framework to separate tools that genuinely move outcomes from tools that just move money around. As a builder, you need to know which questions the serious buyers will ask before they sign, and make sure you can answer them for yourself first.

In either case, we've found the same four principles apply.

Included Health is designed to treat you better: A modern experience designed to treat people better — mind, body, and wallet.

Care for everything, anytime, no matter where you are — it’s all Included.  

1. Start with the mechanism, not the model

Every vendor can describe their model. Far fewer can describe a mechanism. "Improves navigation" and "supports clinicians" are categories, not explanations.

The distinction matters because a precise causal chain is a verifiable one. It tells you exactly what to measure, when the solution isn't working, and who absorbs the harm when it fails. AI tools designed to improve diagnosis and early detection are a strong example of a narrow, defined mechanism: read the test results, surface the hidden risk, and get the patient treatment sooner. In recent cardiology trials, for example, that chain produced measurable improvements in accuracy and time to treatment.

Contrast that with the vague claims we routinely encounter. At a recent roundtable of enterprise employers evaluating vendors, few could describe which clinical bottleneck their AI was supposed to fix. If the mechanism is vague, the value claim is too, and accountability vanishes.

For builders: if you can't describe your mechanism in one sentence, you haven't finished thinking. A clear mechanism isolates failure, proves success, and shows buyers exactly what they're purchasing.

If you're buying →

  • Require a one-sentence mechanism statement and a simple flow diagram in every RFP and demo. If a vendor can't describe it, that's your answer.
  • Ask vendors to map failure modes. How does the system produce false positives and false negatives, and who absorbs the harm in each case: patient, clinician, health plan, or employer?
  • Insist on results from your own population, not just the vendor's benchmark. (Among hospitals using predictive models, less than half check for bias locally.)

2. Be clear on financial incentives and objectives

Many AI deployments (like coding optimization) are logical responses to existing fee-for-service incentives. That's why "AI creates efficiency" is an incomplete claim. Efficiency for whom, captured where? Who benefits financially if this AI works exactly as intended? It's a simple question that providers, insurers, and healthcare purchasers often overlook.

Though the market is focused on time to value, not every AI investment pays back within a budget year. Promising use cases — such as clinical decision support that reduces low-value procedures, or adherence tools that prevent avoidable hospitalizations — can be harder to scale precisely because the return is distributed and downstream, requiring workflow redesign and operational patience.

GLP-1 management shows how complicated the math can get. AI-enhanced GLP-1 programs show promise in boosting engagement and improving outcomes, yet a 2025 Milliman analysis found that for non-diabetic obesity, drug costs from increased adherence exceed medical savings, driving up total cost of care. The tension is real: Multiple large employers we work with believe in the clinical value of GLP-1s but cannot justify the near-term cost.

If you're buying →

  • Build a matrix to include in every AI RFP. Rows are stakeholders (plan, employer, system, clinician, member); columns are revenue, cost, and risk.
  • Pay for outcomes, not clicks. Tie fees to sustained outcomes and avoided costs, not engagement scores.
  • Separate near- and long-term ROI. Don't evaluate a two-year adherence play against a six-month coding-optimization tool on the same spreadsheet.

3. Assess workflows, not features

AI delivers the most value when it lives inside a workflow where someone knows how to act on the output and owns the consequences when it's wrong. As models improve, accuracy is no longer the primary reason a tool fails in the real world. Integration and human accountability are.

The solutions that work best show up at the right time and place, and are clear on what they ask the human to do (or not do). Ambient scribes embedded inside the patient visit deliver measurable reductions in documentation time and burnout. By contrast, a diagnostic aid in the emergency department showed no benefit in a randomized trial; the protocol required physicians to consult the tool during diagnostic workup but didn't specify exactly when.

"Human-in-the-loop" is an essential AI design principle, but it's not the fail-safe backstop we assume. Even skilled clinicians trained in AI literacy are susceptible to algorithmic errors. Thoughtful workflow design is equally important, and helps mitigate that risk.

For operators and builders, this is the uncomfortable truth: If your tool requires a separate login, doesn't show up where clinicians and patients already are, and lacks a named human accountable for acting on its outputs, it will likely be underused and ineffective.

If you're buying →

  • Ask for the click-by-click workflow, not a feature list. Where does the AI surface in the clinician's, member's, or benefits administrator's day-to-day?
  • Pressure-test the handoff. The most important moment in any AI workflow is when AI doesn't know the answer. Ask vendors to demonstrate the escalation path. How many clicks to reach a human, and does context carry over or does the member start from scratch?
  • Name the accountable human for every AI output. For each output type, define who is responsible for acting (or not acting) and make sure the action pathway is viable.

4. Move faster with better judgment, not more process

An HR leader at a major technology company recently expressed concern that employees were using ChatGPT to get answers about their benefits, including one employee who became convinced their vision partner was terrible as a result. "ChatGPT is not the answer," she told us, "but I can't say that until I have the right answer."

The gap between what purchasers want and what their organizations let them deploy is real and growing. In that gap, employees are making their own choices. A recent West Health-Gallup poll estimated that 14 million Americans have skipped a doctor's visit after consulting AI. Meanwhile, employer-sanctioned tools sit in review cycles stretching into 2027.

One of our partners, a large financial services firm, endured months of delays until they reclassified their navigation assistant as a vendor service rather than an AI model. Changing the framing to fit the actual risk profile, without changing the governance standards, cleared the path to launch.

For builders: if your go-to-market depends on buyers who need 12 months of internal approvals, your deployment model is the bottleneck. The companies that ship fastest arrive with governance artifacts already built: local validation data, bias audits, and escalation documentation that a compliance officer can evaluate in a week, not a quarter.

If you're buying →

  • Tier your AI portfolio by risk, and govern each tier differently. Don't govern a summarization tool like a diagnostic one.
  • Set a "governance SLA" for new AI tools. If a low-risk tool takes months to approve, your process is the risk. Publish internal timelines by tier, and track whether you meet them.
  • Pool evaluation with peers. Commission a shared, independent vendor evaluation through a purchaser coalition, rather than each HR team running a custom, under-resourced review.

What it adds up to

These four principles aren't a checklist. They're a lens for understanding what's actually happening when AI is deployed, and what's at stake when it isn't evaluated carefully, whether you're buying or building.

Here's what that lens looks like in practice: A large employer recently discovered that some members were inaccurately claiming a diabetes diagnosis to access GLP-1 coverage. Instead of reaching for an AI-driven solution, they implemented a rules-based system and routed flagged claims to clinicians for review. About 5% of requests appropriately dropped off. No black box denied anyone care. The program got safer and more affordable.

That sums up the essence of this framework. The goal isn't to move faster or slow down. It's to know exactly what you're deploying, why, and who's accountable when it matters most.

Nupur Srivastava is Chief Operating Officer at Included Health. Akin Oyalowo, MD, MS, is Medical Director of Clinical Affairs at Included Health and a board-certified gastroenterologist.

Share this article

Spread the word