Skip to content
FHIR Trail

Trail documentation

Best AI models for healthcare in 2026: a developer’s guide

Compare Claude, GPT, Gemini and open-weight medical models for health apps by capability, HIPAA business associate agreement route, data retention and regulatory fit.

What “best” means for a health app

For a healthcare product, the most capable model on a leaderboard is not automatically the right choice. The questions that decide it are more practical: can you get a business associate agreement (BAA) for the way you plan to call the model, what data retention does that require, which features fall outside the agreement, and how will a clinician check the output?

This guide covers the frontier model families and the main open-weight medical models as of September 28, 2026. Model line-ups change every few months, so treat the vendor pages linked in each section as the source of truth.

Anthropic Claude

Anthropic’s current models are Claude Fable 5.1 for the most demanding reasoning and long-running agent work, Claude Opus 5.5, which its documentation recommends as the starting point for most workloads, and the smaller Sonnet and Haiku models for speed and cost. The models are available through Anthropic’s API and through Amazon Bedrock, Google Cloud and Microsoft Foundry.

Anthropic offers a BAA for its first-party API and for HIPAA-ready Claude Enterprise. Some features are excluded from API coverage, including the Batch and Files APIs, code execution and web fetch, and some models require 30-day retention rather than zero data retention. In January 2026 it launched Claude for Healthcare, with connectors for the CMS Coverage Database, ICD-10, the NPI Registry and PubMed.

Sources: Claude models overview · Anthropic BAA scope · Claude for Healthcare

OpenAI GPT

OpenAI’s current API line-up is GPT-6 Astra, described as its most capable model, GPT-6 Sol and the low-cost GPT-6 Luna. OpenAI reviews API BAA requests case by case through baa@openai.com; HIPAA eligibility for the API depends on the account being provisioned with modified retention, and some products, such as Codex cloud, are excluded.

OpenAI also publishes HealthBench, a benchmark built with 262 physicians across 5,000 conversations, and in January 2026 launched ChatGPT for Healthcare for organizations. HealthBench is useful, but note that when a vendor runs a benchmark on competitors’ models, those scores are vendor-reported rather than independent.

Sources: OpenAI models · OpenAI BAA process · OpenAI HIPAA-eligible products · HealthBench · OpenAI for Healthcare

Google Gemini and MedGemma

Google’s stable Gemini API models are currently in the Flash family, led by Gemini 3.8 Flash, with Gemini 3.1 Pro in preview. There is an important split: the Gemini Developer API terms prohibit use in clinical practice or to provide medical advice. Health products should use Gemini through Google Cloud’s Gemini Enterprise Agent Platform, formerly Vertex AI, which is covered by Google Cloud’s HIPAA BAA.

Google also releases MedGemma, open-weight medical models including a multimodal 4B version and 27B text and multimodal versions, under Health AI Developer Foundations terms, plus MedASR for medical speech-to-text. Google states MedGemma is not yet clinical-grade; it is a foundation to validate and adapt, not a finished product.

Sources: Gemini API models · Gemini API terms · Google Cloud HIPAA · MedGemma · MedGemma 1.5 and MedASR

Open-weight models you can run yourself

Self-hosting keeps data inside your own infrastructure, which can simplify a privacy review but moves security, monitoring and updates onto your team. MedGemma is the most established medical option. General open-weight models such as OpenAI’s gpt-oss-120b, released under Apache 2.0, can also be fine-tuned or paired with retrieval over your own clinical content.

Sources: gpt-oss-120b · MedGemma

Our picks by use case

There is no single winner. These are sensible starting points to test against your own de-identified examples.

AI model starting points for health apps
Use caseStart withBAA routeWatch out for
Note drafting and summarizationClaude Opus 5.5 or GPT-6 SolAnthropic or OpenAI API BAA; or Bedrock, Azure or Google CloudRetention settings and excluded features
Hardest multi-step reasoningClaude Fable 5.1 or GPT-6 AstraSame as aboveCost and latency
High-volume classification or routingClaude Haiku, GPT-6 Luna or Gemini Flash via Google CloudVendor or cloud BAADo not use the consumer Gemini API for clinical work
Imaging or speech research, on your own serversMedGemma or MedASRSelf-hosted; your own safeguardsNot clinical-grade without validation

Cloud routes to a BAA

Many teams reach these models through a cloud provider they already have an agreement with. Amazon Bedrock and Amazon Transcribe, including HealthScribe, are on AWS’s HIPAA-eligible services list. Microsoft includes its HIPAA BAA in its product terms and Data Protection Addendum for in-scope Azure services. Google Cloud’s BAA covers the Gemini Enterprise Agent Platform. Always confirm the specific service and region are in scope.

Sources: AWS HIPAA-eligible services · Microsoft Azure HIPAA · Google Cloud HIPAA

Regulation to plan for

Whether a feature is regulated depends on what it does, not which model powers it. The FDA finalized updated clinical decision support guidance in January 2026, and its guidance on AI-enabled device software functions remains in draft. ONC’s HTI-1 rule added transparency requirements for decision support in certified health IT, and the proposed HTI-5 rule would remove the AI “model card” requirements; it had not been finalized when we checked. The WHO’s guidance on large multi-modal models is a useful governance checklist.

Sources: FDA clinical decision support guidance · FDA AI-enabled device software draft guidance · HTI-5 proposed rule · WHO guidance on large multi-modal models

Shipping an app around the model

The model is one piece. A health app also needs sign-in, role-based access, audit logging, encryption and hosting covered by a BAA. Teams writing their own code often use a HIPAA platform such as Aptible, which also offers an LLM gateway covering Anthropic, OpenAI and Bedrock-hosted models under one BAA.

Teams without engineers can start from an app builder instead. Panaceum turns a plain-language description into a working app, builds apps that handle patient information on HIPAA scaffolding under a BAA, and currently limits those apps to synthetic test data while it is in early access.

Sources: Aptible pricing and LLM gateway · Panaceum

Questions about this guide

Which AI model is best for healthcare?

It depends on the task and your compliance route. For documentation and summarization, the current Claude and GPT flagship models are strong starting points; test them on your own de-identified cases.

Can I send patient data to the Gemini API?

Google’s Gemini Developer API terms prohibit clinical use. Use Gemini through Google Cloud under its HIPAA BAA instead.

Browse all public guides →