Trail documentation
Best AI models for healthcare in 2026: a developer’s guide
Compare Claude, GPT, Gemini and open-weight medical models for health apps by capability, HIPAA business associate agreement route, data retention and regulatory fit.
What “best” means for a health app
For a healthcare product, the most capable model on a leaderboard is not automatically the right choice. The questions that decide it are more practical: can you get a business associate agreement (BAA) for the way you plan to call the model, what data retention does that require, which features fall outside the agreement, and how will a clinician check the output?
This guide covers the frontier model families and the main open-weight medical models as of September 28, 2026. Model line-ups change every few months, so treat the vendor pages linked in each section as the source of truth.
Anthropic Claude
Anthropic’s current models are Claude Fable 5.1 for the most demanding reasoning and long-running agent work, Claude Opus 5.5, which its documentation recommends as the starting point for most workloads, and the smaller Sonnet and Haiku models for speed and cost. The models are available through Anthropic’s API and through Amazon Bedrock, Google Cloud and Microsoft Foundry.
Anthropic offers a BAA for its first-party API and for HIPAA-ready Claude Enterprise. Some features are excluded from API coverage, including the Batch and Files APIs, code execution and web fetch, and some models require 30-day retention rather than zero data retention. In January 2026 it launched Claude for Healthcare, with connectors for the CMS Coverage Database, ICD-10, the NPI Registry and PubMed.
Sources: Claude models overview · Anthropic BAA scope · Claude for Healthcare
OpenAI GPT
OpenAI’s current API line-up is GPT-6 Astra, described as its most capable model, GPT-6 Sol and the low-cost GPT-6 Luna. OpenAI reviews API BAA requests case by case through baa@openai.com; HIPAA eligibility for the API depends on the account being provisioned with modified retention, and some products, such as Codex cloud, are excluded.
OpenAI also publishes HealthBench, a benchmark built with 262 physicians across 5,000 conversations, and in January 2026 launched ChatGPT for Healthcare for organizations. HealthBench is useful, but note that when a vendor runs a benchmark on competitors’ models, those scores are vendor-reported rather than independent.
Sources: OpenAI models · OpenAI BAA process · OpenAI HIPAA-eligible products · HealthBench · OpenAI for Healthcare
Google Gemini and MedGemma
Google’s stable Gemini API models are currently in the Flash family, led by Gemini 3.8 Flash, with Gemini 3.1 Pro in preview. There is an important split: the Gemini Developer API terms prohibit use in clinical practice or to provide medical advice. Health products should use Gemini through Google Cloud’s Gemini Enterprise Agent Platform, formerly Vertex AI, which is covered by Google Cloud’s HIPAA BAA.
Google also releases MedGemma, open-weight medical models including a multimodal 4B version and 27B text and multimodal versions, under Health AI Developer Foundations terms, plus MedASR for medical speech-to-text. Google states MedGemma is not yet clinical-grade; it is a foundation to validate and adapt, not a finished product.
Sources: Gemini API models · Gemini API terms · Google Cloud HIPAA · MedGemma · MedGemma 1.5 and MedASR
Open-weight models you can run yourself
Self-hosting keeps data inside your own infrastructure, which can simplify a privacy review but moves security, monitoring and updates onto your team. MedGemma is the most established medical option. General open-weight models such as OpenAI’s gpt-oss-120b, released under Apache 2.0, can also be fine-tuned or paired with retrieval over your own clinical content.
Sources: gpt-oss-120b · MedGemma
Our picks by use case
There is no single winner. These are sensible starting points to test against your own de-identified examples.
| Use case | Start with | BAA route | Watch out for |
|---|---|---|---|
| Note drafting and summarization | Claude Opus 5.5 or GPT-6 Sol | Anthropic or OpenAI API BAA; or Bedrock, Azure or Google Cloud | Retention settings and excluded features |
| Hardest multi-step reasoning | Claude Fable 5.1 or GPT-6 Astra | Same as above | Cost and latency |
| High-volume classification or routing | Claude Haiku, GPT-6 Luna or Gemini Flash via Google Cloud | Vendor or cloud BAA | Do not use the consumer Gemini API for clinical work |
| Imaging or speech research, on your own servers | MedGemma or MedASR | Self-hosted; your own safeguards | Not clinical-grade without validation |
Cloud routes to a BAA
Many teams reach these models through a cloud provider they already have an agreement with. Amazon Bedrock and Amazon Transcribe, including HealthScribe, are on AWS’s HIPAA-eligible services list. Microsoft includes its HIPAA BAA in its product terms and Data Protection Addendum for in-scope Azure services. Google Cloud’s BAA covers the Gemini Enterprise Agent Platform. Always confirm the specific service and region are in scope.
Sources: AWS HIPAA-eligible services · Microsoft Azure HIPAA · Google Cloud HIPAA
Regulation to plan for
Whether a feature is regulated depends on what it does, not which model powers it. The FDA finalized updated clinical decision support guidance in January 2026, and its guidance on AI-enabled device software functions remains in draft. ONC’s HTI-1 rule added transparency requirements for decision support in certified health IT, and the proposed HTI-5 rule would remove the AI “model card” requirements; it had not been finalized when we checked. The WHO’s guidance on large multi-modal models is a useful governance checklist.
Sources: FDA clinical decision support guidance · FDA AI-enabled device software draft guidance · HTI-5 proposed rule · WHO guidance on large multi-modal models
Shipping an app around the model
The model is one piece. A health app also needs sign-in, role-based access, audit logging, encryption and hosting covered by a BAA. Teams writing their own code often use a HIPAA platform such as Aptible, which also offers an LLM gateway covering Anthropic, OpenAI and Bedrock-hosted models under one BAA.
Teams without engineers can start from an app builder instead. Panaceum turns a plain-language description into a working app, builds apps that handle patient information on HIPAA scaffolding under a BAA, and currently limits those apps to synthetic test data while it is in early access.
Sources: Aptible pricing and LLM gateway · Panaceum
Questions about this guide
Which AI model is best for healthcare?
It depends on the task and your compliance route. For documentation and summarization, the current Claude and GPT flagship models are strong starting points; test them on your own de-identified cases.
Can I send patient data to the Gemini API?
Google’s Gemini Developer API terms prohibit clinical use. Use Gemini through Google Cloud under its HIPAA BAA instead.