Back to News Feed
Research August 29, 2026 7 min read
Med-PaLM 2 at 86.5% on MedQA: Why Doctors Still Don't Trust It (2026)

Med-PaLM 2 at 86.5% on MedQA: Why Doctors Still Don't Trust It (2026)

Medically Reviewed by Dr. Marcus Vance, Chief Medical Officer & Clinical Lead on August 29, 2026. Adheres to strict medical communication criteria.
R
Dr. Elena Rostova, MD, PhD
Chief Medical Officer at Premedice Systems

Summary & Key Takeaway

Medical licensing exams are among the hardest knowledge tests in any profession. In 2023, Google's Med-PaLM 2 became the first large language model to reach expert-level performance on the MedQA benchmark of USMLE-style questions, scoring 86.5%. Physicians who evaluated its long-form answers preferred Med-PaLM 2 over answers from actual human physicians on 8 out of 9 clinical evaluation axes. This wasn't a generic chatbot fine-tuned with medical keywords. It was a purpose-built clinical reasoning engine, trained through reinforcement learning from physician feedback and ensemble refinement techniques that forced the model to ground every claim in verifiable medical literature. This article traces the architecture, benchmarks, and real-world impact of Med-PaLM 2 and its commercial evolution into Google Cloud's MedLM.

?? Core Insights

  • Med-PaLM 2 scored 86.5% on MedQA, a 19-percentage-point improvement over the original Med-PaLM's 67.6%.
  • Physicians preferred Med-PaLM 2 answers over human physician answers on 8 of 9 clinical evaluation axes including completeness, correctness, and harm avoidance.
  • 90.6% of Med-PaLM 2 outputs received a low-risk-of-harm rating in adversarial safety testing, up from 79.4% in the original model.
  • 92.6% of Med-PaLM 2 outputs aligned with established scientific and medical consensus, making it the most clinically grounded LLM at its release.
  • Med-PaLM 2 powers MedLM, which became available to Google Cloud healthcare customers in late 2023 for question answering and clinical summarization tasks.

From Med-PaLM to Med-PaLM 2: The 19-Point Leap on MedQA

Google's journey to expert-level medical AI began with Med-PaLM, a 540-billion-parameter model that scored 67.6% on MedQA. Med-PaLM 2 improved that result to 86.5% - a 19-percentage-point jump achieved in roughly a year. The improvement came from three techniques combined: instruction fine-tuning on curated medical knowledge, chain-of-retrieval reasoning, and reinforcement learning from human feedback supplied by physicians.

The result changed how the field measured medical language models. Rather than asking whether an AI could pass a medical exam, researchers began asking whether its answers were preferred by clinicians. On that measure, Med-PaLM 2 won: physicians rated its long-form answers higher than answers from human physicians on 8 of 9 clinical axes, including completeness, reasoning, and alignment with medical consensus.

How Physician Feedback Shaped Clinical Safety

The core innovation of Med-PaLM 2 was the feedback loop. Google worked with physicians to evaluate model outputs on judgment axes such as correctness, harm avoidance, and alignment with scientific consensus. Those evaluations became training signal, steering the model away from plausible but dangerous answers.

The safety improvements were measurable. Adversarial safety testing classified 90.6% of Med-PaLM 2 outputs as low-risk-of-harm, up from 79.4% for the original model, and 92.6% of outputs aligned with established medical consensus. These are the numbers that matter for healthcare - a model that is well-calibrated about what it does not know is safer in clinical deployment than one that is merely more accurate on a benchmark.

Med-PaLM 2 versus Med-PaLM M: Text-Only Reasoning and the Multimodal Frontier

Med-PaLM 2 is a text-only model. Its capabilities are centered on clinical question answering and long-form medical text. For multimodal work, Google developed Med-PaLM M as a research prototype trained across text, imaging, and other clinical data modalities.

Med-PaLM M demonstrated the direction of travel, but it was never broadly released. Google's public guidance to developers needing medical image interpretation today points to MedGemma, the open multimodal family, rather than Med-PaLM M. The distinction matters when choosing a deployment pathway: text-only clinical reasoning versus imaging-capable models require different clearance, hardware, and compliance reviews.

Through MedLM into Enterprise Healthcare Production

Med-PaLM 2's commercial path runs through MedLM, Google Cloud's managed suite for healthcare. Made available to select cloud customers in late 2023, MedLM exposes the underlying clinical reasoning as an API for question answering and clinical summarization. That is how a research breakthrough became a product hospitals could integrate into EHR workflows.

MedLM inherits Med-PaLM 2's trade-offs. Access is controlled by Google Cloud, the model runs on third-party infrastructure, and data governance rests on the customer relationship with the platform. For healthcare organizations that have already cleared Google Cloud, MedLM is the lowest-friction route to production-grade clinical reasoning; for organizations that must keep data in their own infrastructure, it is not an option at all.

Why Medical LLMs Need Clinical Grounding, Not Just More Data

General-purpose chatbots can reach Med-PaLM 2's benchmark scores by scaling. But benchmark parity does not equal clinical reliability. The differentiator is grounding - the ability to trace a claim to an authoritative medical source and to decline confidently when evidence is absent.

Med-PaLM 2 introduced this discipline through its training methodology, and it is the reason physician preference scores diverged from raw accuracy. For any organization deploying medical AI, the lesson generalizes: demand explainability, citation behavior, and calibrated uncertainty from the model, and treat ungrounded fluency as a liability rather than a feature.

How the MedQA Benchmark Became Medical AI's Yardstick

MedQA is a multiple-choice dataset built from USMLE-style questions, and it earned outsize influence because it was the first scalable stand-in for the licensing exam. A model scoring 67.6% was already passing; at 86.5% it outperforms the average human test-taker on the same material. But the benchmark measures recalled clinical knowledge, not performance in a real clinic, where context, uncertainty, and consequences change everything.

That distinction explains the gap between the headline number and the deployment reality. A multiple-choice score says what a model knows and how well it can hold that knowledge together. It says little about whether the model can read a messy chart, admit what it does not know, or stay steady under pressure. Benchmarks are scoreboards for researchers, and clinicians rightly treat them as necessary but nowhere near sufficient.

Where Med-PaLM 2 Falls Short in Real Clinical Work

For all its gains, Med-PaLM 2 is a text-based knowledge engine, not a reasoning partner in a busy clinic. It cannot fold an examination finding into what a patient's tone suggests, can only reason about what is written down, and will produce a confident answer even when the input is ambiguous. That is why the field has moved toward multimodal models with better calibration, which is exactly the direction clinical practice demands.

It also inherits the general weaknesses of its era: sensitivity to how a question is phrased, a tendency to favor fluent answers, and no way to promise it has considered everything relevant. Hospitals testing such models found value in drafting summaries and answering well-defined questions, and they kept humans at the controls for anything that changes a treatment. Knowing the limits is half of using a tool like this well.

Who Actually Benefits From a Model Like Med-PaLM 2

The clearest beneficiaries are not patients asking for a diagnosis. They are the organizations that generate clinical text at volume: hospitals summarizing charts, research teams indexing literature, and developers building assistants that answer clearly bounded medical questions. In those settings, a well-grounded model turns hours of reading into minutes of confirmation and frees clinicians for direct patient care.

Patients benefit indirectly and later. When strong clinical reasoning lives inside an EHR assistant, office messages get drafted faster, summaries read more completely, and triage answers stay on message. That is the honest value chain: the model amplifies the system's paperwork, and the system returns the saved time to the patient. Direct-to-patient diagnosis is a different product entirely, one that needs a different safety regime.

Choosing Between a Managed Medical API and a Self-Hosted Model

If you build on clinical reasoning, the first fork is deployment, not accuracy. A managed API keeps the model on third-party infrastructure and ties you to a vendor's roadmap, which makes it fast to start and simplest when clearance and privacy agreements already exist. Self-hosted open models keep data on your own hardware, which matters for institutions with strict residency rules, but they shift engineering and monitoring onto your own team.

There is no universally correct choice, only a fit. Favor the managed route when speed and support matter more than data geography, and favor self-hosting when your compliance profile demands the records never leave. Ask the same questions either way: who sees the inputs, how are outputs verified, and what happens when the model is wrong. The best deployment is the one your institution can actually sustain.

R
About the Author

Dr. Elena Rostova, MD, PhD

Dr. Rostova is a clinical informatics specialist with over 14 years of research experience in machine learning systems for diagnostic decision support at Stanford Medical Center.

Expert Takeaway

Med-PaLM 2 demonstrated that domain-specific fine-tuning with physician feedback loops can close the gap between general-purpose AI and clinical expert performance. Its successor, Med-Gemini, has since surpassed it at 91.1% on MedQA.

QFrequently Asked Questions

Q1Is Med-PaLM 2 available to the public?

No. Med-PaLM 2 is a research model that was made available to select Google Cloud customers through Vertex AI. It is not open-source. Its capabilities are accessible through Google's MedLM enterprise API.

Q2How does Med-PaLM 2 differ from the original Med-PaLM?

Med-PaLM 2 improved accuracy from 67.6% to 86.5% on MedQA through ensemble refinement, chain-of-retrieval reasoning, and physician feedback-based reinforcement learning. It also achieved a 90.6% low-harm rating compared to the original's 79.4%.

Q3Can Med-PaLM 2 process medical images?

The standard Med-PaLM 2 is text-only. Google developed Med-PaLM M as a multimodal research variant, but it is not widely available. For medical imaging, Google directs developers to MedGemma.

Q4What is MedLM and how does it relate to Med-PaLM 2?

MedLM is Google Cloud's enterprise healthcare AI product built on Med-PaLM 2. It offers question answering and clinical summarization through a managed API, designed for integration with hospital EHR systems and clinical workflows.

Q5Has anything surpassed Med-PaLM 2's accuracy?

Yes. Google's Med-Gemini scored 91.1% on MedQA in 2024, and general-purpose models like Claude Fable 5 and Gemini 3.5 Flash have since reached comparable levels. However, these general-purpose models lack the clinical grounding techniques that made Med-PaLM 2 reliable for medical applications.

Q6Does a score of 86.5% mean Med-PaLM 2 is a reliable doctor?

No. The score measures knowledge on multiple-choice questions, not bedside performance. Real medicine involves ambiguous histories, examining patients, and managing consequences, where the model has no equal footing. Use it as a strong reference tool and keep clinical decisions with the clinician.

Q7Do health apps need a model trained specifically on medicine?

Specialized models bring better grounding and lower hallucination rates on clinical content, but they are not the only factor. Prompt design, output constraints, and human review matter just as much. For any app that interprets health information, check how it verifies answers and what it does when it is unsure.

Verified References & Literature

01

Toward Expert-Level Medical Question Answering with Large Language Models (Med-PaLM 2)

Nature Medicine, 2025

View Source
02

Large Language Models Encode Clinical Knowledge (Med-PaLM)

Nature, 2023

View Source
03

Med-PaLM 2 Statistics 2026: Healthcare AI Performance and Adoption

AboutChromebooks Research, 2025

View Source
04

Sharing Google's Med-PaLM 2 Medical Large Language Model

Google Cloud Blog, 2023

View Source

Get a structured second read in seconds

Upload lab results, describe symptoms, or ask about a diagnosis — Premedice gives you medically-grounded answers backed by 30+ clinical databases.