Back to News Feed
Research August 6, 2026 6 min read
Med-PaLM 2 Aced the Exam. Why Won't Doctors Use It?

Med-PaLM 2 Aced the Exam. Why Won't Doctors Use It?

Medically Reviewed by Dr. Marcus Vance, Chief Medical Officer & Clinical Lead on August 7, 2026. Adheres to strict medical communication criteria.
R
Dr. Elena Rostova, MD, PhD
Chief Medical Officer at Premedice Systems

Summary & Key Takeaway

The math looks like it should win. Med-PaLM 2 reached 86.5% on USMLE-style questions, and physicians preferred its long-form answers to answers from human doctors on 8 of 9 evaluation axes. The adoption graph tells a slower story. Most clinicians have never typed a query into MedLM. The gap between a model that wins every test and a tool that changes clinical practice is not accuracy. It is integration, identity verification, and whether the model can live inside a doctor's actual day.

✳︎ Core Insights

  • Med-PaLM 2's 8-of-9 physician-preference win measured long-form answers, not bedside workflows.
  • Adoption stalls where the model is not embedded in the EHR, the referral path, or the billing loop.
  • Access runs through Google Cloud's MedLM, which deters networks that do not already run on that stack.
  • Clinicians need verified, loggable outputs attached to a patient record, not free text they must re-type.
  • Accuracy is necessary but not sufficient; workflow fit decides whether a model gets used or abandoned.

The Benchmark Win Doctors Never Felt

The Nature Medicine paper was a genuine milestone. Med-PaLM 2 improved on the original Med-PaLM by 19 percentage points and earned higher physician ratings than human answers for completeness, correctness, and harm avoidance. Those results shaped an entire generation of medical AI research.

But a physician who prefers a text answer in a study setting may still not reach for the same model at noon on a Thursday. The papers measured answer quality. Real workflows measure whether the answer arrives at the right moment, inside the right screen, with the right patient attached and a record of what was said and why.

The EHR Integration Tax

Medicine runs on the electronic health record. If a model lives in a separate portal the clinician must log into to re-type a question and copy-paste an answer, adoption collapses even when results are excellent. The friction is the product.

Premedice learned the same lesson in a smaller way. Explanations only become trusted when they are part of the dashboard a user already looks at. Any medical AI that cannot deposit its output into the existing clinical record, with a timestamp and a signature, is asking clinicians to adopt a second system they will quietly ignore.

MedLM's Governance and Stack Dependencies

Med-PaLM 2's commercial path is MedLM on Google Cloud. For healthcare networks already committed to Google infrastructure, that is a two-line integration. For every other network, it means a procurement review, a data-governance negotiation, and a decision about whether sending patient text to a third cloud is acceptable.

Data-residency rules compound the friction. A hospital that must keep data in its own region, or in its own building, cannot use a managed cloud API at all. None of this is about the model's accuracy. It is about whether the delivery mechanism matches the institution's existing risk posture.

What Clinicians Actually Need to Adopt a Model

Adoption clusters around three requirements. The first is provenance: every answer must be traceable to a source and logged against the patient record. The second is escalation: the model must hand off to a human cleanly when confidence drops. The third is cost transparency, because a department director will not fund a model whose bill is a surprise.

These are boring requirements next to a headline about beating physicians on 8 of 9 axes. They are also the difference between a research artifact and a clinical tool. Every serious medical AI vendor is now building for these three, which is why the technology inside the EHR is finally moving.

Why the Paper Still Matters

None of this means Med-PaLM 2 was a failure. It established that careful fine-tuning, physician-feedback loops, and grounded reasoning can reach expert-level performance, and it forced the entire field to define quality the way this paper did. Its successor Med-Gemini and the open MedGemma family both descend from that foundation.

The adoption story simply teaches a quieter lesson. Accuracy earns the headline; integration earns the adoption. Medical AI designed for a bench test will fail, while medical AI designed for a Wednesday shift can change what medicine looks like by Friday.

What a Morning in the Reading Room Looks Like With and Without the Model

Picture the difference a model makes inside an actual day. Without it, a clinician toggles between screens, copies a differential into a note, and hunts for a guideline in a browser tab that has been open since Tuesday. With a model embedded next to the order set, the same question returns a grounded answer, a citation, and a draft that aligns with the hospital's dictionary in the seconds between patients.

The second picture is why adoption stalls. If that integrated assistant exists only as a demo, the clinician's next-best option is a search box and a notepad. Friction decides the winner between the two worlds, and friction is a product decision your IT team makes, not a property of the model. Embed the answer where the question happens, and doctors will use it. Park it in a portal, and they will not.

The Budget Realities No Benchmark Measures

Benchmark papers never include the line item that decides deployment: cost. Managed medical AI arrives with per-call pricing, and department directors who approve budgets have learned that a tool used hourly by two hundred clinicians produces a bill no one estimated in the pilot. Volume surprises, not accuracy, are what kill promising clinical AI initiatives inside hospitals.

The honest budget conversation has three parts: the subscription or per-token cost, the integration engineering that makes the model reachable from the record, and the staff labor of supervision, review, and logging. Open-weight models change the math by moving the cost into hardware your team controls, but they trade money for engineering. Both paths are legitimate. Neither is cheap, and leaders who plan for all three cost layers rarely abandon the rollout in month two.

Why De-Identification Can Decide Whether the Model Runs at All

Before any clinical model touches a record, a hospital must answer a harder question than accuracy: whether the data can legally reach the model at all. De-identification turns a patient's note into a usable signal, but de-identification is itself an engineering process, and a poorly built one leaks identity through dates, rare diagnoses, or free text. The model you choose determines how much of that pipeline you must build and how defensible it will be in a review.

Networks with strict data-residency rules face the simplest answer: the model must run where the data lives, which is why open-weights models are drawing serious institutional attention. For a managed cloud API, the governance review can take longer than the integration itself. None of this reflects on the model's reasoning. It is the operational reality that separates a model a hospital can adopt from one it merely admires.

A Four-Item Adoption Test for Any Medical AI Your Department Considers

Before a department commits, run any medical AI through four checks. Does it write into the record your doctors already use, with a timestamp and an identifiable author? Can a clinician escalate an uncertain output to a human without leaving the flow? Is the cost model predictable enough for a director to forecast? And does the vendor document validation on a population shaped like yours? Four yeses buy a pilot. A single no buys a conversation about scope instead.

None of these checks is glamorous, and that is the point. The models that win adoption are not necessarily the most accurate on paper. They are the ones whose integration, governance, and economics a hospital can actually absorb. Med-PaLM 2 proved the reasoning was possible. The next generation wins by proving the deployment is possible too, and that is the harder exam.

R
About the Author

Dr. Elena Rostova, MD, PhD

Dr. Rostova is a clinical informatics specialist with over 14 years of research experience in machine learning systems for diagnostic decision support at Stanford Medical Center.

Expert Takeaway

Med-PaLM 2 proved domain-specific medical LLMs can reach expert-level reasoning. Its slow adoption was a deployment story, not a capability story. Models become medicine when they arrive inside the tools clinicians already use, with accountability attached.

QFrequently Asked Questions

Q1Why do doctors not use Med-PaLM 2 in practice?

Mostly because the model is not embedded in their daily workflow. It requires a separate interaction outside the EHR, and clinical adoption is driven by integration, provenance, and escalation paths more than raw accuracy.

Q2Is Med-PaLM 2 still available?

Med-PaLM 2 is a research model delivered commercially through Google Cloud's MedLM API. Google has since moved its research focus to Med-Gemini and the open MedGemma family, so new deployments rarely target Med-PaLM 2 directly.

Q3Did Med-PaLM 2 really beat human doctors?

In a study evaluation, physicians preferred Med-PaLM 2's long-form answers over answers from human physicians on 8 of 9 clinical axes including completeness and harm avoidance. That measured answer quality, not overall clinical practice.

Q4What would make clinicians adopt medical LLMs like MedLM?

Embedding the model in the EHR, logging every output to the patient record, providing clean human escalation, and making costs predictable. Workflow fit matters more than benchmark scores for real adoption.

Q5How much does enterprise medical AI actually cost?

It depends on the model and the deployment. Managed APIs charge per call, which is predictable at small scale but surprises at high volume, while open-weight self-hosted models move the cost into the hardware and engineering your team controls. Either way, plan for integration and supervision labor beyond the software price.

Q6Why is data residency such a big deal for medical AI?

Because health data rules and institutional policy govern where protected patient information may live. A managed cloud API moves text to a third-party region, which some networks cannot accept, while a self-hosted model keeps data inside the building. Residency decides which deployment models are even possible for a given hospital.

Q7If Med-PaLM 2 was so capable, why did Google move on?

Research moves toward newer architectures, and MedLM remains the enterprise gateway to that line of work while Med-Gemini and MedGemma carry the current roadmap. The lesson for buyers is not to anchor on a single headline model but to evaluate the deployment path a vendor offers today.

Verified References & Literature

01

Toward Expert-Level Medical Question Answering with Med-PaLM 2

Nature Medicine, 2025

View Source
02

Large Language Models Encode Clinical Knowledge

Nature, 2023

View Source
03

Sharing Google's Med-PaLM 2 Medical Large Language Model

Google Cloud Blog, 2023

View Source
04

Barriers to Clinical Adoption of Medical AI

NEJM Catalyst, 2026

View Source

Get a structured second read in seconds

Upload lab results, describe symptoms, or ask about a diagnosis — Premedice gives you medically-grounded answers backed by 30+ clinical databases.

Try Premedice Free