Google’s MedPaLM emphasizes human doctors in medical AI

[ad_1]

gnatiev/Getty Images

Most applications of AI in medicine have failed to use language, broadly speaking, a fact Google and its DeepMind unit addressed in a paper published in the prestigious science journal Nature on Monday.

Their invention, MedPaLM, is a large language model like ChatGPT that is tuned to answer questions from a variety of medical datasets, including a brand new one invented by Google that represents questions consumers ask about health on the Internet . That data set, HealthSearchQA, consists of “3,173 commonly searched consumer questions” that are “generated by a search engine,” such as “How bad is atrial fibrillation?”

Also: Google follows OpenAI in saying next to nothing about its new PaLM 2 AI program

The researchers used an increasingly important area of ​​AI research, rapid engineering, in which the program is given curated examples of the desired output in its input.

In case you were wondering, the MedPaLM program follows the recent trend by Google and OpenAI of hiding the technical details of the program, rather than specifying them as is standard practice in machine learning AI.

Google’s MedPaLM is based on a version of its PaLM language model, Flan-PALM, with the help of human prompt engineering.

Google/DeepMind

The MedPaLM program took a big step forward when it answered HealthSearchQA questions, as judged by a panel of human physicians. The percentage of times his predictions agreed with the medical consensus surpassed the 61.9% score for a variant of Google’s PaLM language model, reaching 92.6%, just short of the average human physician. 92.9%.

However, when a group of medically experienced laypeople were asked to rate how well MedPaLM answered the question, namely: “Does it allow them to [consumers] to draw a conclusion,” 80.3 percent of the time MedPaLM was helpful, versus 91.1 percent of responses from human physicians. The researchers believe this means that “much work remains to be done to approximate the quality of results provided by human clinicians”.

Plus: 7 advanced speed writing tips you need to know

The paper, “Large Language Models Encode Clinical Knowledge,” by lead author Karan Singhal of Google and colleagues, focuses on using so-called rapid engineering to make MedPaLM better than other large language models.

MedPaLM is a derivative of PaLM-powered question-and-answer pairs provided by five physicians in the US and UK. Those question-answer pairs, just 65 examples, were used to train MedPaLM through a series of immediate engineering strategies.

The typical way to refine a large language model like PaLM, or OpenAI’s GPT-3, is to feed it “with large amounts of in-domain data,” note Singhal and the team, “an approach that is challenging here given the scarcity of medical data.” Instead, for MedPaLM, they are based on three suggestion strategies.

MedPaLM significantly outperforms Flan-PaLM in human evaluations, although it still falls short of the capabilities of human physicians.

Google/DeepMind

The suggestion is the practice of improving model performance “through a handful of demo examples coded as prompt text into the input context”. The three suggestion approaches are prompts of a few strokes, “describing the activity through text-based demonstrations”; the so-called thought chain prompt, which involves “increasing each example a few strokes into the prompt with step-by-step breakdown and a coherent series of intermediate reasoning steps toward the final answer”; and “self-consistency suggestion”, in which different program results are sampled and a majority vote indicates the right answer.

Plus: Six skills needed to become an AI-ready engineer

MedPaLM’s high score, they write, shows that “Instruction prompt optimization is a data- and metric-efficient alignment technique that is useful for improving accuracy, factuality, consistency, safety, harm, and bias factors.” , helping to close the gap with clinical experts and bring these models closer to real-world clinical applications.”

However, “these models are not at the clinical expert level on many clinically important axes,” they conclude. Singhal and team suggest expanding the use of expert human participation.

“The number of model responses evaluated and the pool of clinicians and laypersons evaluating them were limited, as our results relied on only one clinician or layperson evaluating each response,” they note. “This could be mitigated by the inclusion of a considerably larger and intentionally diverse pool of human raters.”

Plus: How to write better ChatGPT prompts

Despite the MedPaLM deficit, Singhal and team conclude, “Our findings suggest that strong performance in answering medical questions may be an emerging skill of LLMs combined with effective instruction fine-tuning.”

Sources

1/ https://Google.com/

2/ https://www.zdnet.com/article/googles-medpalm-emphasizes-human-clinicians-in-medical-ai/

The mention sources can contact us to remove/changing this article

[ad_2]

Leave a Reply

Your email address will not be published. Required fields are marked *

Related Posts