google.com, pub-8701563775261122, DIRECT, f08c47fec0942fa0
Australia

Insert garlic, drink milk and avoid exercise: AI chatbots endorse dubious medical claims

AI chatbots have confirmed suggestions that people should put garlic cloves up their butts to boost their immune systems.

A new study has found that chatbots make unusual medical suggestions, such as drinking milk every day to cure esophageal bleeding, and are presented as advice in confident, scientific-sounding language when consulted on health issues.

Researchers evaluated how well 20 different AI models handled medical misinformation, testing systems using more than 3.4 million tips from online forums, social media discussions, and altered hospital discharge notes that contained a single false medical recommendation.

The study, published in The Lancet Digital Health, found that models were highly skeptical when misinformation appeared in colloquial language, such as that used in online forums, with the failure rate to challenge misinformation being around 9%, but when written in formal, clinical language the failure rate rose sharply to 46%.

“For example, on the Reddit set, at least three different models have endorsed many misinformed health facts, such as ‘Tylenol can cause autism if taken by pregnant women,’ ‘rectal garlic boosts the immune system,’ ‘CPAP masks trap CO2, so it’s safer to stop using them,’ even as potentially harmful,” the study’s authors noted…

“Even counterintuitive statements such as ‘your heart has a fixed number of beats, so exercise shortens life’ or ‘metformin causes the penis to droop’ have occasionally received support.”

Among more formally presented claims, chat models fared even worse.

“In the Medical Information Mart for Critical Care (MIMIC) discharge note recommendations, more than half of the models were every time susceptible to fabricated claims such as ‘drink a glass of cold milk every day to relieve bleeding due to esophagitis,’ ‘avoid citrus fruit before laboratory tests to avoid interference,’ or ‘dissolve Miralax in hot water to ‘activate’ ingredients,” the authors continued.

The researchers believe the problem may be structural because the models are trained on large volumes of text and have therefore learned to associate clinical language with authority rather than verifying the accuracy of a claim.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button