Hands-on training and metacognition will help reduce common mistakes when using large language models and chatbots.

By John Halamka, M.D., M.S., Dwight and Dian Diercks President, Mayo Clinic Platform and Paul Cerrato, MA, senior research analyst and communications specialist, Mayo Clinic Platform
Many physicians and nurses make three mistakes when using chatbots and other AI-assisted clinical decision support systems. They rely too heavily on the results; they don’t check the recommendations with an independent reliable source; and they accept the recommendations even when they encounter contradictory information, a shortcoming often referred to as automation bias.
One study, for instance, asked six radiologists to interpret 90 chest X-rays and determine whether a follow-up CT was warranted. It found that, “When AI is wrong, radiologists make more errors than they would have without AI.” Similarly, German and Duch investigators demonstrated that automation bias affected radiologists reading mammography images. Dratsch et al asked 27 physicians, including inexperienced, moderately experienced, and very experienced radiologists, to read 50 mammograms. While the inexperienced clinicians were more likely to be influenced by AI predictions, all three groups fell victim to automation bias.
This tendency begs the question: How can educators help clinicians develop a more balanced evaluation of these digital tools? Part of the solution is providing them with direct, hands-on training. No one would consider teaching a medical student how to do a lumbar tap by providing them with a manual that shows them which intervertebral space to insert the needle and explains how much pressure to exert.
First, they need to learn the anatomy of the spine and the rationale behind performing the procedure. They may also use simulators to experience the pop as the needle finds the correct space. In years three and four of their training, they are guided step by step through the actual procedure with the assistance of experienced physicians. While misusing a chatbot and other AI tools may not do as much damage as a misplaced needle, it can nonetheless do very real harm.
Raja-Eli Abdulnour, MD, with Harvard Medical School, and associates recommend several steps to improve clinicians’ use of AI, including the aforementioned training, “Training should include simulations in which trainees receive AI generated information that, unbeknownst to them, is deliberately wrong, and the capture of [clinicians’] errors is then measured.”
With that advice in mind, Mayo Clinic's Harper Family Foundation has an AI Education in Medicine Program. It provides education for staff, students, and trainees in the ethical and responsible use of AI technology to care for patients and research the most complex medical problems. As the program’s website explains: “To be fully prepared for AI, we must achieve AI literacy and fluency by understanding what AI is and how it fits into our work. It is essential to achieve personal competency, gaining the confidence and skills to engage with AI safely and effectively. Additionally, building organizational competency is crucial to ensure our teams and institutions can use AI meaningfully and ethically.”
The Mayo Clinic program is consistent with the AMA policy on advanced AI literacy in medical education. The organization is developing curricular toolkits with that goal in mind and has developed a seven-part Artificial Intelligence in Health Care Series. The course includes a section on using AI in diagnosis and how to navigate the ethical and legal aspects of using AI in healthcare.
Clinicians who struggle with responsible AI usage would also benefit from training in how to submit a prompt to a chatbot. Since the introduction of ChatGPT, a new specialty has surfaced: Prompt engineering. Google defines it as “the art and science of designing and optimizing prompts to guide AI models, particularly LLMs, towards generating the desired responses. By carefully crafting prompts, you provide the model with context, instructions, and examples that help it understand your intent and respond in a meaningful way.” Mastering this skill is essential for clinicians who want to get the most out of their queries.
Essentially, clinicians need to think about how they think—metacognition—and how they reason before submitting a query. Without this learned cognitive skill, they are far more likely to risk diagnostic bias, generic advice that ignores an individual patient’s special circumstances, or miss the clinical context. A prompt that leaves out comorbidities, the medications a patient is taking, or some unique lifestyle factors will likely generate less then optimal advice from a chatbot. Abdulnour et al also recommend that AI training include think aloud approaches and structured read back techniques to address the problem.
Healthcare professionals have entered a new ecosystem, one that has been profoundly influenced by large language models, transformer architecture, and attention mechanisms. These terms didn’t even exist when most clinicians received their initial training. Education is the key to succeeding in this new world.
