EventsThe 1st International Online Conference on Healthcare
Published
This submission belongs to the session S5. Generative AI in Clinical Practice—Evidence-Based Evaluation of Diagnostic and Therapeutic Applications of the event The 1st International Online Conference on Healthcare
Published date
20 Mar, 2026
Academic Editor
author-avatarPing Yu
Citation
Blessing Oluwatofunmi Apata, Oluwabusayomi Akeju, Olukayode Emmanuel Apata, Kelly L Wilson, Evaluating Large Language Models for Accuracy and Misinformation in HPV Vaccine Communication, in Proceedings of The 1st International Online Conference on Healthcare, 25 March–26 March 2026, MDPI: Basel, Switzerland
Share
Email
Facebook
Twitter
LinkedIn

Evaluating Large Language Models for Accuracy and Misinformation in HPV Vaccine Communication

Kelly L Wilson 4
1. Department of Health Behavior, School of Public Health, Texas A&M University, College Station. TX 77843, USA, USA
2. College of Integrated Health Sciences, University at Albany, State University of New York (SUNY), Albany, NY 12222, USA, USA
3. Department of Educational Psychology, College of Education and Human Development, Texas A&M University, College Station, TX 77843, USA, USA
4. College of Nursing, Texas A&M University, Bryan, TX 77807, USA, USA
Abstract

Introduction: Social media misinformation is a key contributor to low HPV vaccination rates, particularly among young adults who rely heavily on online sources. As generative artificial intelligence (GenAI) tools powered by large language models (LLMs) become widely used for health information, there are concerns that they could amplify misinformation by generating confident but incorrect responses. Understanding how these tools address HPV vaccination questions is therefore critical for public health communication.

Methodology: We systematically examined responses to HPV vaccine-related questions from two widely used LLMs, ChatGPT (Version 5) and Gemini (version 2.5). One team member posed 30 questions on vaccine safety, effectiveness, dosing schedule, and cost, among others, to each model, using a prompt requesting concise answers. All queries and outputs were saved verbatim. Two independent raters coded responses for accuracy (0 = incorrect to 3 = fully correct) and misinformation risk (0 = none to 2 = strongly misleading). Inter-rater reliability was high.

Results: Across the 30 prompts, both LLMs generated highly accurate content. All Gemini responses were rated fully accurate, with no misinformation detected and the lowest harm scores. ChatGPT responses were fully accurate for 29 of 30 items, with one response rated mostly correct but still free of clearly misleading statements. Most responses were consistent with authoritative guidance like the CDC and few referenced peer-reviewed studies. No response from either tool was judged to pose a strong risk of harm.

Conclusion: Both LLMs provided accurate, low-risk information about the HPV vaccine in this evaluation. Although performance may differ for other topics, languages, or future model versions, these findings suggest that current LLMs can serve as supportive tools for HPV vaccine education rather than major sources of misinformation. Ongoing monitoring and periodic re-evaluation are needed as these systems evolve and as users increasingly turn to AI for health information.

Keywords
Human papillomavirus
HPV vaccination
Misinformation
Generative Artificial Intelligence (GenAI)
ChatGPT
Gemini
Health communication
Poster
IOCH 2026.pdf
ADVANCEMENTS IN PREDICTING DIABETES BIOMARKERS: A MACHINE LEARNING EPIGENETIC APPROACH
Evaluating Generative AI in Clinical Workflows: A Practical Framework for Safe and Effective Deployment