Skip to main navigation Skip to search Skip to main content

Readability and information quality of LLM-Generated HIV content: a methodological content evaluation

  • Guihua Chen
  • , Chuan Lin
  • , Jae Heon Kim
  • , Fuxi Du
  • , Zhao Luo*
  • , Yu Seob Shin*
  • , Xianxin Li*
  • *Corresponding author for this work
  • Jeonbuk National University
  • Soonchunhyang University

Research output: Contribution to journalJournal articlepeer-review

Abstract

Background: Large language models (LLMs) are increasingly used to draft public-facing HIV materials, but the balance between readability and information quality remains unclear. This study provides a methodological, content-level evaluation of unedited LLM outputs. Methods: Fifteen high-frequency, HIV-relevant lay queries were drawn from Google Trends (using “AIDS” only as the entry term for salience) and mapped to MeSH. From March 13–20, 2025, each query was posed to ChatGPT-4, DeepSeek-R1, and Grok-3 across three conversational rounds (15 × 3 × 3 = 135 texts). Readability of raw outputs was computed using ARI, FRE, GFI, FKGL, and TSI. Information quality was rated with DISCERN by two non-specialist clinicians for all texts and by two infectious-disease specialists for a randomized subset (45 texts). Inter-rater reliability used two-way random-effects ICC; group differences used one-way ANOVA with Tukey’s post hoc or Kruskal–Wallis with Dunn’s post hoc, as appropriate. Results: ARI and GFI showed no between-model differences, whereas FRE, FKGL, and TSI indicated that Grok-3 required higher reading levels than ChatGPT-4 and DeepSeek-R1. In non-specialist ratings, Grok-3 scored higher on DISCERN (56.5 ± 7.7) than ChatGPT-4 (36.3 ± 5.2) and DeepSeek-R1 (36.3 ± 4.9) (both p < 0.0001). Among specialists, Grok-3 (71 [69–73]) exceeded ChatGPT-4 (43.5 [38–49], p < 0.0001) and DeepSeek-R1 (54 [49–59], p = 0.0006); ChatGPT-4 vs. DeepSeek-R1 was not significant (p = 0.2055). ICC among non-specialists was excellent for DeepSeek-R1 (0.95) and Grok-3 (0.98) but moderate for ChatGPT-4 (0.58). Among specialists, ICCs were moderate (≈ 0.59–0.72) with wide confidence intervals. Conclusions: A consistent readability–quality tension emerged: outputs rated higher for information quality tended to be harder to read, whereas more readable texts were less complete. These findings suggest that LLMs may be best suited for upstream use as drafting aids, paired with expert review for accuracy, stigma sensitivity, and audience fit, then edited to target literacy bands and localized, rather than as direct-to-patient systems. This methodological study does not assess reach, behavior change, or clinical outcomes; future work should test comprehension and actionability in target populations and explore low-bandwidth, localized deployments to determine whether content improvements translate into meaningful public-health benefits.

Original languageEnglish
Article number1624
JournalBMC Infectious Diseases
Volume25
Issue number1
DOIs
StatePublished - 2025.12

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 3 - Good Health and Well-being
    SDG 3 Good Health and Well-being

Keywords

  • DISCERN
  • HIV
  • Information quality
  • Inter-rater reliability (ICC)
  • Large language models (LLMs)
  • Readability

Fingerprint

Dive into the research topics of 'Readability and information quality of LLM-Generated HIV content: a methodological content evaluation'. Together they form a unique fingerprint.

Cite this