Abstract
Background: Large language models (LLMs) are increasingly used to draft public-facing HIV materials, but the balance between readability and information quality remains unclear. This study provides a methodological, content-level evaluation of unedited LLM outputs. Methods: Fifteen high-frequency, HIV-relevant lay queries were drawn from Google Trends (using “AIDS” only as the entry term for salience) and mapped to MeSH. From March 13–20, 2025, each query was posed to ChatGPT-4, DeepSeek-R1, and Grok-3 across three conversational rounds (15 × 3 × 3 = 135 texts). Readability of raw outputs was computed using ARI, FRE, GFI, FKGL, and TSI. Information quality was rated with DISCERN by two non-specialist clinicians for all texts and by two infectious-disease specialists for a randomized subset (45 texts). Inter-rater reliability used two-way random-effects ICC; group differences used one-way ANOVA with Tukey’s post hoc or Kruskal–Wallis with Dunn’s post hoc, as appropriate. Results: ARI and GFI showed no between-model differences, whereas FRE, FKGL, and TSI indicated that Grok-3 required higher reading levels than ChatGPT-4 and DeepSeek-R1. In non-specialist ratings, Grok-3 scored higher on DISCERN (56.5 ± 7.7) than ChatGPT-4 (36.3 ± 5.2) and DeepSeek-R1 (36.3 ± 4.9) (both p < 0.0001). Among specialists, Grok-3 (71 [69–73]) exceeded ChatGPT-4 (43.5 [38–49], p < 0.0001) and DeepSeek-R1 (54 [49–59], p = 0.0006); ChatGPT-4 vs. DeepSeek-R1 was not significant (p = 0.2055). ICC among non-specialists was excellent for DeepSeek-R1 (0.95) and Grok-3 (0.98) but moderate for ChatGPT-4 (0.58). Among specialists, ICCs were moderate (≈ 0.59–0.72) with wide confidence intervals. Conclusions: A consistent readability–quality tension emerged: outputs rated higher for information quality tended to be harder to read, whereas more readable texts were less complete. These findings suggest that LLMs may be best suited for upstream use as drafting aids, paired with expert review for accuracy, stigma sensitivity, and audience fit, then edited to target literacy bands and localized, rather than as direct-to-patient systems. This methodological study does not assess reach, behavior change, or clinical outcomes; future work should test comprehension and actionability in target populations and explore low-bandwidth, localized deployments to determine whether content improvements translate into meaningful public-health benefits.
| Original language | English |
|---|---|
| Article number | 1624 |
| Journal | BMC Infectious Diseases |
| Volume | 25 |
| Issue number | 1 |
| DOIs | |
| State | Published - 2025.12 |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 3 Good Health and Well-being
Keywords
- DISCERN
- HIV
- Information quality
- Inter-rater reliability (ICC)
- Large language models (LLMs)
- Readability
Fingerprint
Dive into the research topics of 'Readability and information quality of LLM-Generated HIV content: a methodological content evaluation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver