Performance comparison of large language models in interpreting clinical guidelines for migraine prevention: A multidimensional analysis
DOI:
https://doi.org/10.54029/2026cssKeywords:
large language models, migraine, artificial intelligenceAbstract
Background & Objective: Large language models (LLMs) such as DeepSeek-R1, Gemini-2.5 Pro, ChatGPT-5 Thinking, and Grok-4 Expert are increasingly applied in medical contexts, yet their reliability in evidence-based clinical domains like migraine prophylaxis remains uncertain. This study aimed to compare the performance of four leading LLMs in interpreting and applying the International Headache Society’s global practice recommendations for preventive pharmacological treatment of migraine.
Methods: Sixteen standardized clinical scenario questions derived from the IHS guideline were presented identically to each model. Responses were evaluated by blinded expert raters across five dimensions—Accuracy, Overconclusiveness, Supplementary Value, Incompleteness, and Readability—using a 10-point Likert scale. Readability was further analyzed using composite indices from readabilityformulas.com.
Results: No significant inter-model differences were observed in Accuracy (P = 0.856), Overconclusiveness (P = 0.400), or Incompleteness (P = 0.531). However, DeepSeek-R1 provided significantly more Supplementary Information than Gemini-2.5 Pro (P = 0.010) and Grok-4 Expert (P = 0.030). Readability analysis further revealed substantial variation across models (P < 0.001), with DeepSeek-R1 generating the most accessible outputs.
Conclusion: While all four models exhibited comparable adherence to guideline-based content, DeepSeek-R1 demonstrated superior performance in supplementary informational value and readability. These findings highlight the importance of evaluating LLMs not only for accuracy but also for their capacity to enhance clinical communication and decision support.