Performance comparison of large language models in interpreting clinical guidelines for migraine prevention: A multidimensional analysis

Authors

  • Li Xu The Forth Clinical Medical College , Zhejiang Chinese Medicine University, Hangzhou 315000, Zhejiang Province, Hangzhou, China
  • Xu Qiu 18226218711
  • Jiayi Deng Department of Pain, Wuxi Xishan People's Hospital, Wuxi 214105, China.
  • Chengqi Dong The Forth Clinical Medical College , Zhejiang Chinese Medicine University, Hangzhou 315000, Zhejiang Province, Hangzhou, China
  • Dong liang The Fourth School of Clinical Medicine, Zhejiang Chinese Medical University, Hangzhou First People’s Hospital, Hangzhou, China
  • Keyang Wu The Fourth School of Clinical Medicine, Zhejiang Chinese Medical University, Hangzhou First People’s Hospital, Hangzhou, China
  • Xiaoxue Dong National Neuroscience Institute of Singapore, 11 Jalan Tan Tock Seng, Singapore 308433, Singapore.
  • Tao Mei Department of Pain, The Affiliated Hangzhou First People's Hospital, Westlake University School of Medicine, Hangzhou, China.
  • Shi Chen Department of Pain, The Affiliated Hangzhou First People's Hospital, Westlake University School of Medicine, Hangzhou, China.
  • Yali Wu Department of Pain, The Affiliated Hangzhou First People's Hospital, Westlake University School of Medicine, Hangzhou, China
  • Yuan Cheng Department of Pain, The Affiliated Hangzhou First People's Hospital, Westlake University School of Medicine, Hangzhou, China
  • Jianliang Sun Department of Pain, the Affiliated Hangzhou First People's Hospital, Westlake University School of Medicine, Hangzhou, China.
  • Liang Yu Department of Pain, the Affiliated Hangzhou First People's Hospital, Westlake University School of Medicine, Hangzhou, China.
  • Hanbing Wang Department of Pain, the Affiliated Hangzhou First People's Hospital, Westlake University School of Medicine, Hangzhou, China.
  • Qinghua Li Department of Pain, the Affiliated Hangzhou First People's Hospital, Westlake University School of Medicine, Hangzhou, China.

DOI:

https://doi.org/10.54029/2026css

Keywords:

large language models, migraine, artificial intelligence

Abstract

Background & Objective: Large language models (LLMs) such as DeepSeek-R1, Gemini-2.5 Pro, ChatGPT-5 Thinking, and Grok-4 Expert are increasingly applied in medical contexts, yet their reliability in evidence-based clinical domains like migraine prophylaxis remains uncertain. This study aimed to compare the performance of four leading LLMs in interpreting and applying the International Headache Society’s global practice recommendations for preventive pharmacological treatment of migraine.

Methods: Sixteen standardized clinical scenario questions derived from the IHS guideline were presented identically to each model. Responses were evaluated by blinded expert raters across five dimensions—Accuracy, Overconclusiveness, Supplementary Value, Incompleteness, and Readability—using a 10-point Likert scale. Readability was further analyzed using composite indices from readabilityformulas.com.

Results: No significant inter-model differences were observed in Accuracy (P = 0.856), Overconclusiveness (P = 0.400), or Incompleteness (P = 0.531). However, DeepSeek-R1 provided significantly more Supplementary Information than Gemini-2.5 Pro (P = 0.010) and Grok-4 Expert (P = 0.030). Readability analysis further revealed substantial variation across models (P < 0.001), with DeepSeek-R1 generating the most accessible outputs.

Conclusion: While all four models exhibited comparable adherence to guideline-based content, DeepSeek-R1 demonstrated superior performance in supplementary informational value and readability. These findings highlight the importance of evaluating LLMs not only for accuracy but also for their capacity to enhance clinical communication and decision support.

Published

2026-09-18

Issue

Section

Original Article