arXiv:2607.03704cs.CL2026-07

用温度优化提升大模型对药物不良反应的因果判断准确率

Optimizing Large Language Models for Causality Assessment in Pharmacovigilance: Developing a Performance Metric as Objective for Bayesian Hyperparameter Optimization

  • 设计可适配高斯过程优化的评估指标,聚焦关键判断项
  • 温度调优使判断一致性从45%提升至72%,尤其改善模糊案例
  • 适合关注药监自动化评估的临床与算法研究者

背景:个例安全报告(ICSR)数量激增,亟需可扩展的自动化因果评估。大型语言模型(LLM)虽有潜力,但在临床任务中表现仍不理想,且推理时超参数优化尚未被研究。目标:开发适用于高斯过程(GP)优化的评估目标,并探究温度调节是否能提升GPT-5.2在FAERS ICSRs上对Naranjo因果评估的一致性。方法:对723个分层的FAERS病例进行专家评估,使用链式思维(CoT)提示测试OpenAI的GPT-5.2。构建四种复合指标:加权余弦相似度(WCS)、信息加权一致得分(IWAS)、熵加权一致与余弦相似度(EWACS)、共识加权余弦相似度(CWCS),并采用基于概率改进(PoI)获取策略的贝叶斯优化,在温度[0, 2]范围内进行搜索。结果:在基线温度(T=0)下,GPT-5.2优于先前生物医学大模型,对Naranjo量表第5题达成74.1%一致率,第10题为65.4%。熵分析显示仅第5、10题为有效优化目标。温度在群体层面无系统性影响(η = 0.002,p = 0.959)。以EWACS为指导的贝叶斯优化将因果分类一致性从45.0%提升至72.0%(+27个百分点),其中“可疑”类别的提升最大(+42.9个百分点)。结论:EWACS是最佳的GP兼容指标。尽管不存在普适最优温度,但基于案例的温度选择显著提升性能,支持在药物警戒中采用温度优化。

原文摘要 · Abstract (English)

Background: Growing individual case safety report (ICSR) volumes have intensified demand for scalable automated causality assessment. Large Language Models (LLMs) show promise, yet performance on clinically demanding tasks remains suboptimal and inference-time hyperparameter optimization has not been investigated. Objective: To develop a Gaussian Process (GP)-compatible optimization objective and investigate whether temperature optimization improves LLM-expert agreement on Naranjo causality assessment of FAERS ICSRs. Methods: Expert causality assessments were performed on 723 stratified FAERS cases. OpenAI's GPT-5.2 was evaluated using chain-of-thought (CoT) prompting. Four composite metrics were developed: Weighted Cosine Similarity (WCS), Information-Weighted Agreement Score (IWAS), Entropy-Weighted Agreement and Cosine Similarity Score (EWACS), and Consensus-Weighted Cosine Similarity (CWCS) and Bayesian optimization using a GP surrogate with Probability of Improvement (PoI) acquisition was applied across temperature [0, 2]. Results: GPT-5.2 outperformed prior biomedical LLMs at baseline (T = 0), achieving 74.1% agreement on question 5 and 65.4% on question 10 of Naranjo algorithm. Entropy analysis identified these as the sole informative optimization targets. Temperature showed no systematic population-level effect (\b{eta} = 0.002, p = 0.959). EWACS-guided Bayesian optimization improved causality classification agreement from 45.0% to 72.0% (+27 pp), with the largest gain in Doubtful cases (+42.9 pp). Conclusion: EWACS was identified as the optimal GP-compatible metric. The absence of a universal temperature optimum indicates LLM performance is driven primarily by ICSR content, yet case-specific temperature selection produced meaningful improvements, supporting temperature optimization for LLM-assisted pharmacovigilance.

大模型优化药物警戒因果推断贝叶斯优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。