arXiv:2501.12394cs.CYcs.LG2025-01被引 12

为医疗经济研究中使用大模型提供可操作的报告指南。

ELEVATE-GenAI: Reporting Guidelines for the Use of Large Language Models in Health Economics and Outcomes Research: an ISPOR Working Group on Generative AI Report

  • 构建十项涵盖模型特性与公平性的报告框架。
  • 在两项真实研究中验证了指南的适用性与实用性。
  • 适合科研人员、期刊编辑及审稿人参考使用。

生成式人工智能,特别是大语言模型(LLMs),在健康经济与结果研究(HEOR)中具有巨大潜力,但缺乏标准化的报告规范。本文提出专门针对HEOR中使用LLMs的研究的ELEVATE GenAI框架与检查清单。该框架基于文献综述和国际健康政策研究学会(ISPOR)生成式AI工作组的专家意见,包含模型特征、准确性、可复现性、公平性与偏见等十大领域。配套检查清单将框架转化为具体可执行的报告条目。通过在两项已发表的HEOR研究中应用——一项系统文献回顾任务,另一项经济建模——验证了框架在不同场景下的相关性与可用性。尽管框架提供了稳健的报告指导,仍需进一步实证测试以评估其有效性、完整性、易用性及跨应用场景的普适性。结论指出,ELEVATE GenAI框架填补了关键空白,有助于提升LLM辅助HEOR研究的透明度、准确性和可复现性。未来工作将聚焦于大规模测试与验证,推动更广泛采纳与优化。

原文摘要 · Abstract (English)

Introduction: Generative artificial intelligence (AI), particularly large language models (LLMs), holds significant promise for Health Economics and Outcomes Research (HEOR). However, standardized reporting guidance for LLM-assisted research is lacking. This article introduces the ELEVATE GenAI framework and checklist - reporting guidelines specifically designed for HEOR studies involving LLMs. Methods: The framework was developed through a targeted literature review of existing reporting guidelines, AI evaluation frameworks, and expert input from the ISPOR Working Group on Generative AI. It comprises ten domains, including model characteristics, accuracy, reproducibility, and fairness and bias. The accompanying checklist translates the framework into actionable reporting items. To illustrate its use, the framework was applied to two published HEOR studies: one focused on systematic literature review tasks and the other on economic modeling. Results: The ELEVATE GenAI framework offers a comprehensive structure for reporting LLM-assisted HEOR research, while the checklist facilitates practical implementation. Its application to the two case studies demonstrates its relevance and usability across different HEOR contexts. Limitations: Although the framework provides robust reporting guidance, further empirical testing is needed to assess its validity, completeness, usability, as well as its generalizability across diverse HEOR use cases. Conclusion: The ELEVATE GenAI framework and checklist address a critical gap by offering structured guidance for transparent, accurate, and reproducible reporting of LLM-assisted HEOR research. Future work will focus on extensive testing and validation to support broader adoption and refinement.

大模型健康经济报告指南

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。