自动化评估能准确区分AI对住院问题的回答好坏
Automated Evaluation can Distinguish the Good and Bad AI Responses to Patient Questions about Hospitalization
- 设计精细的自动评估框架,基于临床参考答案打分
- 28个AI系统在100个病例中表现,自动评分与医生评分高度一致
- 适合需要快速比较AI医疗回答质量的研究者和开发者
自动化方法用于回答患者提出的健康问题日益增多,但系统间的优劣评估仍依赖耗时的人工专家评审,难以规模化。现有自动指标常与人类判断不一致且受上下文影响。为验证自动化评估在住院相关患者提问场景下的可行性,我们开展了一项大规模系统性研究。在100个患者案例中,收集了28个AI系统生成的2800条回复,并从三方面评估:是否回答问题、是否恰当使用病历证据、是否运用通用医学知识。以临床医生编写的参考答案为基准,自动化评分与人工评分高度吻合。结果表明,精心设计的自动化评估可实现对AI系统的高效对比,助力医患沟通。
原文摘要 · Abstract (English)
Automated approaches to answer patient-posed health questions are rising, but selecting among systems requires reliable evaluation. The current gold standard for evaluating the free-text artificial intelligence (AI) responses--human expert review--is labor-intensive and slow, limiting scalability. Automated metrics are promising yet variably aligned with human judgments and often context-dependent. To address the feasibility of automating the evaluation of AI responses to hospitalization-related questions posed by patients, we conducted a large systematic study of evaluation approaches. Across 100 patient cases, we collected responses from 28 AI systems (2800 total) and assessed them along three dimensions: whether a system response (1) answers the question, (2) appropriately uses clinical note evidence, and (3) uses general medical knowledge. Using clinician-authored reference answers to anchor metrics, automated rankings closely matched human ratings. Our findings suggest that carefully designed automated evaluation can scale comparative assessment of AI systems and support patient-clinician communication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。