针对临床问答评估成本高难题,提出高效可靠的新框架
CQA-Eval: Designing Reliable Evaluations of Multi-paragraph Clinical QA under Resource Constraints
- 按句子细粒度评估正确性,按答案粗粒度评估相关性
- 仅标注少量句子即可达到与全量标注相当的可靠性
- 适合医疗AI评估、资源有限的研究团队使用
评估多段落临床问答系统成本高且困难:准确判断需医学专业知识,对多段文本达成一致的人类评判也较难。我们提出CQA-Eval评估框架与建议,适用于资源有限且需高专业性的场景。基于300个真实患者问题的医生标注(涵盖医生与大模型的回答),我们比较了粗粒度答案级与细粒度句子级评估在正确性、相关性和风险披露三个维度的表现。结果显示:细粒度标注提升正确性的一致性,粗粒度标注提升相关性的一致性,而风险披露判断仍不一致。此外,仅标注少量句子即可获得与粗粒度标注相当的可靠性,显著降低评估成本与工作量。
原文摘要 · Abstract (English)
Evaluating multi-paragraph clinical question answering (QA) systems is resource-intensive and challenging: accurate judgments require medical expertise and achieving consistent human judgments over multi-paragraph text is difficult. We introduce CQA-Eval, an evaluation framework and set of evaluation recommendations for limited-resource and high-expertise settings. Based on physician annotations of 300 real patient questions answered by physicians and LLMs, we compare coarse answer-level versus fine-grained sentence-level evaluation over the dimensions of correctness, relevance, and risk disclosure. We find that inter-annotator agreement (IAA) varies by dimension: fine-grained annotation improves agreement on correctness, coarse improves agreement on relevance, and judgments on communicates-risks remain inconsistent. Additionally, annotating only a small subset of sentences can provide reliability comparable to coarse annotations, reducing cost and effort.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。