arXiv:2510.10415cs.CLcs.AI2025-10被引 1

针对临床问答评估成本高难题,提出高效可靠的新框架

CQA-Eval: Designing Reliable Evaluations of Multi-paragraph Clinical QA under Resource Constraints

  • 按句子细粒度评估正确性,按答案粗粒度评估相关性
  • 仅标注少量句子即可达到与全量标注相当的可靠性
  • 适合医疗AI评估、资源有限的研究团队使用

评估多段落临床问答系统成本高且困难:准确判断需医学专业知识,对多段文本达成一致的人类评判也较难。我们提出CQA-Eval评估框架与建议,适用于资源有限且需高专业性的场景。基于300个真实患者问题的医生标注(涵盖医生与大模型的回答),我们比较了粗粒度答案级与细粒度句子级评估在正确性、相关性和风险披露三个维度的表现。结果显示:细粒度标注提升正确性的一致性,粗粒度标注提升相关性的一致性,而风险披露判断仍不一致。此外,仅标注少量句子即可获得与粗粒度标注相当的可靠性,显著降低评估成本与工作量。

原文摘要 · Abstract (English)

Evaluating multi-paragraph clinical question answering (QA) systems is resource-intensive and challenging: accurate judgments require medical expertise and achieving consistent human judgments over multi-paragraph text is difficult. We introduce CQA-Eval, an evaluation framework and set of evaluation recommendations for limited-resource and high-expertise settings. Based on physician annotations of 300 real patient questions answered by physicians and LLMs, we compare coarse answer-level versus fine-grained sentence-level evaluation over the dimensions of correctness, relevance, and risk disclosure. We find that inter-annotator agreement (IAA) varies by dimension: fine-grained annotation improves agreement on correctness, coarse improves agreement on relevance, and judgments on communicates-risks remain inconsistent. Additionally, annotating only a small subset of sentences can provide reliability comparable to coarse annotations, reducing cost and effort.

临床问答评估框架医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。