对比多种评估方法,发现大模型能更好判断医学报告的因果解释质量。
Evaluating Causal Explanation in Medical Reports with LLM-Based and Human-Aligned Metrics
- 用大模型和人工标准对比六种评估指标的可靠性
- GPT-Black在识别逻辑合理、临床有效的因果叙述上表现最佳
- 相似性指标与临床推理质量不匹配,权重策略影响结果
本研究探讨不同评估指标在自动生成诊断报告中因果解释质量评价上的准确性。我们在两种输入类型(基于观察和多选题)下比较六种指标:BERTScore、余弦相似度、BioSentVec、GPT-White、GPT-Black以及专家定性评估。采用两种加权策略:一种反映任务优先级,另一种赋予所有指标同等权重。结果显示,GPT-Black在识别逻辑连贯且临床有效的因果叙述方面具有最强判别力;GPT-White也与专家评估高度一致,而基于相似性的指标则偏离临床推理质量。这些发现强调了评估指标选择与加权策略对结果的影响,支持在需要可解释性和因果推理的任务中使用大模型评估。
原文摘要 · Abstract (English)
This study investigates how accurately different evaluation metrics capture the quality of causal explanations in automatically generated diagnostic reports. We compare six metrics: BERTScore, Cosine Similarity, BioSentVec, GPT-White, GPT-Black, and expert qualitative assessment across two input types: observation-based and multiple-choice-based report generation. Two weighting strategies are applied: one reflecting task-specific priorities, and the other assigning equal weights to all metrics. Our results show that GPT-Black demonstrates the strongest discriminative power in identifying logically coherent and clinically valid causal narratives. GPT-White also aligns well with expert evaluations, while similarity-based metrics diverge from clinical reasoning quality. These findings emphasize the impact of metric selection and weighting on evaluation outcomes, supporting the use of LLM-based evaluation for tasks requiring interpretability and causal reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。