arXiv:2608.03675cs.CL2026-08

为兽医长文本问答设计风险加权事实验证方法,提升生成内容可靠性。

VetScore: Risk-Weighted Fact Verification for Veterinary Long-Form QA with Citations

论文配图:VetScore: Risk-Weighted Fact Verification for Veterinary Long-Form QA with Citations
图 1 · 摘自论文原文
  • 分步评估生成主张:拆解输出、评估危害性与引用一致性。
  • 与兽医专家评分高度相关,小模型也能达到良好效果。
  • 适用于高风险医疗问答场景,支持多维度可解释性分析。

引用片段可用于提升生成结果的可靠性及其对来源的忠实度,这在人类和兽医医学等高风险领域尤为重要。然而,这并不能保证生成主张与所提供片段一致。我们提出VetScore,一种用于兽医长文本问答的多步骤评估方法,旨在衡量生成主张在给定引用片段中的支持程度,并根据每个主张的危害潜力进行加权。VetScore首先将输出分割并分解为独立主张,然后针对每个主张评估其危害潜力并检验其与源引用片段的一致性,最后计算整体风险调整得分。我们构建了一个专家标注的元评估数据集,使用多种判别模型评估该方法,结果显示即使使用小型判别模型,其与兽医专家评分仍具有高度相关性,同时提供多维度可解释性。

原文摘要 · Abstract (English)

Citation excerpts can be used to increase the reliability of generated outputs and their faithfulness to cited sources, which is especially important in high-stakes domains such as human and veterinary medicine. However, this does not guarantee that generated claims are faithful to the provided excerpts. We present VetScore, a multi-step evaluation method for veterinary long-form question answering, designed to assess how well are generated claims supported by the provided excerpts, weighing this information by each claim's harm potential. VetScore first segments the output and decomposes it into individual claims, then scores each claim with respect to its harm potential and evaluates its faithfulness to source excerpts, and finally calculates the overall risk-adjusted score. We collect an expert-annotated meta-evaluation dataset, evaluate our approach with a range of judge models, and show that it achieves high correlations with veterinary experts even with small judge models, while offering explainability across multiple dimensions.

事实验证兽医AI可解释性风险评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。