arXiv:2509.02198cs.CL2025-09被引 1

构建医疗领域细粒度文本生成评估基准,提升LLM事实准确性判断能力。

FActBench: A Benchmark for Fine-grained Automatic Evaluation of LLM-Generated Text in the Medical Domain

  • 设计多任务医疗文本生成评估框架,覆盖四大生成任务。
  • 联合使用CoT提示与NLI技术,一致评分与专家评估相关性最高。
  • 适合医学AI研究者、临床数据生成系统开发者参考使用。

大语言模型在专业领域表现受限,其中事实准确性尤为关键。为应对这一挑战,我们构建了FActBench——一个覆盖四个生成任务、六种先进大模型的医疗领域细粒度事实核查基准。采用链式思维提示(CoT Prompting)与自然语言推理(NLI)两种前沿事实核查技术,实验表明:两者通过一致投票获得的评分与领域专家评估的相关性最优,验证了该基准的有效性。

原文摘要 · Abstract (English)

Large Language Models tend to struggle when dealing with specialized domains. While all aspects of evaluation hold importance, factuality is the most critical one. Similarly, reliable fact-checking tools and data sources are essential for hallucination mitigation. We address these issues by providing a comprehensive Fact-checking Benchmark FActBench covering four generation tasks and six state-of-the-art Large Language Models (LLMs) for the Medical domain. We use two state-of-the-art Fact-checking techniques: Chain-of-Thought (CoT) Prompting and Natural Language Inference (NLI). Our experiments show that the fact-checking scores acquired through the Unanimous Voting of both techniques correlate best with Domain Expert Evaluation.

医疗AI事实核查大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。