arXiv:2509.12440cs.CLcs.AI2025-09中稿 · The Fifth Workshop…

测试大模型在中文医学文本中的事实核查能力,发现其定位错误能力弱且易误判正确信息。

MedFact: Benchmarking the Fact-Checking Capabilities of Large Language Models on Chinese Medical Texts

  • 构建包含2116条专家标注的中文医学事实核查数据集,覆盖13个专科与5级难度。
  • 20个主流大模型在真伪判断上表现尚可,但错误定位准确率远低于人类。
  • 发现模型存在'过度批评'现象,高级推理技术反而加剧误判风险。

将大语言模型(LLMs)部署于医疗场景需具备事实核查能力以保障患者安全和合规性。本文提出MedFact,一个具有挑战性的中文医学事实核查基准,包含2,116条来自多样化真实文本的专家标注实例,涵盖13个专科、8类错误类型、4种写作风格及5个难度等级。数据构建采用混合式人机框架,通过迭代专家反馈优化AI驱动的多准则筛选流程,确保高质量与高难度。我们评估了20个领先的LLMs在真伪分类与错误定位上的表现,结果显示模型虽能判断文本是否存在错误,但在精确定位错误方面表现不佳,顶尖模型仍显著落后于人类水平。分析揭示了'过度批评'现象——模型倾向于将正确信息误判为错误,且这一问题在多代理协作与推理时长扩展等先进推理技术下更为严重。MedFact凸显了医疗LLM部署的挑战,并为开发事实可靠医疗AI系统提供了资源支持。

原文摘要 · Abstract (English)

Deploying Large Language Models (LLMs) in medical applications requires fact-checking capabilities to ensure patient safety and regulatory compliance. We introduce MedFact, a challenging Chinese medical fact-checking benchmark with 2,116 expert-annotated instances from diverse real-world texts, spanning 13 specialties, 8 error types, 4 writing styles, and 5 difficulty levels. Construction uses a hybrid AI-human framework where iterative expert feedback refines AI-driven, multi-criteria filtering to ensure high quality and difficulty. We evaluate 20 leading LLMs on veracity classification and error localization, and results show models often determine if text contains errors but struggle to localize them precisely, with top performers falling short of human performance. Our analysis reveals the "over-criticism" phenomenon, a tendency for models to misidentify correct information as erroneous, which can be exacerbated by advanced reasoning techniques such as multi-agent collaboration and inference-time scaling. MedFact highlights the challenges of deploying medical LLMs and provides resources to develop factually reliable medical AI systems.

医学AI事实核查大模型评测中文数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。