用多智能体辩论机制评估大模型长文本事实准确性
MAD-Fact: A Multi-Agent Debate Framework for Long-Form Factuality Evaluation in LLMs
- 设计多智能体辩论系统,模拟专家间交叉验证
- 在中文长文本上构建新数据集,提升评估针对性
- 适合关注大模型可信输出的科研与医疗从业者
大型语言模型在生物医学、法律、教育等高风险领域的广泛应用引发了对其输出事实准确性的严重担忧。现有短文本评估方法难以应对长文本中复杂的推理链、交织的观点和累积信息的问题。为此,我们提出一种整合大规模长文本数据集、多智能体验证机制和加权评估指标的系统性方法。构建了中文长文本事实性数据集LongHalluQA;开发了基于辩论的多智能体验证系统MAD-Fact。引入事实重要性层级,捕捉长文本中不同主张的重要性差异。在两个基准测试上的实验表明,更大规模的模型通常保持更高事实一致性,而国产模型在中文内容上表现更优。本工作为评估和提升长文本生成的事实可靠性提供了结构化框架,有助于推动大模型在敏感领域中的安全部署。
原文摘要 · Abstract (English)
The widespread adoption of Large Language Models (LLMs) raises critical concerns about the factual accuracy of their outputs, especially in high-risk domains such as biomedicine, law, and education. Existing evaluation methods for short texts often fail on long-form content due to complex reasoning chains, intertwined perspectives, and cumulative information. To address this, we propose a systematic approach integrating large-scale long-form datasets, multi-agent verification mechanisms, and weighted evaluation metrics. We construct LongHalluQA, a Chinese long-form factuality dataset; and develop MAD-Fact, a debate-based multi-agent verification system. We introduce a fact importance hierarchy to capture the varying significance of claims in long-form texts. Experiments on two benchmarks show that larger LLMs generally maintain higher factual consistency, while domestic models excel on Chinese content. Our work provides a structured framework for evaluating and enhancing factual reliability in long-form LLM outputs, guiding their safe deployment in sensitive domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。