arXiv:2504.05607cs.CLcs.AI2025-04被引 4

用多智能体自动生成真假问题数据,提升长文本问答模型准确率。

FactGuard: Leveraging Multi-Agent Systems to Generate Answerable and Unanswerable Questions for Enhanced Long-Context LLM Extraction

  • 构建多智能体协作框架,自动生成可回答与不可回答问题。
  • 创建包含25,220条样本的FactGuard-Bench数据集,上下文长达8K至128K。
  • 揭示主流LLM在长文本中仅61.79%准确率,强调识别无答案题的重要性。

抽取式阅读理解系统旨在从给定文本中定位正确答案。然而,如何在保持高准确率的同时可靠识别无答案问题,仍是长期挑战。尽管大语言模型在阅读理解方面取得显著进展,但随着上下文长度持续扩展,该问题愈发突出。为此,我们提出一种基于多智能体协作框架的创新数据增强方法。不同于需昂贵人工标注的SQuAD 2.0等数据集,本方法可自主生成基于证据的问题-答案对,并系统构造不可回答问题。基于此,我们构建了FactGuard-Bench数据集,包含25,220个可回答与不可回答问题实例,上下文长度为8K至128K。在七个主流大语言模型上的实验表明,即使最先进的模型整体准确率也仅达61.79%。此外,强调模型对无答案问题的推理能力至关重要,以避免生成看似合理却错误的答案。通过在多智能体框架中高效进行数据选择与生成,本方法大幅降低传统人工标注成本,为LLM训练与优化提供重要参考。

原文摘要 · Abstract (English)

Extractive reading comprehension systems are designed to locate the correct answer to a question within a given text. However, a persistent challenge lies in ensuring these models maintain high accuracy in answering questions while reliably recognizing unanswerable queries. Despite significant advances in large language models (LLMs) for reading comprehension, this issue remains critical, particularly as the length of supported contexts continues to expand. To address this challenge, we propose an innovative data augmentation methodology grounded in a multi-agent collaborative framework. Unlike traditional methods, such as the costly human annotation process required for datasets like SQuAD 2.0, our method autonomously generates evidence-based question-answer pairs and systematically constructs unanswerable questions. Using this methodology, we developed the FactGuard-Bench dataset, which comprises 25,220 examples of both answerable and unanswerable question scenarios, with context lengths ranging from 8K to 128K. Experimental evaluations conducted on seven popular LLMs reveal that even the most advanced models achieve only 61.79% overall accuracy. Furthermore, we emphasize the importance of a model's ability to reason about unanswerable questions to avoid generating plausible but incorrect answers. By implementing efficient data selection and generation within the multi-agent collaborative framework, our method significantly reduces the traditionally high costs associated with manual annotation and provides valuable insights for the training and optimization of LLMs.

问答系统多智能体数据增强长文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。