arXiv:2510.24303cs.AI2025-10中稿 · AAMAS 2026被引 7

用多智能体协作+证据推理提升未来事件预测准确性

Retrieval- and Argumentation-Enhanced Multi-Agent LLMs for Judgmental Forecasting (Extended Version with Supplementary Material)

  • 设计多智能体框架,各智能体基于正反证据生成量化论证
  • 三智能体协作比双智能体提升预测准确率,且结果可解释
  • 适合需要可信推理的决策类预测任务

判断性预测是基于人类判断对未来事件进行预测的任务,可视为一种命题验证,其中命题对应未来事件,任务为评估该事件的合理性。本文提出一种新型多智能体框架用于命题验证,不同智能体可对命题真实性持不同意见,并提供具体证据支持或反驳,以量化双向论证框架(QBAFs)表示。我们在此框架下实现多种基于大语言模型(LLMs)的智能体:(1) ArgLLM,现有命题验证方法,能生成并评估QBAFs;(2) RbAM,利用外部源中基于关系的论证挖掘(RbAM)生成QBAFs;(3) RAG-ArgLLM,将ArgLLM扩展为从外部源检索增强生成论证。我们在两个标准判断性预测数据集上进行了实验,采用两种或三种智能体,由六种不同基础LLM驱动。结果表明,融合多个智能体的证据可提升预测准确率,尤其在三智能体情况下效果更显著,且提供可解释的证据整合方式。

原文摘要 · Abstract (English)

Judgmental forecasting is the task of making predictions about future events based on human judgment. This task can be seen as a form of claim verification, where the claim corresponds to a future event and the task is to assess the plausibility of that event. In this paper, we propose a novel multi-agent framework for claim verification, whereby different agents may disagree on claim veracity and bring specific evidence for and against the claims, represented as quantitative bipolar argumentation frameworks (QBAFs). We then instantiate the framework for supporting claim verification, with a variety of agents realised with Large Language Models (LLMs): (1) ArgLLM agents, an existing approach for claim verification that generates and evaluates QBAFs; (2) RbAM agents, whereby LLM-empowered Relation-based Argument Mining (RbAM) from external sources is used to generate QBAFs; (3) RAG-ArgLLM agents, extending ArgLLM agents with a form of Retrieval-Augmented Generation (RAG) of arguments from external sources. Finally, we conduct experiments with two standard judgmental forecasting datasets, with instances of our framework with two or three agents, empowered by six different base LLMs. We observe that combining evidence from agents can improve forecasting accuracy, especially in the case of three agents, while providing an explainable combination of evidence for claim verification.

多智能体推理增强预测建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。