arXiv:2602.00996cs.CLcs.AI2026-02

多智能体协作问答框架,用自然语言日志实现去中心化推理。

DeALOG: Decentralized Multi-Agents Log-Mediated Reasoning Framework

  • 各智能体分工处理文本、表格、图像,通过共享日志协同工作。
  • 在6个数据集上表现接近顶尖模型,验证了日志与验证机制的有效性。
  • 适合需要可解释性与模块扩展性的复杂多模态问答任务。

跨文本、表格和图像的复杂问题回答需要整合多种信息源。亟需一种支持专业化处理、协调机制和可解释性的框架。我们提出DeALOG,一个用于多模态问题回答的去中心化多智能体框架。该框架包含专门处理表格、上下文、视觉内容、摘要和验证的智能体,通过共享的自然语言日志作为持久化记忆进行通信。这种基于日志的机制实现了无需中心控制的协作式错误检测与验证,提升了系统鲁棒性。在FinQA、TAT-QA、CRT-QA、WikiTableQuestions、FeTaQA和MultiModalQA六个数据集上的评估表明其性能具有竞争力。分析证实共享日志、智能体专业化和验证机制对准确率至关重要。DeALOG通过模块化组件与自然语言通信,提供了可扩展的解决方案。

原文摘要 · Abstract (English)

Complex question answering across text, tables and images requires integrating diverse information sources. A framework supporting specialized processing with coordination and interpretability is needed. We introduce DeALOG, a decentralized multi-agent framework for multimodal question answering. It uses specialized agents: Table, Context, Visual, Summarizing and Verification, that communicate through a shared natural-language log as persistent memory. This log-based approach enables collaborative error detection and verification without central control, improving robustness. Evaluations on FinQA, TAT-QA, CRT-QA, WikiTableQuestions, FeTaQA, and MultiModalQA show competitive performance. Analysis confirms the importance of the shared log, agent specialization, and verification for accuracy. DeALOG, provides a scalable approach through modular components using natural-language communication.

多智能体问答系统可解释性去中心化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。