arXiv:2601.19151cs.AIcs.MA2026-01被引 1

让多个专家模型辩论,提升大模型零样本时间序列推理能力

Multimodal Collaborative Debate for Zero-Shot Time Series Reasoning

  • 设计多模态专家协作辩论机制,分工处理文本、图像和数值信息
  • 在20个任务上超越基线,尤其擅长全局结构与跨模态推理
  • 无需微调,适合需要可靠推理的金融、医疗等时序数据分析场景

大型语言模型(LLMs)被广泛用作结构化数据的自然语言接口,但在时间序列推理中仍表现脆弱。视觉模式可能误导判断,数值陈述可能虚构,文本上下文可能压倒信号证据。本文将零样本时间序列推理视为多模态证据仲裁问题,提出TS-Debate——一种无需任务特定微调的推理时多智能体协议。该方法先提取相关领域知识,再分配专精于文本上下文、视觉模式和数值信号的智能体,并通过验证-冲突-校准流程协调其交互。评审智能体利用轻量级代码执行与数值查表验证关键主张,解决跨模态分歧并校准最终答案。相比通用多智能体辩论或无约束工具使用,TS-Debate明确定义了证据暴露方式、可验证主张及验证结果如何影响融合。在三个公开基准的20个任务中,TS-Debate显著提升分类与问答性能,揭示辩论最有益于全局结构与跨视图推理,而非局部值重建。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used as natural-language interfaces to structured data, yet they remain brittle when reasoning over time series. Visual patterns can be misleading, numerical claims can be hallucinated, and textual context can override evidence from the signal. We study zero-shot time-series reasoning as a multimodal evidence arbitration problem for LLM agents. We propose TS-Debate, an inference-time multi-agent protocol that requires no task-specific fine-tuning. TS-Debate first elicits relevant domain knowledge, then assigns modality-specialized agents to textual context, visual patterns, and numerical signals, and coordinates their interaction through a verification-conflict-calibration procedure. Reviewer agents check decision-critical claims with lightweight code execution and numerical lookup, resolve cross-modal disagreement, and calibrate the final answer. Unlike generic multi-agent debate or unconstrained tool use, TS-Debate specifies how evidence is exposed, which claims are checkable, and how verification outcomes shape synthesis. Across 20 tasks from three public benchmarks, TS-Debate improves classification and question answering performance over strong baselines, while revealing that debate is most useful for global-structure and cross-view reasoning rather than local value reconstruction.

时间序列多模态推理智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。