arXiv:2605.27904cs.AIcs.LG2026-05被引 1

测试智能体主动找对预测有用信息的能力,发现现有方法表现很差。

Dr-CiK: A Testbed for Foresight-Driven Agents

论文配图:Dr-CiK: A Testbed for Foresight-Driven Agents
图 1 · 摘自论文原文
  • 设计新基准Dr-CiK,评估智能体从杂乱文档中找预测证据的能力。
  • 多数现有模型只找到不到5%的关键证据,被超过80%的干扰项误导。
  • 适合关注智能体自主决策、信息筛选与未来预测的研究者。

真实场景中的时间序列预测不仅依赖历史数据,还需主动从噪声大、异构的信息源中发现外部上下文。然而,现有上下文辅助预测的评测基准通常假设上下文已提供,未检验智能体能否自行识别。为此,我们提出Dr-CiK,一个评估智能体能否从文档语料库中检索与预测相关的支持性上下文、过滤干扰项、提炼出可用证据并生成有依据预测的基准。通过上下文消融实验和对先进深度研究与预测方法的联合评估,我们发现高质量上下文可显著提升预测性能。但大多数现有深度研究代理仅能获取不足5%的真实支撑证据,常被超过80%的干扰引用误导,且使用检索到的上下文反而导致预测更差。结果表明亟需发展能主动搜寻关键信息以预判未来的前瞻性智能体。

原文摘要 · Abstract (English)

Time series forecasting in real-world settings often depends not only on historical observations, but also on external context that must be actively discovered from noisy, heterogeneous information sources. Yet existing context-aided forecasting benchmarks typically assume that the supporting context is already provided, leaving open whether agents can identify it on their own. Therefore, we introduce Dr-CiK, a benchmark for evaluating whether agents can retrieve forecasting-relevant supporting context from a document corpus, filter out distractors, distill the retrieved context into forecast-useful evidence, and generate forecasts supported by that evidence. Through context ablations and evaluations of state-of-the-art deep research and forecasting methods paired together, we show that high-quality context substantially improves forecasting performance in Dr-CiK. However, most existing DR agents recover only a small fraction of the ground-truth supporting evidence (usually <5%), are frequently misled by distractors (>80% distractor citations), and can cause forecasters to perform worse with retrieved context than without context. Our results motivate research on foresight-driven agents that search for the right context to predict the future.

时间序列智能体预测上下文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。