arXiv:2605.30802cs.MAcs.AI2026-05

多智能体系统提升预测市场裁决准确率,自动生成高可信答案并识别需人工介入的难题。

Design and Evaluation of Multi-Agent AI Oracle Systems for Prediction Market Resolution

论文配图:Design and Evaluation of Multi-Agent AI Oracle Systems for Prediction Market Resolution
图 1 · 摘自论文原文
  • 采用共享证据层的多智能体协同,通过置信度加权投票提升裁决精度
  • 独立聚合达83.43%准确率,优于最优单模型1.01个百分点
  • 提出混合路由策略,仅自动处理高置信一致问题,其余交由人工

预测市场依赖可靠的结果裁决来发挥价值,现有预言机系统在自动化速度与人工仲裁准确性之间权衡。单大模型预言机虽有可观准确率,但缺乏自我纠错能力,易继承底层模型全部缺陷。本文评估多智能体大模型架构在预测市场裁决中的表现,对比独立聚合与协商共识策略,基于1,189个来自KalshiBench的已解决议题,在统一证据层(Exa)支持下,通过发布日期过滤确保推理质量。独立聚合以置信度加权投票实现83.43%准确率,较最佳单模型(如GPT-5 Nano、DeepSeek V3、Llama-3.3-70B)高出1.01个百分点;协商共识则降至约76%,低于所有单模型基线,因错误在辩论中传播导致正确判断被反转。模型间误差相关性(0.529–0.689)解释了集成方法无法达到理论上的康多塞上限。部分问题难以被任何多智能体架构修正,因此建议升级至人机混合系统:仅对全体一致且高置信的问题自动裁决,可实现97.87%准确率,覆盖数据集47%,其余由智能体分歧标记后转交人工处理。

原文摘要 · Abstract (English)

Prediction markets aggregate collective intelligence to forecast uncertain events, but their utility depends on reliable outcome resolution. Existing oracle systems tradeoff fast but brittle automation against accurate but costly human arbitration. Single-LLM oracles achieve meaningful accuracy but inherit all failure modes of their underlying model with no self-correction mechanism. We evaluate whether multi-agent LLM architectures can improve oracle resolution accuracy over single-model baselines. We compare independent aggregation and deliberative consensus against single-LLM baselines (GPT-5 Nano, DeepSeek V3, and Llama-3.3-70B) on 1,189 resolved prediction market questions from KalshiBench. All agents share a common evidence layer through Exa, with retrieval filtered by publication date to isolate reasoning from retrieval quality. Independent aggregation with confidence-weighted voting achieves the highest accuracy at 83.43 percent, outperforming the best individual model by 1.01 percentage points. Deliberative consensus degrades accuracy to approximately 76 percent, below every single-model baseline, attributed to error propagation during debate where confidently wrong models flip correct ones. Error correlations across models (0.529-0.689) explain why aggregation gains fall short of the theoretical Condorcet ceiling, placing a fundamental limit on ensemble approaches. Many questions resist correction by any multi-agent architecture, motivating escalation to human arbitration. We propose routing criteria for hybrid AI-human oracle systems: auto-resolving only unanimous, high-confidence questions yields 97.87 percent accuracy on 47 percent of the dataset, with inter-agent disagreement flagging the remainder for human review.

预测市场多智能体自动裁决人机协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。