arXiv:2603.00324cs.CV2026-03被引 1

PoP框架让AI推理可验证,用工具+置信区间减少幻觉。

Proof-of-Perception: Certified Tool-Using Multimodal Reasoning with Compositional Conformal Guarantees

  • 将多模态推理构造成带可信度的可执行图
  • 在多个基准上比现有方法更准更省算力
  • 适合需要可解释性与可靠性保障的场景

我们提出Proof-of-Perception(PoP),一种使用工具的多模态推理框架,将推理过程建模为带显式可靠性保证的可执行图。每个感知或逻辑节点输出一个置信集,实现校准的分步不确定性;轻量级控制器基于这些证书在预算内分配计算资源,仅在必要时调用额外工具,否则提前终止。该方法使答案基于可验证证据,减少误差累积与幻觉,支持合理的准确率-算力权衡。在文档、图表和多图像问答基准上,PoP在性能与可靠性方面均优于强基线方法(如链式思维、ReAct、程序化思维),同时计算效率更高。

原文摘要 · Abstract (English)

We present Proof-of-Perception (PoP), a tool-using framework that casts multimodal reasoning as an executable graph with explicit reliability guarantees. Each perception or logic node outputs a conformal set, yielding calibrated, stepwise uncertainty; a lightweight controller uses these certificates to allocate compute under a budget, expanding with extra tool calls only when needed and stopping early otherwise. This grounds answers in verifiable evidence, reduces error compounding and hallucinations, and enables principled accuracy-compute trade-offs. Across document, chart, and multi-image QA benchmarks, PoP improves performance and reliability over strong chain-of-thought, ReAct-style, and program-of-thought baselines while using computation more efficiently.

多模态推理可验证性工具使用置信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。