让视觉问答机器人每一步都可追溯,确保推理有据可依。
LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger

- 用结构化证据账本记录每步操作来源,限制后续推理只能引用已有证据。
- 在多个基准上提升答案准确率与推理过程可信度,尤其减少虚假推理。
- 适合需要高可信度推理的场景,如医疗、金融等严谨领域应用。
多模态智能体在视觉问答中逐步演变为包含感知、检索与推理的多步轨迹,但评估仍仅关注最终答案正确性。这一整体指标无法区分正确答案是基于真实证据、语言先验还是偶然误差抵消。为此,我们提出将多模态智能体轨迹视为带溯源约束的状态机:工具输出被规范化为结构化证据账本(Structured Evidence Ledger),作为轨迹状态;下游推理与决策仅可引用活跃账本条目,实体与数值层面均进行可追溯性验证;修复通过类型化的状态转移实现,禁止引入无工具来源的内容。我们据此构建LedgerMind系统,包含三层接地协议、自适应双路径调度器(按问题复杂度匹配推理深度)及事件触发式验证-修复引擎(具备形式化溯源不扩增保证)。实验覆盖多个多模态推理基准与主流多模态大模型,验证其有效缓解了最终答案准确率掩盖的四大典型失败模式:中间推理无依据、基于引用的实体幻觉(幻影溯源)、简单问题过度推理、修复时信息冗余放大。结果表明,LedgerMind同时提升了答案准确率与轨迹级忠实度。
原文摘要 · Abstract (English)
Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answer accuracy. This aggregate signal cannot tell whether a correct answer was reached through grounded evidence, language priors, or accidental error cancellation. We propose to treat a multimodal agent trajectory as a provenance-constrained state machine: tool outputs are normalized into a Structured Evidence Ledger that serves as the trajectory state, downstream reasoning and decision claims may cite only active ledger entries, grounding is checked at the entity and numeric level, and repair is realized as typed state transitions that cannot introduce content without tool-produced provenance. We instantiate this design as LedgerMind (Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger), augmented by a Three-Layer Grounding Protocol, an Adaptive Dual-Path Dispatcher that matches reasoning depth to question complexity, and an Event-Triggered Verification-and-Repair engine with a formal provenance non-amplification guarantee. We use LedgerMind to target four recurring failure patterns that final-answer accuracy tends to obscure: unsupported intermediate reasoning, citation-backed entity hallucination (Phantom Grounding), over-reasoning on simple queries, and repair-time amplification. Experiments across multiple multimodal reasoning benchmarks and backbone MLLMs show that LedgerMind improves both answer accuracy and trajectory-level faithfulness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。