让AI推理过程更可信,通过能量模型约束视觉解释与实际计算一致。
DiffuseAgent-MI: Distributionally-Grounded,Tool-Integrated Self-Evolving Agents for Faithful Visual Reasoning

- 用能量模型约束视觉特征分布,确保解释与内部计算对齐。
- 在多个数据集上提升准确率最高5.1点,互信息忠实度翻倍。
- 适合需要高可信推理的医疗、金融等安全敏感场景。
工具集成的视觉语言智能体在组合性与多步视觉推理上取得显著进展,但其输出常出现不忠实问题:陈述的推理路径与实际生成答案的计算过程不符,影响安全关键应用的可靠性。本文提出DiffuseAgent-MI,一种自我演化的智能体,其感知基础由特征单元上的KL最小化能量模型驱动,提供视觉机制可解释性的分布视角。该模型学习一个能量景观,软性约束生成样本靠近选定可解释单元的原始先验,缩小解释与内部表征之间的差距。验证器提供轨迹级忠实度奖励,当检测到不忠实步骤时,修复分支重新调节能量。在GeoQA、SciVis、VQA-v2及自建多模态推理数据集上,DiffuseAgent-MI相较先前自演化智能体最高提升5.1点准确率,互信息忠实度与人类可解释性一致性均超两倍。分析表明,能量项与验证器互补:前者保障分布忠实度,后者保障轨迹级忠实度,二者结合才能同时弥合双重差距。
原文摘要 · Abstract (English)
Tool-integrated vision-language agents have made remarkable progress on compositional and multi-step visual reasoning. Yet their outputs frequently exhibit unfaithfulness: the stated reasoning path diverges from the computation that actually produced the answer, undermining reliability in safety-critical applications. We present DiffuseAgent-MI, a self-evolving agent whose perceptual grounding is governed by a KL-minimal energy model over feature units, providing a distributional view of visual mechanistic interpretability. The agent learns an energy landscape that softly constrains generated samples to lie near the native prior conditioned on the chosen interpretable unit, closing the gap between the explanation and the internal representation. A verifier then supplies trajectory-level faithfulness rewards, and a repair branch re-conditions the energy when the verifier flags an unfaithful step. On GeoQA, SciVis, VQA-v2 and an in-house multimodal reasoning set, DiffuseAgent-MI improves accuracy by up to 5.1 points over prior self-evolving agents while more than doubling mutual-information faithfulness and human-interpretability agreement. Our analysis shows the energy term and the verifier are complementary: the former guarantees distributional faithfulness, the latter trajectory-level faithfulness, and only their combination closes both gaps.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。