用能量机制动态干预,提升多模态生成的可靠性。
MODE-RAG: Manifold Outlier Diagnosis and Energy-based Retrieval-Augmented Generation Evaluation

- 基于变分自由能与注意力状态,动态决定是否干预生成
- 在MultiVent数据集上将幻觉率显著降低,逻辑错误减少40%
- 适合需要高可信度生成的医疗、法律等专业场景
尽管多模态检索增强生成(M-RAG)提升了大视觉语言模型的能力,但仍极易产生跨模态幻觉、因果捏造和奉承倾向。现有缓解方法常陷入干预悖论:静态规则会误伤准确生成,而完全不干预则使已有错配演变为严重逻辑谬误。为此,我们提出基于变分自由能(VFE)和内部注意力状态的多智能体系统MODE-RAG,动态控制干预时机。高风险查询被路由至五个阶段专用智能体,结合蒙特卡洛树搜索(MCTS)进行严格因果推导,并通过对数几率扰动抑制奉承行为。专门的修正与监督智能体确保格式稳定并执行事后事实验证。为客观评估,我们引入ModeVent——从MultiVent数据集中提取的挑战性子集。大量实验表明,该系统有效降低了幻觉率与逻辑虚构,显著提升了M-RAG系统的鲁棒性。
原文摘要 · Abstract (English)
While Multimodal Retrieval-Augmented Generation (M-RAG) enhances Large Vision-Language Models, it remains highly susceptible to cross-modal hallucinations, causal fabrications, and sycophancy. Furthermore, existing mitigation pipelines often face an intervention paradox: static rules tend to unnecessarily disrupt accurate generations, whereas leaving the multi-modal reasoning completely unguided allows existing mismatches to cascade into severe logical fabrications. To quantify and mitigate these hallucinations, we propose a Multi-Agent system, MODE-RAG, driven by Variational Free Energy (VFE) and internal attention states to dynamically gate interventions. High-risk queries are routed to five stage-specific agents, integrating Monte Carlo Tree Search (MCTS) for rigorous causal derivation and logit perturbations to penalize sycophancy. Dedicated Correction and Overseer agents ensure formatting stability and perform post-hoc factual verification. To objectively evaluate our approach, we introduce ModeVent, a challenging subset derived from the MultiVent dataset. Extensive experiments indicate that our system effectively reduces hallucination rates and logical fabrication, significantly improving the robustness of M-RAG systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。