提出TIGER框架,用图结构修复多模态生成中的事实错误
Self-Correction Can Amplify Hallucinations: Fact-Level Repair with Graph-Based Evidence Routing in Multimodal Generation

- 构建输入观察图与输出声明图,量化每条事实的风险
- 修复高风险事实后,幻觉内容减少37%且任务质量不变
- 适用于图像、音频、视频到文本等多种跨模态场景
我们研究多模态生成中的事实级修复问题,即流畅输出中可能包含未被输入支持的具体事实。现有推理时修复方法通常联合条件于输入和当前输出生成反馈,存在两个局限:输出中的幻觉陈述会误导模型对输入的理解,且自由形式的反馈无法在事实层面进行排序或调度。我们提出TIGER,一种重新设计反馈机制的推理时框架。TIGER独立从输入中提取观察图,从输出中提取声明图,并基于支持与冲突关系为每个声明分配图条件化的风险评分。模型仅修复高风险声明,同时保持主干网络冻结。我们提供收敛性分析,表明在弱假设下,期望总风险以几何速度下降至明确的渐近界。在四种跨模态路径(图像到文本、图像+文本到文本、音频到文本、视频到文本)上的实验表明,TIGER在保留任务质量的同时减少了未支持内容。该效果在多个主干模型上一致,且危机事实(CrisisFACTS)案例研究显示,相同修复机制可提升多源场景下的事实一致性。
原文摘要 · Abstract (English)
We study fact-level repair for multimodal generation, where a fluent output may contain specific facts that are not supported by the input. Existing inference-time repair methods often generate feedback by jointly conditioning on the input and the current output. This design has two limitations: hallucinated claims in the output can bias the model's interpretation of the input, and free-form feedback cannot be ranked or scheduled at the fact level. We present TIGER, an inference-time framework that redesigns feedback for localized repair. TIGER independently extracts an observation graph from the input and a claim graph from the current output, then assigns each claim a graph-conditioned risk score based on support and conflict. The model repairs selected high-risk claims while keeping the backbone frozen. We provide a convergence analysis showing that the expected total risk decreases geometrically to an explicit asymptotic bound under mild assumptions. Experiments across four cross-modal paths, including image-to-text, image+text-to-text, audio-to-text, and video-to-text, show that TIGER reduces unsupported content while preserving task quality. The gains hold across multiple backbones, and a CrisisFACTS case study suggests that the same repair mechanism can improve grounding in multi-source settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。