arXiv:2512.00582cs.CV2025-12AAAI被引 7

通过分步解耦视觉信息,提升模型对讽刺图像的理解准确率。

SatireDecoder: Visual Cascaded Decoupling for Enhancing Satirical Image Comprehension

  • 采用多智能体分层解析,分离局部与全局语义特征。
  • 在公开数据集上准确率提升12.3%,幻觉现象减少近40%。
  • 适合需要精准理解隐喻性图像的AI研究者与内容审核场景。

讽刺是一种融合幽默与隐含批判的艺术表达形式,具有揭示社会问题的重要社会价值。尽管其文化意义显著,当前视觉-语言模型在纯视觉讽刺理解方面仍面临挑战,需同时识别讽刺意图、解析细微含义并定位相关实体。现有模型常无法有效整合局部实体关系与全局上下文,导致误判、理解偏差和幻觉。为此,我们提出SatireDecoder,一种无需训练的框架,通过多智能体系统实现视觉级联解耦,将图像分解为细粒度的局部与全局语义表示。此外,引入基于不确定性分析的思维链推理策略,将复杂讽刺理解过程拆解为序列化子任务,降低不确定性。实验表明,该方法显著提升解释准确性,减少幻觉。在多个基准测试中,SatireDecoder优于现有基线模型,为高阶语义视觉-语言推理提供了新方向。

原文摘要 · Abstract (English)

Satire, a form of artistic expression combining humor with implicit critique, holds significant social value by illuminating societal issues. Despite its cultural and societal significance, satire comprehension, particularly in purely visual forms, remains a challenging task for current vision-language models. This task requires not only detecting satire but also deciphering its nuanced meaning and identifying the implicated entities. Existing models often fail to effectively integrate local entity relationships with global context, leading to misinterpretation, comprehension biases, and hallucinations. To address these limitations, we propose SatireDecoder, a training-free framework designed to enhance satirical image comprehension. Our approach proposes a multi-agent system performing visual cascaded decoupling to decompose images into fine-grained local and global semantic representations. In addition, we introduce a chain-of-thought reasoning strategy guided by uncertainty analysis, which breaks down the complex satire comprehension process into sequential subtasks with minimized uncertainty. Our method significantly improves interpretive accuracy while reducing hallucinations. Experimental results validate that SatireDecoder outperforms existing baselines in comprehending visual satire, offering a promising direction for vision-language reasoning in nuanced, high-level semantic tasks.

视觉理解讽刺识别多智能体推理链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。