多模态少样本分割新框架,融合视觉文本音频提升语义理解
DFR: A Decompose-Fuse-Reconstruct Framework for Multi-Modal Few-Shot Segmentation
- 分-融-重构三步走:拆解多模态信息,动态融合增强语义一致性
- 在真实与合成数据上均超越现有方法,显著提升分割精度
- 适合需要多模态感知的少样本场景,如跨模态图像理解任务
本文提出DFR(Decompose, Fuse and Reconstruct)框架,解决少样本分割中有效利用多模态引导的根本挑战。现有方法多依赖视觉支持样本或文本描述,单一或双模态范式限制了真实场景中丰富感知信息的挖掘。为此,该方法基于分割一切模型(SAM),系统整合视觉、文本与音频模态以增强语义理解。核心创新包括:1)多模态分解:通过SAM提取视觉区域建议,将文本语义扩展为细粒度描述符,并处理音频特征实现上下文丰富;2)多模态对比融合:采用对比学习保持跨模态一致性,同时实现前景与背景特征间的动态语义交互;3)双路径重构:自适应融合三模态融合令牌的语义引导与多模态位置先验的几何线索。在合成与真实设置下,跨视觉、文本、音频模态的大量实验表明,DFR显著优于当前最优方法。
原文摘要 · Abstract (English)
This paper presents DFR (Decompose, Fuse and Reconstruct), a novel framework that addresses the fundamental challenge of effectively utilizing multi-modal guidance in few-shot segmentation (FSS). While existing approaches primarily rely on visual support samples or textual descriptions, their single or dual-modal paradigms limit exploitation of rich perceptual information available in real-world scenarios. To overcome this limitation, the proposed approach leverages the Segment Anything Model (SAM) to systematically integrate visual, textual, and audio modalities for enhanced semantic understanding. The DFR framework introduces three key innovations: 1) Multi-modal Decompose: a hierarchical decomposition scheme that extracts visual region proposals via SAM, expands textual semantics into fine-grained descriptors, and processes audio features for contextual enrichment; 2) Multi-modal Contrastive Fuse: a fusion strategy employing contrastive learning to maintain consistency across visual, textual, and audio modalities while enabling dynamic semantic interactions between foreground and background features; 3) Dual-path Reconstruct: an adaptive integration mechanism combining semantic guidance from tri-modal fused tokens with geometric cues from multi-modal location priors. Extensive experiments across visual, textual, and audio modalities under both synthetic and real settings demonstrate DFR's substantial performance improvements over state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。