用文本提示统一生成可见光-红外-标签三元组,提升少样本分割性能。
UniTriGen: Unified Triplet Generation of Aligned Visible-Infrared-Label for Few-Shot RGB-T Semantic Segmentation

- 通过共享潜空间与扩散过程,联合建模可见光、红外与标签的一致性。
- 在有限真实数据下生成高质量对齐三元组,使多个分割模型性能提升。
- 适合少样本场景下的多模态语义分割研究者使用。
RGB-T语义分割需要严格对齐的可见光-红外-标签三元组;然而真实场景中此类数据常稀缺。现有生成增强方法通常采用级联生成范式,将联合三元组生成分解为局部条件过程,导致可见光、红外与标签在空间结构、语义内容和跨模态细节上难以保持一致。为此,我们提出UniTriGen,一种统一三元组生成框架,在文本提示引导下直接生成空间对齐、语义一致且模态互补的可见光-红外-标签三元组。UniTriGen首次引入统一三元组生成机制,将可见光、红外与标签联合编码至共享潜空间,并通过扩散过程建模以保证全局跨模态一致性。进一步集成轻量级模态特定残差适配器,以适应各模态成像特性和输出格式。为缓解有限配对三元组中场景与类别分布不均带来的生成偏差,还采用场景平衡且类别感知的少样本采样策略,提升生成三元组的场景与类别多样性。实验表明,UniTriGen能从有限真实配对数据中生成高质量对齐三元组,从而在多种RGB-T语义分割模型上实现稳定性能提升。
原文摘要 · Abstract (English)
RGB-T semantic segmentation requires strictly aligned VIS-IR-Label triplets; however, such aligned triplet data are often scarce in real-world scenarios. Existing generative augmentation methods usually adopt cascaded generation paradigms, decomposing joint triplet generation into local conditional processes. As a result, consistency among VIS, IR, and Label in spatial structure, semantic content, and cross-modal details cannot be reliably maintained. To address this issue, we propose UniTriGen, a unified triplet generation framework that directly generates spatially aligned, semantically consistent, and modality complementary VIS-IR-Label triplets under the guidance of text prompts. UniTriGen first introduces a unified triplet generation mechanism, where VIS, IR, and Label are jointly encoded into a shared latent space and modeled with a diffusion process to enforce global cross-modal consistency. Lightweight modality-specific residual adapters are further integrated into this mechanism to accommodate modality-specific imaging characteristics and output formats. To mitigate generation bias caused by imbalanced scene and class distributions in limited paired triplets, UniTriGen also employs a scene-balanced and class-aware few-shot sampling strategy, which induces a more balanced sampling distribution and enhances the scene and class diversity of generated triplets. Experiments show that UniTriGen generates high-quality aligned triplets from limited real paired data, thereby achieving consistent performance improvements across various RGB-T semantic segmentation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。