arXiv:2508.06104cs.CV2025-08

通过多层级自适应修正与对齐,提升2D-3D跨模态检索在噪声标签下的性能。

MCA: 2D-3D Retrieval with Noisy Labels via Multi-level Adaptive Correction and Alignment

  • 利用多模态历史预测联合建模标签一致性,实现更可靠的标签修正。
  • 在多个基准上达到最优效果,噪声环境下仍保持稳定性能。
  • 适合需要处理标注不全或错误的跨模态检索研究者使用。

随着2D和3D数据的日益丰富,跨模态检索领域取得了显著进展。然而,不完善的标注带来了巨大挑战,亟需在噪声标签条件下实现鲁棒的2D-3D跨模态检索。现有方法通常在各模态内独立划分样本,易受污染标签过拟合。为此,本文提出多层级跨模态自适应修正与对齐框架(MCA)。首先引入多模态联合标签修正(MJC)机制,利用多模态历史自预测联合建模模态预测一致性,实现可靠标签优化;其次设计多层级自适应对齐(MAA)策略,有效增强跨模态特征语义与区分度。大量实验表明,MCA在常规及真实噪声3D基准上均达领先性能,验证了其通用性与有效性。

原文摘要 · Abstract (English)

With the increasing availability of 2D and 3D data, significant advancements have been made in the field of cross-modal retrieval. Nevertheless, the existence of imperfect annotations presents considerable challenges, demanding robust solutions for 2D-3D cross-modal retrieval in the presence of noisy label conditions. Existing methods generally address the issue of noise by dividing samples independently within each modality, making them susceptible to overfitting on corrupted labels. To address these issues, we propose a robust 2D-3D \textbf{M}ulti-level cross-modal adaptive \textbf{C}orrection and \textbf{A}lignment framework (MCA). Specifically, we introduce a Multimodal Joint label Correction (MJC) mechanism that leverages multimodal historical self-predictions to jointly model the modality prediction consistency, enabling reliable label refinement. Additionally, we propose a Multi-level Adaptive Alignment (MAA) strategy to effectively enhance cross-modal feature semantics and discrimination across different levels. Extensive experiments demonstrate the superiority of our method, MCA, which achieves state-of-the-art performance on both conventional and realistic noisy 3D benchmarks, highlighting its generality and effectiveness.

跨模态检索噪声标签2D-3D自适应对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。