提出结构化对齐方法,缓解文本视频检索中的遗忘问题。
StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval
- 用等角紧框架几何结构统一建模跨模态对齐
- 在多个数据集上超越现有持续检索方法
- 适合需要长期学习新类别又不丢旧知识的多模态系统
持续文本到视频检索(CTVR)是一项挑战性的多模态持续学习任务,模型需在增量学习新语义类别时保持对已学内容的准确对齐,易产生灾难性遗忘。核心挑战包括:模态内特征漂移和跨模态非合作性特征漂移导致的模态错位。为此,我们提出 StructAlign,一种面向 CTVR 的结构化跨模态对齐方法。首先,引入单纯形等角紧框架(ETF)几何作为统一几何先验,以缓解模态错位;在此基础上,设计跨模态 ETF 对齐损失,将文本与视频特征对齐至类别级 ETF 原型,促使表示近似形成单纯形 ETF 几何。此外,为抑制模态内特征漂移,设计跨模态关系保持损失,利用互补模态保留跨模态相似性关系,提供稳定的关系监督。联合解决跨模态与模态内漂移,有效缓解了 CTVR 中的灾难性遗忘。在基准数据集上的大量实验表明,该方法持续优于当前最优的持续检索方法。
原文摘要 · Abstract (English)
Continual Text-to-Video Retrieval (CTVR) is a challenging multimodal continual learning setting, where models must incrementally learn new semantic categories while maintaining accurate text-video alignment for previously learned ones, thus making it particularly prone to catastrophic forgetting. A key challenge in CTVR is feature drift, which manifests in two forms: intra-modal feature drift caused by continual learning within each modality, and non-cooperative feature drift across modalities that leads to modality misalignment. To mitigate these issues, we propose StructAlign, a structured cross-modal alignment method for CTVR. First, StructAlign introduces a simplex Equiangular Tight Frame (ETF) geometry as a unified geometric prior to mitigate modality misalignment. Building upon this geometric prior, we design a cross-modal ETF alignment loss that aligns text and video features with category-level ETF prototypes, encouraging the learned representations to form an approximate simplex ETF geometry. In addition, to suppress intra-modal feature drift, we design a Cross-modal Relation Preserving loss, which leverages complementary modalities to preserve cross-modal similarity relations, providing stable relational supervision for feature updates. By jointly addressing non-cooperative feature drift across modalities and intra-modal feature drift, StructAlign effectively alleviates catastrophic forgetting in CTVR. Extensive experiments on benchmark datasets demonstrate that our method consistently outperforms state-of-the-art continual retrieval approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。