arXiv:2510.03247cs.LGcs.AI2025-10中稿 · Transactions on Ma…被引 3

提出首个面向非对齐多模态数据的主动学习框架,大幅降低标注成本。

Towards Multimodal Active Learning: Efficient Learning with Limited Paired Data

  • 设计模态感知算法,同时考虑不确定性和多样性,动态选择最有价值样本。
  • 在ColorSwap数据集上减少40%标注量,性能保持不变。
  • 适用于池式与流式场景,计算高效,适合真实多模态应用。

主动学习(AL)是一种有效降低深度学习标注成本的策略。然而,现有算法几乎只针对单模态数据,忽略了多模态学习中高昂的对齐标注代价。本文提出首个面向非对齐多模态数据的主动学习框架,学习者需主动获取跨模态对齐而非预对齐样本的标签。该设定捕捉了现代多模态流程中的实际瓶颈:单模态特征易得,但高质量对齐成本高。我们开发了一种结合不确定性与多样性的模态感知算法,实现线性时间样本选择,且可无缝应用于池式与流式设置。在基准数据集上的大量实验表明,该方法显著降低多模态标注需求并维持性能;例如,在ColorSwap数据集上,标注量最多减少40%而精度无损失。

原文摘要 · Abstract (English)

Active learning (AL) is a principled strategy to reduce annotation cost in data-hungry deep learning. However, existing AL algorithms focus almost exclusively on unimodal data, overlooking the substantial annotation burden in multimodal learning. We introduce the first framework for multimodal active learning with unaligned data, where the learner must actively acquire cross-modal alignments rather than labels on pre-aligned pairs. This setting captures the practical bottleneck in modern multimodal pipelines, where unimodal features are easy to obtain but high-quality alignment is costly. We develop a new algorithm that combines uncertainty and diversity principles in a modality-aware design, achieves linear-time acquisition, and applies seamlessly to both pool-based and streaming-based settings. Extensive experiments on benchmark datasets demonstrate that our approach consistently reduces multimodal annotation cost while preserving performance; for instance, on the ColorSwap dataset it cuts annotation requirements by up to 40% without loss in accuracy.

主动学习多模态标注效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。