arXiv:2502.15834cs.LGstat.ML2025-02

多模态深度预测中,核心数据选择面临新挑战。

Challenges of Multi-Modal Coreset Selection for Depth Prediction

  • 将单模态最优核心集选取方法拓展至多模态场景
  • 发现融合嵌入与降维策略效果受限于模态间关系
  • 适合研究多模态学习与高效训练的学者参考

核心集选取方法能有效加速训练并降低内存开销,但在多模态应用中仍缺乏探索。本文将一种最先进的核心集选取技术适配到多模态数据,聚焦于深度预测任务。通过实验对比嵌入聚合与维度缩减方法,揭示了将单模态算法推广至多模态场景所面临的挑战,凸显了需设计专门方法以更好捕捉模态间关联性的必要性。

原文摘要 · Abstract (English)

Coreset selection methods are effective in accelerating training and reducing memory requirements but remain largely unexplored in applied multimodal settings. We adapt a state-of-the-art (SoTA) coreset selection technique for multimodal data, focusing on the depth prediction task. Our experiments with embedding aggregation and dimensionality reduction approaches reveal the challenges of extending unimodal algorithms to multimodal scenarios, highlighting the need for specialized methods to better capture inter-modal relationships.

多模态核心集深度预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。