解决多模态传感器学习不平衡问题,提升系统可靠性。
Balancing Multi-modal Sensor Learning via Multi-objective Optimization
- 将多模态学习建模为多目标优化,优先提升最差模态性能。
- 在标准数据集上实现更均衡的性能,优于现有最优方法。
- 计算效率高,子程序耗时减少约20倍,适合资源受限场景。
依赖多感官(如视觉、音频、语言等)感知与决策支持的学习型控制系统日益增多。其关键挑战在于多模态传感器训练动态常出现不平衡:易学模态主导优化过程,而难学模态被忽视,导致感知扰动下系统可靠性下降。现有平衡策略多为启发式,且需高计算成本的子程序。本文将多模态传感器学习重构为多目标优化(MOO)问题,显式优先优化最差模态,同时保留原有的多模态融合目标。为此提出简单梯度方法MIMO(基于MOO的多模态传感器学习),提供收敛性保证,并在标准多模态基准上进行评估。结果表明,相比当前最优的平衡多模态学习与MOO基线,MIMO实现更优的均衡性能,且子程序计算时间最多降低约20倍,凸显其在资源受限控制流水线中的适用性。
原文摘要 · Abstract (English)
Learning-enabled control systems increasingly rely on multiple sensing modalities (e.g., vision, audio, language, etc.) for perception and decision support. A key challenge is that multi-modal sensor training dynamics are often imbalanced: fast-to-learn sensing channels dominate optimization, while slower channels remain underutilized, degrading reliability under sensing perturbations. Existing balancing strategies are largely heuristic and can require computationally intensive subroutines. In this paper, we reformulate multi-modal sensor learning as a multi-objective optimization (MOO) problem that explicitly prioritizes the worst-performing modality while retaining the nominal multi-modal sensor fusion objective. We then propose a simple gradient-based method, MIMO (multi-modal sensor learning via MOO), for the resulting formulation. We provide convergence guarantees and evaluate the method on standard multi-modal benchmarks. Results show improved balanced performance over state-of-the-art balanced multi-modal learning and MOO baselines, together with up to ~20x reduction in subroutine computation time, highlighting the suitability of MIMO for resource-constrained control pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。