用多任务学习和模型集成提升多视角视频动作定位精度
Multi-task Learning with Extended Temporal Shift Module for Temporal Action Localization
- 扩展TSM模块,引入背景类并分段分类动作区间
- 联合优化场景识别与动作定位,利用环境上下文增强效果
- 加权集成多模型预测,显著提升结果鲁棒性
我们针对ICCV 2025的BinEgo-360挑战赛提出解决方案,聚焦于多视角、多模态视频中的时序动作定位(TAL)。数据集包含全景、第三人称及第一人称视频片段,标注了细粒度动作类别。方法基于时序移位模块(TSM),通过引入背景类并划分固定长度不重叠区间来适配TAL任务。采用多任务学习框架,联合优化场景分类与动作定位,利用动作与环境间的上下文线索提升性能。最终通过加权集成策略融合多个模型,增强预测一致性与鲁棒性。本方法在初赛与延展赛中均排名第一,验证了多任务学习、高效骨干网络与集成学习结合在TAL任务中的有效性。
原文摘要 · Abstract (English)
We present our solution to the BinEgo-360 Challenge at ICCV 2025, which focuses on temporal action localization (TAL) in multi-perspective and multi-modal video settings. The challenge provides a dataset containing panoramic, third-person, and egocentric recordings, annotated with fine-grained action classes. Our approach is built on the Temporal Shift Module (TSM), which we extend to handle TAL by introducing a background class and classifying fixed-length non-overlapping intervals. We employ a multi-task learning framework that jointly optimizes for scene classification and TAL, leveraging contextual cues between actions and environments. Finally, we integrate multiple models through a weighted ensemble strategy, which improves robustness and consistency of predictions. Our method is ranked first in both the initial and extended rounds of the competition, demonstrating the effectiveness of combining multi-task learning, an efficient backbone, and ensemble learning for TAL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。