针对动作引导视频分割中的标注噪声问题,提出首个鲁棒性基准与解决方案。
Segment-to-Act: Label-Noise-Robust Action-Prompted Video Segmentation Towards Embodied Intelligence
- 构建动作引导分割的噪声模拟机制,涵盖文本与掩码两类噪声。
- 提出并验证多种抗噪学习策略,揭示其在前景背景间的权衡特性。
- 设计并行掩码头机制,有效缓解边界模糊带来的标注噪声影响。
具身智能依赖于对交互中物体的精确分割。基于动作的视频对象分割通过关联语义与动作提升准确性,但严重依赖大规模标注,而这些标注存在成本高、不一致及多模态噪声(如不精确掩码、指代歧义)等问题。本文首次系统研究动作引导视频分割中的标注噪声,提出两类噪声:文本提示噪声(类别误标与同类别名词替换)和掩码标注噪声(边界扰动以模拟不精确监督)。贡献包括:1)定义两种噪声类型;2)构建首个噪声环境下的基准数据集 ActiSeg-NL,适配六种抗噪学习策略,并制定文本、边界及混合噪声下的评估协议;3)通过全面分析揭示噪声类型与失败模式的关系,提出并行掩码头机制(PMHM)应对掩码噪声。定性评估显示,边界扰动导致边界泄漏与定位错误,文本误标引发身份混淆。对比分析表明不同策略具有不同鲁棒性特征,由前景-背景权衡决定,部分策略平衡表现,部分则以牺牲背景精度换取前景准确率。相关基准与代码将公开发布。
原文摘要 · Abstract (English)
Embodied intelligence relies on accurately segmenting objects actively involved in interactions. Action-based video object segmentation addresses this by linking segmentation with action semantics, but it depends on large-scale annotations and prompts that are costly, inconsistent, and prone to multimodal noise such as imprecise masks and referential ambiguity. To date, this challenge remains unexplored. In this work, we take the first step by studying action-based video object segmentation under label noise, focusing on two sources: textual prompt noise (category flips and within-category noun substitutions) and mask annotation noise (perturbed object boundaries to mimic imprecise supervision). Our contributions are threefold. First, we introduce two types of label noises for the action-based video object segmentation task. Second, we build up the first action-based video object segmentation under a label noise benchmark ActiSeg-NL and adapt six label-noise learning strategies to this setting, and establish protocols for evaluating them under textual, boundary, and mixed noise. Third, we provide a comprehensive analysis linking noise types to failure modes and robustness gains, and we introduce a Parallel Mask Head Mechanism (PMHM) to address mask annotation noise. Qualitative evaluations further reveal characteristic failure modes, including boundary leakage and mislocalization under boundary perturbations, as well as occasional identity substitutions under textual flips. Our comparative analysis reveals that different learning strategies exhibit distinct robustness profiles, governed by a foreground-background trade-off where some achieve balanced performance while others prioritize foreground accuracy at the cost of background precision. The established benchmark and source code will be made publicly available at https://github.com/mylwx/ActiSeg-NL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。