分析苹果视频分割中因光照变化导致的实例漂移问题
Understanding Image2Video Domain Shift in Food Segmentation: An Instance-level Analysis on Apples
- 从实例跟踪视角分析图像训练模型在视频中的表现
- 光照变化引发遮罩闪烁,计数误差达30%以上
- 适合关注视频场景下食品分割可靠性的研究者
基于静态图像训练的食品分割模型在基准数据集上表现良好,但在视频场景中可靠性未知。实际应用如食品监控和实例计数需要时间一致性,但图像训练模型在视频中常失效。本文以苹果为例,采用实例分割与匹配跟踪框架,在视频序列上评估模型,揭示高帧级分割准确率无法保证实例身份稳定。光照变化、镜面反射和纹理模糊导致遮罩闪烁与身份分裂,使苹果计数误差显著。传统图像指标严重高估真实视频性能。研究还探索了无需全视频监督的改进方法,包括后处理时序正则化与自监督时序一致性目标。结果表明,失败根源在于忽视时序一致性的图像中心训练目标,而非模型容量。该研究揭示了食品分割研究中的关键评估缺口,推动面向视频的时序感知学习与评估协议。
原文摘要 · Abstract (English)
Food segmentation models trained on static images have achieved strong performance on benchmark datasets; however, their reliability in video settings remains poorly understood. In real-world applications such as food monitoring and instance counting, segmentation outputs must be temporally consistent, yet image-trained models often break down when deployed on videos. In this work, we analyze this failure through an instance segmentation and tracking perspective, focusing on apples as a representative food category. Models are trained solely on image-level food segmentation data and evaluated on video sequences using an instance segmentation with tracking-by-matching framework, enabling object-level temporal analysis. Our results reveal that high frame-wise segmentation accuracy does not translate to stable instance identities over time. Temporal appearance variations, particularly illumination changes, specular reflections, and texture ambiguity, lead to mask flickering and identity fragmentation, resulting in significant errors in apple counting. These failures are largely overlooked by conventional image-based metrics, which substantially overestimate real-world video performance. Beyond diagnosing the problem, we examine practical remedies that do not require full video supervision, including post-hoc temporal regularization and self-supervised temporal consistency objectives. Our findings suggest that the root cause of failure lies in image-centric training objectives that ignore temporal coherence, rather than model capacity. This study highlights a critical evaluation gap in food segmentation research and motivates temporally-aware learning and evaluation protocols for video-based food analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。