用自监督方法让合成数据训练的光流模型更好适应真实世界
SciFlow: Semantic Cross Interference for Self-Supervised Optical Flow Domain Generalization

- 训练时将真实图像语义干扰注入合成图像,融合域间特征
- 无需真实世界标注,在多个真实场景上光流精度提升显著
- 适合需要低成本部署的视频理解系统开发者
物体与场景的运动在视频理解中蕴含关键信息,为动态环境与交互提供丰富线索。然而,由于高质量像素级光流标注成本高、稀缺,运动估计模型通常在合成域训练却部署于真实世界。解决从合成到真实世界的域泛化问题对开放世界应用至关重要。本文提出SciFlow,一种简单而有效的、不依赖网络结构的训练方法,通过自监督学习实现跨合成与真实域的运动估计泛化。具体而言,SciFlow在训练中将真实世界图像的语义干扰施加于合成图像,混合域内特征与跨域干扰,使网络适应真实场景。同时,利用几何一致性保证自监督的有效性。实验表明,SciFlow不仅显著提升模型在域变化下的鲁棒性,且在无需任何真实世界真值的情况下,大幅实现合成到真实域的泛化。
原文摘要 · Abstract (English)
Motions of objects and scenes carry essential intelligence in video understanding, offering rich cues for interpreting dynamic settings and interactions. Due to the cost and scarcity of high-quality annotation or ground truth of pixel-wise optical flow, however, motion estimation models are typically trained in synthetic domains while deployed in real-world domains. Addressing synthetic-to-real domain generalization challenges has been crucial for developing practical solutions in diverse open-world use cases. This paper introduces SciFlow, a simple yet effective, network-agnostic, training-based approach that leverages self-supervised learning to generalize motion estimation across synthetic and open-world domains. Specifically, SciFlow imposes semantic interference from open-world images onto synthetic images during training, blending indomain features with cross-domain interference, which enables the network to adapt to the real-world domains. Additionally, SciFlow utilizes geometric consistency to ensure validity of the self-supervision. Our experiment results show that SciFlow not only significantly enhances model robustness amidst domain variations, but also remarkably enables synthetic-to-real domain generalization without requiring any ground truth in the open world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。