构建挑战场景下的语义对应基准,揭示现有方法鲁棒性短板。
Towards Robust Semantic Correspondence: A Benchmark and Insights
- 设计14种极端成像条件的评测数据集,涵盖几何畸变、模糊等
- 所有方法在恶劣条件下性能显著下降,大模型微调反降低鲁棒性
- DINO比Stable Diffusion更鲁棒,融合效果最佳,通用增强无效
语义对应旨在识别不同图像间的语义关联,是计算机视觉的基础任务,支撑3D重建、目标追踪与图像编辑等应用。尽管大规模视觉模型使该任务在理想条件下表现优异,但其在复杂场景下的鲁棒性仍缺乏系统研究。本文建立了一个新基准,包含14类典型成像难题,如几何失真、图像模糊、数字伪影与环境遮挡。大量实验揭示关键洞察:(1)现有方法在恶劣条件下均出现明显性能下降;(2)使用大规模视觉模型可提升整体鲁棒性,但微调后相对鲁棒性反而下降;(3)DINO模型在相对鲁棒性上优于Stable Diffusion,二者融合实现更优绝对鲁棒性。此外,评估常见鲁棒性增强策略发现,通用数据增强无效,亟需任务专用设计。结果在本数据集与真实世界基准上均一致。
原文摘要 · Abstract (English)
Semantic correspondence aims to identify semantically meaningful relationships between different images and is a fundamental challenge in computer vision. It forms the foundation for numerous tasks such as 3D reconstruction, object tracking, and image editing. With the progress of large-scale vision models, semantic correspondence has achieved remarkable performance in controlled and high-quality conditions. However, the robustness of semantic correspondence in challenging scenarios is much less investigated. In this work, we establish a novel benchmark for evaluating semantic correspondence in adverse conditions. The benchmark dataset comprises 14 distinct challenging scenarios that reflect commonly encountered imaging issues, including geometric distortion, image blurring, digital artifacts, and environmental occlusion. Through extensive evaluations, we provide several key insights into the robustness of semantic correspondence approaches: (1) All existing methods suffer from noticeable performance drops under adverse conditions; (2) Using large-scale vision models can enhance overall robustness, but fine-tuning on these models leads to a decline in relative robustness; (3) The DINO model outperforms the Stable Diffusion in relative robustness, and their fusion achieves better absolute robustness; Moreover, We evaluate common robustness enhancement strategies for semantic correspondence and find that general data augmentations are ineffective, highlighting the need for task-specific designs. These results are consistent across both our dataset and real-world benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。