arXiv:2506.03709cs.CV2025-06中稿 · CVPR被引 1

构建跨空地视角的多角度红外分割基准,评估模型泛化能力

AetherVision-Bench: An Open-Vocabulary RGB-Infrared Benchmark for Multi-Angle Segmentation across Aerial and Ground Perspectives

  • 提出空地视角下多角度红外图像的开放词汇分割评测基准
  • 发现视角差异显著影响零样本迁移模型性能,验证了跨域泛化瓶颈
  • 为具身智能感知系统提供可复现的鲁棒性测试平台

开放词汇语义分割(OVSS)基于文本描述为图像每个像素分配标签,依赖如CLIP等世界模型。然而其在跨域泛化中面临显著挑战,限制了实际应用效果。本文提出AetherVision-Bench,一个面向空中与地面视角的多角度红外分割基准,支持对不同观测角度和传感器模态下的性能进行系统评估。我们在该基准上测试了当前最先进的OVSS模型,并探究了影响零样本迁移模型性能的关键因素。本工作首次构建了鲁棒性评测基准,为未来研究提供了重要基础与洞见。

原文摘要 · Abstract (English)

Open-vocabulary semantic segmentation (OVSS) involves assigning labels to each pixel in an image based on textual descriptions, leveraging world models like CLIP. However, they encounter significant challenges in cross-domain generalization, hindering their practical efficacy in real-world applications. Embodied AI systems are transforming autonomous navigation for ground vehicles and drones by enhancing their perception abilities, and in this study, we present AetherVision-Bench, a benchmark for multi-angle segmentation across aerial, and ground perspectives, which facilitates an extensive evaluation of performance across different viewing angles and sensor modalities. We assess state-of-the-art OVSS models on the proposed benchmark and investigate the key factors that impact the performance of zero-shot transfer models. Our work pioneers the creation of a robustness benchmark, offering valuable insights and establishing a foundation for future research.

开放词汇分割多视角感知红外图像具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。