arXiv:2510.00818cs.CV2025-10中稿 · ICCV被引 1

首个支持语言短语与立体图像区域对应标注的数据集

PhraseStereo: The First Open-Vocabulary Stereo Image Segmentation Dataset

  • 基于单视图数据生成立体图像对,实现双视角语义分割
  • 包含5,000组带短语标注的立体图像与对齐分割掩码
  • 适合研究多模态、几何感知与3D语义理解的学者

理解自然语言短语与图像中特定区域的对应关系是多模态语义分割的关键挑战。当前的短语定位进展主要局限于单视图图像,忽略了立体视觉提供的丰富几何线索。为此,我们提出PhraseStereo,首个将短语-区域分割引入立体图像对的数据集。PhraseStereo在PhraseCut基础上,利用GenStereo从现有单视图数据生成精确的右视图图像,从而将短语定位扩展至立体域。该新设置为多模态学习带来独特挑战与机遇,尤其在利用深度信息实现更精准、上下文感知的定位方面。通过提供带有对齐分割掩码和短语标注的立体图像对,PhraseStereo为语言、视觉与三维感知交叉领域的未来研究奠定基础,推动能联合推理语义与几何的模型发展。该数据集将在论文接受后公开发布。

原文摘要 · Abstract (English)

Understanding how natural language phrases correspond to specific regions in images is a key challenge in multimodal semantic segmentation. Recent advances in phrase grounding are largely limited to single-view images, neglecting the rich geometric cues available in stereo vision. For this, we introduce PhraseStereo, the first novel dataset that brings phrase-region segmentation to stereo image pairs. PhraseStereo builds upon the PhraseCut dataset by leveraging GenStereo to generate accurate right-view images from existing single-view data, enabling the extension of phrase grounding into the stereo domain. This new setting introduces unique challenges and opportunities for multimodal learning, particularly in leveraging depth cues for more precise and context-aware grounding. By providing stereo image pairs with aligned segmentation masks and phrase annotations, PhraseStereo lays the foundation for future research at the intersection of language, vision, and 3D perception, encouraging the development of models that can reason jointly over semantics and geometry. The PhraseStereo dataset will be released online upon acceptance of this work.

多模态立体视觉语义分割语言理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。