arXiv:2504.16072cs.CVcs.AI2025-04ICCV被引 101

让AI精准描述图像视频中任意区域的细节,效果远超现有方法。

Describe Anything: Detailed Localized Image and Video Captioning

论文配图:Describe Anything: Detailed Localized Image and Video Captioning
图 1 · 摘自论文原文
  • 用焦点提示和局部视觉主干,兼顾细节与上下文信息。
  • 在7个基准上达到新最好性能,涵盖关键词到多句描述。
  • 自建数据流水线解决高质量标注数据稀缺问题,适合做细粒度视觉理解。

为图像和视频中特定区域生成详细准确描述仍是视觉语言模型的核心挑战。我们提出描述任何内容模型(DAM),专用于详细定位描述(DLC)。DAM通过两项关键创新保持局部细节与全局上下文:焦点提示确保目标区域高分辨率编码,局部视觉主干将精确定位与整体语境融合。为应对高质量DLC数据稀缺问题,我们设计基于半监督学习的数据流水线(DLC-SDP),从已有分割数据集出发,利用半监督学习扩展至未标注网络图像。我们提出DLC-Bench,一个不依赖参考描述的DLC评估基准。DAM在涵盖关键词级、短语级及详细多句描述的7个基准上均达到新最佳表现。

原文摘要 · Abstract (English)

Generating detailed and accurate descriptions for specific regions in images and videos remains a fundamental challenge for vision-language models. We introduce the Describe Anything Model (DAM), a model designed for detailed localized captioning (DLC). DAM preserves both local details and global context through two key innovations: a focal prompt, which ensures high-resolution encoding of targeted regions, and a localized vision backbone, which integrates precise localization with its broader context. To tackle the scarcity of high-quality DLC data, we propose a Semi-supervised learning (SSL)-based Data Pipeline (DLC-SDP). DLC-SDP starts with existing segmentation datasets and expands to unlabeled web images using SSL. We introduce DLC-Bench, a benchmark designed to evaluate DLC without relying on reference captions. DAM sets new state-of-the-art on 7 benchmarks spanning keyword-level, phrase-level, and detailed multi-sentence localized image and video captioning.

图像描述视频理解定位生成半监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。