构建开放树状结构视觉分解数据集,支持任意粒度的图像组件分解。
COCOTree: A Dataset and Benchmark for Open Tree-Structured Visual Decomposition

- 用大视觉语言模型与SAM结合实现全自动标注,突破人工瓶颈。
- 生成21,000张图像、180万结构节点,覆盖3500+标签,捕捉长尾复杂组合。
- 提出OTQ评估指标,兼顾掩码精度、标签准确与结构一致性,适合视觉理解研究者。
我们正式提出并实现了开放树状结构视觉分解任务,将图像分割为具有自由粒度和灵活性的层次化视觉组件树。首先,通过融合大视觉语言模型(LVLMs)的语义推理与SAM 3的精确几何定位,开发出全自动生成流水线,克服了手动标注带来的认知与物理瓶颈。其次,基于该流水线构建了大规模基准COCOTree,包含超过21,000张图像和180万个结构节点,涵盖超过3,500个独特标签,有效捕捉复杂物理组合的长尾分布。严谨的人工评估表明,生成标注与人类结构判断高度一致。第三,我们提出标准化评估协议,引入开放树质量(OTQ)指标,综合评估掩码精度、标签准确率和结构一致性。相关数据集与代码已开源至https://github.com/melonkick3090/COCOTree。
原文摘要 · Abstract (English)
We formalize and enable the task of open tree decomposition, which segments an image into hierarchical trees of visual components with unconstrained granularity and flexibility. Specifically, we provide the foundation benchmark for this new paradigm with the following three key contributions. First, we overcome the prohibitively high cognitive and physical bottlenecks of manual annotation by developing a fully automated generation pipeline that synergizes the semantic reasoning of Large Vision-Language Models (LVLMs) with the precise geometric grounding of SAM 3. Second, leveraging this pipeline, we construct COCOTree, a massive-scale benchmark featuring over 21K images and 1.8M structural nodes. By embracing an open-vocabulary space of over 3.5K unique labels, it successfully captures the long-tail distribution of complex physical assemblies. Notably, rigorous human evaluation confirms our generated annotations demonstrate strong alignment with human structural judgment. Third, we establish a standardized evaluation protocol by proposing the Open Tree Quality (OTQ) metric, which jointly assesses mask precision, label accuracy, and structural consistency. We release our dataset and benchmark code at https://github.com/melonkick3090/COCOTree.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。