让3D场景生成结构化文字描述,实现视觉与语言的精准映射。
Holo-Captioning: Toward the Text Equivalent of 3D Scenes

- 构建实例感知的解耦生成流程,逐个描述物体及其关系。
- 在1.5万+3D场景上训练,显著超越现有3D图文模型。
- 专设评估指标与人工标注集,确保结果可信可靠。
本文提出全新任务holo-captioning,旨在为3D场景生成结构化文本描述,全面涵盖场景中所有实体的语义标签、空间位置、属性及相互关系。为此,我们设计了一个高效的描述生成引擎,用于生成单个实体和实体对的详细描述,并构建了包含超过15,000个3D场景的大规模基准数据集。在此基础上,提出HoloScribe模型,采用实例感知的解耦生成架构,并引入锚点感知的实例关联机制以识别关系实体对。同时,提出HoloScore评估指标,并提供人工标注测试集以保证评估可靠性。实验表明,HoloScribe显著优于当前最优的3D密集描述生成器和3D大语言模型通用系统,验证了该方法的有效性。
原文摘要 · Abstract (English)
This work introduces holo-captioning, a novel task that strives to seek the text equivalent of 3D scenes. As the initial step, we formulate holo-captioning as generating a structured textual description that comprehensively depicts all entities within a 3D scene -- including their semantic tags, spatial locations, attributes, and inter-entity relations. To tackle this challenging task, we first develop an effective captioning engine to produce detailed descriptions of individual entity instances and instance pairs, and contribute a large-scale benchmark comprising over 15K scenes for training and evaluation. Building upon this foundation, we propose HoloScribe, a novel model that features an instance-aware decoupled pipeline for generating structured holo-captions, and further incorporates anchor-aware instance linking to identify relational instance pairs. Additionally, we propose a comprehensive evaluation metric named HoloScore, and provide a human-curated test set to ensure reliable model assessment. Experimental results demonstrate that HoloScribe significantly outperforms state-of-the-art 3D dense captioners and 3D LLM generalists, underscoring the effectiveness of our approach. Project page: https://visual-ai.github.io/holocap/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。