arXiv:2509.24192cs.CVcs.AI2025-09中稿 · EMNLP

通过拆解句子成分,提升语言驱动目标检测对复杂查询的解析能力。

Talk in Pieces, See in Whole: Disentangled and Hierarchical Representation Learning in Language-based Object Detection

  • 将文本分解为对象、属性、关系三类组件,分层构建语义表示。
  • 在OmniLabel基准上性能提升24%,显著改善复杂描述词处理能力。
  • 适合研究多模态理解、语言结构建模的开发者参考使用。

视觉语言模型(VLMs)在多模态感知方面取得进展,尤其在使用简单语言查询进行开放词汇目标检测方面表现优异。然而,当前最先进的VLMs仍难以处理包含描述性属性和关系子句的复杂查询。为此,本文提出根据句子内部的层次关系重构语言表征:关键洞察是将文本标记解耦为对象、属性和关系三个核心成分,并聚合为分层结构的句子级表征。基于此,我们提出了TaSe(Talk in Pieces, See in Whole)框架,包含三项主要贡献:(1) 构建了一个涵盖三个层级(类别名至描述性句子)的分层合成标题数据集;(2) 设计一种三成分解耦模块,由新颖的解耦损失函数引导,将文本嵌入转换为子空间组合;(3) 在提出的分层目标指导下,将解耦后的成分聚合为分层结构嵌入。在OmniLabel基准上的实验结果显示,性能提升24%,验证了语言组合性的重要性。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have advanced multimodal perception, demonstrated by open-vocabulary object detection with simple language queries. State-of-the-art VLMs still struggle to handle complex queries involving descriptive attributes and relational clauses. To address this problem, we propose restructuring linguistic representations according to the hierarchical relations within sentences for language-based object detection. A key insight is that textual tokens should be disentangled into core components-objects, attributes, and relations-and aggregated into hierarchically structured sentence-level representations. Building on this principle, we introduce the TaSe (Talk in Pieces, See in Whole) framework with three main contributions: (1) a hierarchical synthetic captioning dataset spanning three tiers from category names to descriptive sentences; (2) the three-component disentanglement module guided by a novel disentanglement loss function, transforms text embeddings into subspace compositions; and (3) aggregating disentangled components into hierarchically structured embeddings guided by the proposed hierarchical objectives. Experimental results under the OmniLabel benchmark show a 24% performance improvement, demonstrating the importance of linguistic compositionality.

多模态目标检测语言理解表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。