解决视觉定位模型泛化差问题,提升跨数据集表现。
Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding

- 设计模态注意力与互补梯度机制,保持多语义表征多样性。
- 在多个基准上达到顶尖性能,且无需特定数据微调。
- 适合需要统一视觉定位模型的研究者与开发者。
指代表达理解(REC)通常在特定数据集上进行微调,导致模型成为专精于单一数据集的专家,缺乏跨数据集泛化能力。本文从统一开放词汇定位视角重新审视REC,识别出表征退化是构建通用模型的关键障碍。为此,提出数据-模型协同设计框架:架构上引入模态注意力-对比头(mACH),实现高效的细粒度视觉-语言对齐;设计文本条件JEPA辅助流,提供互补梯度支持,在不增加推理开销的前提下保护对齐活跃表征。数据层面,构建Objects365-Caption,为Objects365添加上下文感知的指代表达,实现大规模语言监督。理论分析表明,互补梯度子空间可保持对齐能力,从而扩展表征多样性。大量实验证明,该单检查点框架在标准REC基准上表现优异,且在异构接地数据集间展现出强大泛化能力,无需针对特定基准调整。
原文摘要 · Abstract (English)
Referring Expression Comprehension (REC) is commonly studied under dataset-specific fine-tuning, resulting in specialist models with limited cross-dataset generalization. In this work, we revisit REC from the perspective of unified open-vocabulary grounding and identify representation degeneration as a key obstacle to scaling a single generalist model. To preserve representation diversity, we propose a holistic data-model co-design framework. Architecturally, we introduce the Modulated Attention-Contrastive Head (mACH) for efficient token-level vision-language alignment and a text-conditioned JEPA auxiliary stream that provides complementary gradient support to preserve alignment-active representations without inference overhead. On the data side, we introduce Objects365-Caption, enriching Objects365 with context-aware referring expressions for large-scale language supervision. We further provide a theoretical analysis showing that complementary gradient subspaces preserve alignment capacity and thereby scale representation diversity. Extensive experiments demonstrate that our single-checkpoint framework achieves highly competitive performance on standard REC benchmarks while exhibiting strong generalization across heterogeneous grounding datasets without benchmark-specific adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。