用物体为中心的紧凑表示,让视觉令牌压缩更高效准确。
CORE: Compact Object-centric REpresentations as a New Paradigm for Token Merging in LVLMs
- 基于物体掩码生成高语义先验,指导视觉令牌合并
- 极端压缩下仅用2.2%令牌仍保持97.4%性能
- 适合追求高效推理的LVLM部署场景
大型视觉语言模型(LVLMs)因图像分辨率增加导致视觉令牌数量呈二次增长,计算与内存开销巨大。现有令牌压缩方法多缺乏高层语义理解,常造成合并效果不佳、信息冗余或上下文丢失。为此,我们提出CORE(Compact Object-centric REpresentations),一种全新的视觉令牌压缩范式。CORE通过高效分割解码器生成物体掩码,作为高层语义先验,引导将视觉令牌合并为一组紧凑的物体中心表示。此外,创新的质心引导排序机制恢复了合并后令牌的连贯空间顺序,有效保留关键位置信息。大量实验表明,CORE不仅在六个权威基准上实现固定率压缩的新最优表现,还在自适应率设置中取得显著效率提升。即使在极端压缩下,仅保留全部视觉令牌的2.2%,仍可维持97.4%的基线性能。本工作证明了物体中心表示在高效且有效的LVLM处理中的优越性。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) usually suffer from prohibitive computational and memory costs due to the quadratic growth of visual tokens with image resolution. Existing token compression methods, while varied, often lack a high-level semantic understanding, leading to suboptimal merges, information redundancy, or context loss. To address these limitations, we introduce CORE (Compact Object-centric REpresentations), a new paradigm for visual token compression. CORE leverages an efficient segmentation decoder to generate object masks, which serve as a high-level semantic prior to guide the merging of visual tokens into a compact set of object-centric representations. Furthermore, a novel centroid-guided sorting mechanism restores a coherent spatial order to the merged tokens, preserving vital positional information. Extensive experiments show that CORE not only establishes a new state-of-the-art on six authoritative benchmarks for fixed-rate compression, but also achieves dramatic efficiency gains in adaptive-rate settings. Even under extreme compression, after aggressively retaining with only 2.2% of all visual tokens, CORE still maintains 97.4% of baseline performance. Our work demonstrates the superiority of object-centric representations for efficient and effective LVLM processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。