arXiv:2607.26554cs.CV2026-07

无需训练,通过三重线索自适应压缩3D医学图像冗余特征

MedARC: Training-Free Adaptive Redundancy Compression of Visual Tokens for 3D Medical Vision-Language Models

论文配图:MedARC: Training-Free Adaptive Redundancy Compression of Visual Tokens for 3D Medical Vision-Language Models
图 1 · 摘自论文原文
  • 融合注意力、文本相关性和结构差异三类信号评估视觉标记重要性
  • 在CT-RATE和MR-RATE上减少90%以上视觉令牌数,推理提速超2倍
  • 适合大模型部署,尤其适用于需要高精度的临床诊断场景

将3D医学影像与视觉语言模型(VLMs)结合有望推动辅助诊断发展。然而,体数据生成极长的视觉标记序列,存在显著的空间与跨切片冗余。现有压缩方法通常采用统一降采样或依赖单一重要性信号,易误删临床相关或结构独特的区域。为此,我们提出MedARC:一种面向3D医学视觉语言模型的统一、无需训练的自适应冗余压缩框架。MedARC通过融合三种互补线索估计标记重要性:来自视觉编码器的自注意力,反映模型内在关注;投影后视觉标记与文本嵌入的相似性,识别查询相关区域;局部基础模型特征与整体特征中心的偏差,突出结构独特解剖。重要性分布引导显著性感知合并策略,保留信息量高的标记,整合冗余部分而非直接丢弃。在CT-RATE和MR-RATE上的实验表明,MedARC显著降低视觉标记开销与推理时间,同时保持或提升诊断性能。其多线索评分成本被处理更少标记带来的收益所覆盖,对更大语言模型更具优势。

原文摘要 · Abstract (English)

Integrating 3D medical images with vision-language models (VLMs) holds substantial promise for computer-aided diagnosis. However, volumetric images generate prohibitively long visual-token sequences with considerable spatial and inter-slice redundancy. Existing token compression methods typically apply uniform reduction or rely on a single importance signal, increasing the risk of removing regions that are clinically relevant to the query or structurally distinctive. To address this limitation, we propose MedARC, a unified, training-free framework for Adaptive Redundancy Compression of visual tokens in 3D medical VLMs. MedARC estimates token importance by integrating three complementary cues: self-attention from the VLM vision encoder, which reflects the model's intrinsic visual focus; similarity between projected visual tokens and text embeddings, which identifies query-relevant regions; and deviations of local visual foundation model features from the volume-level feature center, which highlight structurally distinctive anatomy. The resulting importance distribution guides a saliency-aware merging strategy that preserves informative tokens while consolidating redundant ones rather than simply discarding them. Experiments on CT-RATE and MR-RATE show that MedARC reduces visual-token overhead and inference time while preserving or improving diagnostic performance. Its multi-cue scoring cost is outweighed by the savings from processing fewer tokens, with greater benefits expected for larger language models.

3D医学视觉语言特征压缩自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。