用可微搜索自动发现高效多模态模型的令牌压缩方法
Differentiable Efficient Operator Search

- 将多种手动压缩操作统一为可微搜索空间
- 在极端视觉令牌压缩下仍保持高精度和效率
- 适合追求高效多模态推理的开发者使用
高效多模态基础模型通常依赖于手动设计的令牌压缩算子,如剪枝、合并、池化与自适应加权。尽管这些算子形式不同,我们发现它们可被解释为共享算子空间中的不同范式。基于此,我们提出可微高效的算子搜索(Efficient Operator Search),联合搜索何时压缩、保留多少令牌以及如何处理压缩后的信息。该搜索空间参数化层激活、保留预算及算子行为,搜索策略在单侧预算与成本约束下优化任务性能。该框架能恢复典型的手动设计基线,并发现超越单一设计的混合算子。在多模态基准上的实验表明,所搜算子在极端视觉令牌压缩下仍实现有竞争力的精度-效率权衡,表明高效多模态推理可从手工设计转向可微算子搜索。
原文摘要 · Abstract (English)
Efficient multimodal foundation models often rely on manually designed token-reduction operators, such as pruning, merging, pooling, and adaptive reweighting. Although these operators appear different, we show that they can be interpreted as distinct regimes of a shared operator space. Based on this view, we introduce Efficient Operator Search, a differentiable framework that jointly searches where to reduce tokens, how many tokens to retain, and how reduced token information should be processed. The proposed search space parameterizes layer activation, retention budget, and operator behavior, while the search policy optimizes task performance under one-sided budget and cost constraints. This formulation recovers representative hand-designed baselines as special cases and further discovers hybrid operators beyond isolated manual designs. Experiments on multimodal benchmarks show that the searched operators achieve competitive accuracy-efficiency trade-offs, especially under aggressive visual-token reduction. These results suggest that efficient multimodal inference can be reframed from manual operator design to differentiable operator search.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。