提出一种新型字典学习方法,可自动平衡稀疏性、存储与精度。
On the Sparsity-Storage-Accuracy Tradeoff in Parsimoniously Activated Dictionary Learning

- 通过引入隐变量建立全局激活模式的生成模型
- 揭示稀疏性、存储成本与重建精度间的定量权衡关系
- 无需调参,加速视觉-语言模型推理
字典学习长期从优化和概率视角研究。尽管逐元素稀疏正则化(如基于L1的稀疏编码)有明确的概率解释,但许多施加全局约束的结构化变体缺乏清晰可计算的生成视图。本文重新审视一类实用有效但理论研究不足的方法——对激活字典原子数量施加简单全局正则化,称为稀疏激活字典学习(PADL)。我们证明PADL可等价为在结构化生成模型下的最大后验估计,其辅助隐变量控制全局激活模式。该形式使我们获得原形式难以得到的一般化保证,并实现对稀疏性、存储开销与重建精度之间权衡关系的解析刻画,支持数据驱动地估计最优超参数。基于此,我们设计了一种高效且可解释的PADL算法,避免人工调参,在视觉基准上实现相近稀疏度下的更好重建性能,并进一步展示其在加速视觉-语言模型推理中的实际价值。
原文摘要 · Abstract (English)
Dictionary learning has long been studied from both optimization and probabilistic perspectives. While formulations with element-wise sparsity regularization (e.g., L1-based sparse coding) admit well-established probabilistic interpretations, many structured variants that impose global constraints lack a clear and tractable generative view. In this paper, we revisit a class of practically effective yet theoretically under-explored dictionary learning methods that impose a simple global regularization on the number of activated dictionary atoms, which we term parsimoniously activated dictionary learning (PADL). We show that PADL admits an equivalent formulation as maximum a posteriori estimation under a structured generative model, with auxiliary latent variables that govern global activation patterns. This formulation allows us to derive generalization guarantees that are difficult to obtain under the original formulation. More importantly, it yields an analytical characterization of the tradeoff between sparsity, storage cost, and reconstruction accuracy, enabling data-driven estimation of optimal hyperparameters. Based on this connection, we develop an efficient and interpretable PADL algorithm that eliminates manual hyperparameter tuning, achieving improved reconstruction performance under comparable sparsity levels on visual benchmarks. We further demonstrate its practical utility in accelerating inference for vision-language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。