arXiv:2512.24617cs.LGcs.AI2025-12被引 12

让大模型动态识别语义概念,提升推理效率。

Dynamic Large Concept Models: Latent Reasoning in an Adaptive Semantic Space

  • 基于隐空间发现可变长度语义概念,将计算从词元转向压缩概念空间。
  • 在相同算力下,12个零样本基准平均提升2.69%性能。
  • 提出首个压缩感知缩放定律,支持高效算力分配与跨架构迁移。

大语言模型对所有词元采用统一计算,但语言信息密度极不均匀。这种词元均质化计算方式在局部可预测区域浪费算力,而对语义关键转换则计算不足。本文提出动态大概念模型(DLCM),一种分层语言建模框架,通过隐表示学习语义边界,将计算从词元转移至压缩的概念空间,使推理更高效。DLCM端到端发现可变长度概念,无需预定义语言单元。层级压缩从根本上改变缩放规律,我们首次提出压缩感知缩放定律,解耦词元级容量、概念级推理能力与压缩比,实现固定浮点运算量下的合理算力分配。为稳定训练该异构结构,我们进一步设计解耦μP参数化方法,支持跨宽度与压缩率的零样本超参迁移。在实际设置(压缩比R=4,平均每概念4个词元)下,DLCM将约三分之一推理算力重分配至更高容量的推理主干网络,在匹配推理浮点运算量条件下,12个零样本基准平均提升2.69%。

原文摘要 · Abstract (English)

Large Language Models (LLMs) apply uniform computation to all tokens, despite language exhibiting highly non-uniform information density. This token-uniform regime wastes capacity on locally predictable spans while under-allocating computation to semantically critical transitions. We propose $\textbf{Dynamic Large Concept Models (DLCM)}$, a hierarchical language modeling framework that learns semantic boundaries from latent representations and shifts computation from tokens to a compressed concept space where reasoning is more efficient. DLCM discovers variable-length concepts end-to-end without relying on predefined linguistic units. Hierarchical compression fundamentally changes scaling behavior. We introduce the first $\textbf{compression-aware scaling law}$, which disentangles token-level capacity, concept-level reasoning capacity, and compression ratio, enabling principled compute allocation under fixed FLOPs. To stably train this heterogeneous architecture, we further develop a $\textbf{decoupled $μ$P parametrization}$ that supports zero-shot hyperparameter transfer across widths and compression regimes. At a practical setting ($R=4$, corresponding to an average of four tokens per concept), DLCM reallocates roughly one-third of inference compute into a higher-capacity reasoning backbone, achieving a $\textbf{+2.69$\%$ average improvement}$ across 12 zero-shot benchmarks under matched inference FLOPs.

大模型语义压缩推理优化缩放定律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。