arXiv:2607.26825cs.CL2026-07

让大模型从隐含概念转向显式设计,提升可控性与人类认知对齐。

From Found to Designed: Concepts as a Design Axis for Large Language Models

  • 将概念结构从训练后挖掘转为模型设计时主动构建
  • 揭示现有方法多分散于不同阶段且缺乏统一框架
  • 适合关注模型可解释性与可控生成的研究者

大型语言模型(LLMs)蕴含丰富的类概念信息,但这些信息以分布式统计关联的形式隐式编码,而非显式的、结构化的组合概念。因此,概念结构通常是‘发现’而非‘设计’的:需在训练后通过探测或词典学习恢复,且无法保证其稳定性、组合性、可控性或与人类概念组织的一致性。本文沿两个维度对概念感知干预进行分类:概念结构是内部诱发还是外部赋予,以及引入时机。该分类揭示三类模式:推理阶段的方法仍相对未受重视;各阶段相关研究发展孤立;外部赋予方法贯穿整个流程,但常被不同术语描述。上述观察推动我们从从训练后恢复概念结构,转向在设计阶段即构建显式概念表示。

原文摘要 · Abstract (English)

Large language models (LLMs) encode rich concept-like information, but represent it implicitly through distributed statistical associations rather than as explicit, structured, compositional concepts. Consequently, concept-level structure is typically \emph{found} rather than \emph{designed}: it is recovered after training through probing or dictionary learning, with no architectural guarantee of stability, compositionality, controllability, or alignment with human conceptual organization. We organize concept-aware interventions along two dimensions: whether concept structure is internally induced or externally grounded, and the stage of the pipeline where it is introduced. This taxonomy reveals three broad patterns: inference-time approaches remain comparatively underexplored, related ideas have developed largely in isolation across pipeline stages, and externally grounded methods span the entire pipeline despite often being described under different terminology. Together, these observations motivate moving beyond recovering concept-like structure from trained models toward designing LLMs with explicit conceptual representations.

概念设计大模型可解释性结构化表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。