arXiv:2601.14007cs.CL2026-01

揭示大模型如何将抽象价值落地为具体行为,为可控对齐提供新机制。

BACH-V: Bridging Abstract and Concrete Human-Values in Large Language Models

  • 分三步检验大模型对抽象价值的理解:抽象理解、具象锚定、具体应用。
  • 探测发现抽象值能跨层级识别具体事件中的价值观,说明存在结构化表征。
  • 干预实验证明可改变具体决策但难改抽象认知,表明抽象值是稳定锚点。

大语言模型是否真正理解抽象概念,还是仅将其当作统计模式处理?我们提出一种抽象-具象对齐框架,将概念理解分解为三类能力:抽象-抽象(A-A)、抽象-具象(A-C)和具象-具象(C-C)。以人类价值观为测试基准——因其语义丰富且关乎对齐——我们采用探测(检测内部激活中价值痕迹)与调控(修改表征以改变行为)方法。在六种开源大模型与十种价值维度上,探测显示:仅用抽象价值描述训练的诊断探针,可准确识别具体事件叙述与决策推理中的相同价值观,证明跨层级迁移存在。调控实验揭示不对称性:干预价值表征能因果改变具体判断与决策(A-C, C-C),却无法改变抽象解释(A-A),表明编码的抽象价值作为稳定锚点而非可变激活。研究结果表明,大模型保持结构化的价值表征,能连接抽象与行动,为构建可解释、可泛化、可控制的价值驱动自主系统提供了机制基础。

原文摘要 · Abstract (English)

Do large language models (LLMs) genuinely understand abstract concepts, or merely manipulate them as statistical patterns? We introduce an abstraction-grounding framework that decomposes conceptual understanding into three capacities: interpretation of abstract concepts (Abstract-Abstract, A-A), grounding of abstractions in concrete events (Abstract-Concrete, A-C), and application of abstract principles to regulate concrete decisions (Concrete-Concrete, C-C). Using human values as a testbed - given their semantic richness and centrality to alignment - we employ probing (detecting value traces in internal activations) and steering (modifying representations to shift behavior). Across six open-source LLMs and ten value dimensions, probing shows that diagnostic probes trained solely on abstract value descriptions reliably detect the same values in concrete event narratives and decision reasoning, demonstrating cross-level transfer. Steering reveals an asymmetry: intervening on value representations causally shifts concrete judgments and decisions (A-C, C-C), yet leaves abstract interpretations unchanged (A-A), suggesting that encoded abstract values function as stable anchors rather than malleable activations. These findings indicate LLMs maintain structured value representations that bridge abstraction and action, providing a mechanistic and operational foundation for building value-driven autonomous AI systems with more transparent, generalizable alignment and control.

价值观对齐大模型机制抽象具象

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。