测试大模型对几何约束的编码与执行能力,发现能解码却不等于能生成或控制。
Encoded but Not Actionable: Auditing the Decode-Generate-Steer Gap in Frozen LLMs for Geometric Constraints

- 用参数化CAD约束测试冻结LLM的隐藏状态,区分局部关系与整体约束状态。
- 预训练显著提升局部几何关系的可解码性,但整体自由度状态几乎无需预训练即可解码。
- 解码信息常无法转化为生成结果,说明模型存在‘懂但不会用’的问题,适合研究模型可解释性者阅读。
大型语言模型在结构化推理任务中表现优异,但其编码内容及其对行为的影响尚不明确。本文通过几何推理任务,以参数化CAD约束为可控测试平台,分离局部成对关系与草图级自由度(DOF)状态。对六个冻结的仅解码器型LLM的隐藏状态进行探测,考察四个属性:线性可解码性、强制选择生成、激活水平影响及行为可调控性。预训练显著提升了局部几何关系的解码能力,且该优势在随机打乱顺序的对照实验中仍存在。相比之下,草图级自由度状态在随机初始化表示中已高度可解码,预训练仅带来小幅提升,表明其解码性能主要来自非学习权重。进一步分析显示,可解码信息并不总是可行动:生成常无法体现该信息;在两个干预测试的模型中,修复特定位置激活后,其影响随深度消失,而解码性保持不变;均值差异调节也无法可靠控制输出。结果表明,在该设置下,解码性、生成性、激活影响与可调控性可能分化。本审计提供了一种区分几何结构编码失败与表达/控制失败的可控方法。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated strong performance on structured reasoning tasks, but what they encode and whether it informs model behavior remain unclear. We investigate this question through geometric reasoning, using parametric CAD constraints as a controlled testbed for separating local pairwise relations from sketch-level constraint status. By probing the hidden states of six frozen decoder-only LLMs, we examine four properties: linear decodability, forced-choice generation, activation-level influence, and behavioral steerability. Pretraining substantially improves the decoding of local geometric relations, and this advantage persists after accounting for positional cues with shuffled-order controls. In contrast, sketch-level DOF status is already highly decodable from randomly initialized representations and improves only modestly with pretraining, indicating that much of its probe performance is available without learned weights. Further analyses show that decodable information is not always actionable. Generation often fails to express this information, and on the two intervention-tested backbones, activation-restoration effects at the patched entity position vanish while decodability persists across depth. Mean-difference steering also does not reliably control outputs. These results show that decodability, generation, activation-level influence, and steerability can diverge in the tested setting. The audit provides a controlled way to distinguish failures to encode geometric structure from failures to express or control encoded information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。