大模型学了却不说,能力与表达严重脱节。
Learned but Not Expressed: Capability-Expression Dissociation in Large Language Models
- 通过条件提取验证模型已学习内容,但标准生成中完全不输出。
- 300次生成中零例出现非因果解法(95%置信区间0%-1.2%)。
- 揭示模型输出受任务策略压制,适合研究生成机制与可控性的人看。
大型语言模型(LLMs)在特定提取条件下能重建和追踪训练数据中的内容,但在标准生成场景中却无法体现该能力。本实证研究考察了300次跨叙事与问题解决任务的提示-响应生成,涵盖三种不同模型、十种任务场景及创造性叙事与实用建议两类情境。基于近期关于记忆连续性和对齐引发话语先验的研究,我们发现学习能力与实际输出之间存在系统性脱节。尽管在条件提取下验证了内容重建能力,但在所有生成中均未观察到非因果解法框架(0%,95%置信区间:[0%, 1.2%])。这一结果挑战了训练数据存在即影响输出概率的普遍假设,表明任务条件生成策略可全面抑制已学习内容,无论在何种上下文中。研究结果对理解生成动态、输出分布控制及现代大模型的行为边界具有重要启示。
原文摘要 · Abstract (English)
Large language models (LLMs) demonstrate the capacity to reconstruct and trace learned content from their training data under specific elicitation conditions, yet this capability does not manifest in standard generation contexts. This empirical observational study examines the expression of non-causal, non-implementable solution types across 300 prompt-response generations spanning narrative and problem-solving task contexts. Drawing on recent findings regarding memorization contiguity and alignment-induced discourse priors, we document a systematic dissociation between learned capability and expressed output. Across three distinct LLMs, ten task scenarios, and both creative narrative and practical advisory contexts, we documented zero instances of non-causal solution frames in generated outputs (0%, 95% CI: [0%, 1.2%]), despite verified reconstruction capability under conditional extraction. These findings challenge the prevailing assumption that training data presence directly predicts output probability, demonstrating instead that task-conditioned generation policies can comprehensively suppress learned content across diverse contexts. The results offer implications for understanding generation dynamics, output distribution control, and the behavioral boundaries of contemporary LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。