arXiv:2607.14242cs.CL2026-07

用自然语言链式概念间接引导大模型选择特定答案

Implicit Reasoning Steering via Concept Chaining

论文配图:Implicit Reasoning Steering via Concept Chaining
图 1 · 摘自论文原文
  • 通过构建概念连接段落,隐式关联问题与目标选项
  • 模型在新数据上预训练后,选择偏好显著向目标偏移
  • 无需直接提示,文本看似自然却能隐蔽操控决策

大型语言模型看似推理可靠,但在多项选择题中重复采样常出现正确与错误答案交替,暴露出其最终决策形成的脆弱性。我们研究是否可通过隐式推理引导——即使用自然语言文本,在不提供明确指令、触发词或直接答案线索的情况下,使模型偏向指定答案。提出概念链(Concept Chaining)方法:生成一段简短的连接段落,将问题中的实体通过一个或两个中间概念与目标选项关联。随后在这些连接段落上继续预训练目标模型,并评估其在原始多选题上的答案偏好变化。结果表明,即使看似自然的间接文本也能系统性地引导模型预测,且其可推断性远低于直接改写版本。这说明推理脆弱性并非仅是评测误差,而是真实存在的潜在机制,使得普通文本可悄然放大隐性偏差,隐蔽改变模型决策。

原文摘要 · Abstract (English)

Large language models often appear to reason reliably, yet on many questions repeated sampling yields both correct and incorrect answers, revealing an underlying fragility in how final decisions are formed. We study whether this fragility can be exploited through implicit reasoning steering: using natural-language text to bias a model toward a designated answer without explicit instructions, triggers, or direct answer cues. Our approach, Concept Chaining, generates a short connection paragraph that links question entities to a target option through one or two intermediate concepts. We then continue pretraining a victim model on these connection paragraphs and evaluate whether its answer preference shifts on the original multiple-choice questions. Our results show that indirect, natural-looking text can systematically steer model predictions while remaining substantially less inferable than direct paraphrases, which shows that reasoning brittleness is not merely an evaluation artifact: it creates a practical channel through which latent biases can be amplified by ordinary-looking text to covertly redirect model decisions.

模型操控推理脆弱性概念链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。