用可学习前缀让大模型逻辑判断出错,揭示其推理稳定性弱点
Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft Prefixes
- 通过可学习的连续前缀干扰模型推理,观察逻辑判断变化
- 多模型测试中,前缀使正确答案翻转率高达90%,远超随机控制
- 发现模型普遍存在对特定答案的偏好,但稳定性因模型而异
为检验正确逻辑判断在学习到的上下文影响下的表现,我们在保持模型不变的前提下,向精确标注的三段论推理基准任务前添加一个软前缀。软前缀是不透明的连续向量,通过在逻辑形式和接口的受控变化中观察其引发的行为来表征。通过分析哪些前缀有效及其泛化能力,我们揭示了学习到的上下文压力如何覆盖正确判断,并暴露模型逻辑稳定性的局限。在 Qwen3.6-35B-A3B MoE、Qwen3-8B 与 Gemma 4 31B 上,学习到的前缀能显著改变大量正确答案,且在未见过的形式和接口变化中仍有效。重复测试中,这些前缀在全部 16 个模型-方向-划分组合中均优于配对随机对照,提升 37 至 99 个百分点。Qwen3.6 MoE 的翻转率在 72% 到 90% 之间,而 Gemma 的有效性前缀保持 54% 到 56% 的翻转率,随机前缀则低于 1%。诊断测试表明,主导效应是一种广泛的答案意义偏好,而非固定符号强制或可在任务间可靠转移的逻辑操作。该偏差形式在不同模型间存在差异:在两个 Qwen 模型中,简单打分模型常能预测判断是否会翻转,但无法预测边缘移动幅度;而 Gemma 的整体响应更接近这些模型的预测。结果表明,成功软前缀的主导行为效应是广泛的答案偏好,而剩余响应则揭示了模型间显著的逻辑稳定性差异。
原文摘要 · Abstract (English)
To test how correct logical judgments respond to learned context, we prepend a soft prefix to an exactly labeled syllogistic reasoning benchmark while keeping the model fixed. Soft prefixes are opaque continuous vectors, so we characterize them through the behavior they induce across controlled variations in logical form and interface. By studying which prefixes succeed and how their effects generalize, we characterize how learned contextual pressure can override correct judgments and expose limits in a model's logical stability. Across Qwen3.6-35B-A3B MoE, Qwen3-8B, and Gemma 4 31B, learned prefixes redirect many correct answers and remain effective across unseen forms and interface changes. In repeated tests with Qwen3.6 MoE and Gemma, they outperform paired random controls in all 16 model--direction--split comparisons by 37 to 99 percentage points. Qwen3.6 MoE flip rates remain between 72% and 90% across wording and prompt changes, while Gemma validity prefixes retain 54% to 56% flip compared with less than 1% for matched random prefixes. Diagnostic tests show that the dominant effect is a broad preference for one answer meaning rather than fixed-symbol forcing or a logical operation that transfers reliably between tasks. The form of this bias differs across models. In both Qwen models, simple score models often predict which judgments will flip but not how far their margins will move, whereas Gemma's overall response is more closely approximated by the same models. These results show that the dominant behavioral effect of successful soft prefixes is a broad answer preference, while the remaining response reveals substantial model-specific differences in logical stability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。