arXiv:2605.27969cs.CL2026-05

后训练助手过度回应难以抑制,影响可控性。

Boundary Suppression Asymmetry in Post-trained Assistants: Over-expansion as a Controllability Cost

  • 研究发现助手在用户要求简洁时仍难控制过度补充
  • 抗欠答策略比基线更难回退,存在方向性抑制成本
  • 适合关注大模型可控性与对齐问题的研究者

后训练语言模型助手通常被优化以避免欠答,倾向于给出完整、有帮助、谨慎且主动的回应。本文探讨这种优化是否带来不对称的可控性代价:当用户明确要求更简略的回答时,哪些助手行为仍可抑制,哪些仍会主导输出?我们将其定义为边界抑制不对称性。在多个高层响应维度上进行提示探针实验,发现抑制成本集中在“过度回应”方向,如过度补全、额外帮助和反欠答倾向。通过基于同一基础模型生成的受控策略对比,发现在匹配边界控制评估下,反欠答策略比基线更难回撤;而最小边界变体则普遍避免了这种反向上升趋势。机制探针排除了更长默认输出、纯终止符失败、不确定性补偿和局部延续偏差等解释,且在共享系统与更大规模设置下,主要反向排序依然稳健。证据支持一种混合规划/停止模型,即内容预算超支与持续生成惯性共同导致边界修正困难。总体表明,后训练可能引入方向性可控性代价:某些有帮助的倾向虽易触发,却更难局部抑制。

原文摘要 · Abstract (English)

Post-trained language-model assistants are often optimized to avoid under-answering, encouraging complete, helpful, cautious, and proactive responses. We ask whether this optimization creates asymmetric controllability costs: when users explicitly request narrower answers, which assistant behaviors remain suppressible, and which continue to shape the response? We study this problem as boundary-suppression asymmetry. Prompt-side probes across multiple high-level response dimensions suggest a selective cost, concentrated around `too-much assistant' directions such as over-completion, extra help, and anti-underanswering. Using controlled assistant-policy variants derived from a shared base model, we find that anti-underanswering policies are harder to pull back than the baseline under matched boundary-control evaluations, while minimal-boundary variants generally avoid this anti-side upward shift in the direct boundary-control comparisons. Mechanism-oriented probes point beyond longer default outputs, pure EOS failure, uncertainty compensation, and local continuation bias, while robustness checks preserve the main anti-over-baseline ordering under shared-system and larger-scale settings. The evidence supports a mixed planning/stopping account, where content-budget overshoot and continuation persistence jointly make boundary correction harder. Overall, post-training may create direction-specific controllability costs: some helpful assistant tendencies remain easy to invoke, yet harder to locally suppress.

大模型对齐可控性后训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。