arXiv:2603.18329cs.AI2026-03

提出新基准,测试大模型推理时控制的可靠性与鲁棒性。

FaithSteer-BENCH: A Deployment-Aligned Stress-Testing Benchmark for Inference-Time Steering

  • 设计三重门控评估标准,模拟真实部署环境
  • 发现现有方法在扰动下易失效,存在虚假可控性
  • 适合关注模型部署可靠性的研究者和工程师

推理时控制被广泛认为是轻量且无参数的大语言模型行为调控机制,以往研究常认为仅通过激活层干预即可稳定实现目标行为改变。然而这些结论多基于宽松评估设置,忽略了部署约束、能力权衡和真实世界鲁棒性。为此,我们提出 extbf{FaithSteer-BENCH},一个面向部署场景的应力测试基准,通过三重门控标准——可控性、效用保持性和鲁棒性——在固定部署操作点评估多种主流控制方法。在多个模型和代表性方法上,我们揭示了若干系统性失败模式,这些在常规评估中被掩盖:包括虚假可控性、对无关能力造成显著认知负担,以及在轻微指令扰动、角色提示、编码变换和数据稀缺下的严重脆弱性。门控评估结果表明,现有方法在面向部署的实际场景中并不可靠。机制诊断进一步显示,多数方法诱导的是提示条件对齐而非稳定的潜在方向偏移,解释了其在压力下的不稳定性。FaithSteer-BENCH为未来方法设计、可靠性评估与部署导向研究提供了统一基准与更清晰的分析视角。

原文摘要 · Abstract (English)

Inference-time steering is widely regarded as a lightweight and parameter-free mechanism for controlling large language model (LLM) behavior, and prior work has often suggested that simple activation-level interventions can reliably induce targeted behavioral changes. However, such conclusions are typically drawn under relatively relaxed evaluation settings that overlook deployment constraints, capability trade-offs, and real-world robustness. We therefore introduce \textbf{FaithSteer-BENCH}, a stress-testing benchmark that evaluates steering methods at a fixed deployment-style operating point through three gate-wise criteria: controllability, utility preservation, and robustness. Across multiple models and representative steering approaches, we uncover several systematic failure modes that are largely obscured under standard evaluation, including illusory controllability, measurable cognitive tax on unrelated capabilities, and substantial brittleness under mild instruction-level perturbations, role prompts, encoding transformations, and data scarcity. Gate-wise benchmark results show that existing methods do not necessarily provide reliable controllability in deployment-oriented practical settings. In addition, mechanism-level diagnostics indicate that many steering methods induce prompt-conditional alignment rather than stable latent directional shifts, further explaining their fragility under stress. FaithSteer-BENCH therefore provides a unified benchmark and a clearer analytical lens for future method design, reliability evaluation, and deployment-oriented research in steering.

推理控制模型评估部署可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。