arXiv:2609.05800cs.AI2026-09

让大模型同时对多个价值观精准调控,避免干扰。

Spillover-Aware Multi-Value Steering for Pluralistic LLM Alignment

论文配图:Spillover-Aware Multi-Value Steering for Pluralistic LLM Alignment
图 1 · 摘自论文原文
  • 通过修正方向间几何纠缠,实现多维度无溢出控制
  • 在气候议题上将调控效果从+5.9%提升至+14.0%
  • 无需微调或人工提示,自动完成价值识别与调节

激活值导向在推理时通过向隐藏状态添加学习到的方向来控制大模型行为,但现有方法一次仅处理一个概念。多元对齐需求不同利益相关方对不同价值的强调,需同时操控多个维度。我们发现简单导向会产生显著溢出:意图作用于某一价值的效果会泄漏至其他价值。这类似于因果推断中的处理效应与溢出效应分解。我们发现溢出源于导向方向间的几何纠缠,由其格拉姆矩阵刻画,并从激活范数惩罚目标中导出零成本修正,可精确解耦各方向贡献。我们的端到端流程无需微调、无需奖励模型、也无需人工提示工程:仅需领域问题即可自动发现价值维度、提取方向、诊断纠缠并应用修正后的导向。在气候话语任务中,该修正使净调控效果从+5.9%提升至+14.0%,经十万级成对判断验证。

原文摘要 · Abstract (English)

Activation steering controls LLM behavior at inference time by adding learned directions to hidden states, but existing methods handle one concept at a time. Pluralistic alignment, where different stakeholders need different value emphases, requires steering multiple dimensions simultaneously. We show that naive steering produces substantial spillover: the effect intended for one value leaks into others. This parallels the treatment-versus-spillover decomposition in causal inference. We trace spillover to geometric entanglement of steering directions, captured by their Gram matrix, and derive a zero-cost correction from an activation-norm-penalized objective that decouples each direction's contribution exactly. Our end-to-end pipeline requires no fine-tuning, no reward model, and no manual prompt engineering: given only domain questions, it automatically discovers value dimensions, extracts directions, diagnoses entanglement, and applies corrected steering. On climate discourse, the correction improves the net steering effect from +5.9% to +14.0%, validated over 100,000 pairwise judgments.

大模型对齐激活调控多值控制无监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。