arXiv:2601.12758cs.CLcs.AI2026-01被引 1

让大模型同时表达多元价值观,无需训练即可控制输出倾向。

VISPA: Pluralistic Alignment via Automatic Value Selection and Activation

  • 通过动态选择和内部激活控制,实现无训练的多视角对齐。
  • 在医疗等多个领域验证,能有效支持所有类型的多元对齐模式。
  • 适配不同模型、价值取向与启动方式,可扩展性强。

随着大语言模型在高风险领域的广泛应用,其输出应反映人类偏好的多样性而非单一平均值。然而,实现这种多元性仍具挑战:现有方法仅涵盖有限价值或依赖提示干预,缺乏对价值表达的有效控制与表征。为此,我们提出VISPA——一种无需训练的多元对齐框架,通过动态选择与模型内部激活的引导,实现对价值表达的直接控制。在涵盖多个模型与评估场景的广泛实证研究中,VISPA在医疗等领域均表现出色,支持所有类型的多元对齐模式。进一步分析表明,VISPA可适应不同的引导方式、模型及价值体系。结果说明,通过内部激活机制即可实现多元对齐,为构建服务多元群体的语言模型提供了可扩展路径。

原文摘要 · Abstract (English)

As large language models are increasingly used in high-stakes domains, it is essential that their outputs reflect not average} human preference, rather range of varying perspectives. Achieving such pluralism, however, remains challenging. Existing approaches consider limited values or rely on prompt-level interventions, lacking value control and representation. To address this, we introduce VISPA, a training-free pluralistic alignment framework, that enables direct control over value expression by dynamic selection and internal model activation steering. Across extensive empirical studies spanning multiple models and evaluation settings, we show VISPA is performant across all pluralistic alignment modes in healthcare and beyond. Further analysis reveals VISPA is adaptable with different steering initiations, model, and/or values. These results suggest that pluralistic alignment can be achieved through internal activation mechanisms, offering a scalable path toward language models that serves all.

多元对齐价值控制无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。