arXiv:2511.16688cs.CLcs.AI2025-11

用提示词动态引导大模型生成符合人类价值观的文本

Prompt-Based Value Steering of Large Language Models

  • 设计评分方法量化生成文本中目标价值观的体现程度
  • 对比提示词后,价值引导效果提升显著(未明确具体数值)
  • 无需微调模型或优化提示,适合实时价值调整场景

大型语言模型在需要与人类价值观对齐的应用中日益重要。尽管模型微调常被用于确保安全响应,但该方法静态且难以应对日常中动态变化的价值观与偏好。本文提出一种实用、可复现且不依赖模型的评估流程,用于判断提示词是否能有效引导生成文本向特定人类价值观靠拢,并通过量化方法衡量目标价值观在输出中的存在与增益。我们以基于Wizard-Vicuna的语言模型为例,结合舒瓦茨的基本人类价值观理论,在对话数据集上进行结构化评估。对比基线提示与显式包含价值观条件的提示,结果表明即使不修改模型或动态优化提示,也能实现价值引导。

原文摘要 · Abstract (English)

Large language models are increasingly used in applications where alignment with human values is critical. While model fine-tuning is often employed to ensure safe responses, this technique is static and does not lend itself to everyday situations involving dynamic values and preferences. In this paper, we present a practical, reproducible, and model-agnostic procedure to evaluate whether a prompt candidate can effectively steer generated text toward specific human values, formalising a scoring method to quantify the presence and gain of target values in generated responses. We apply our method to a variant of the Wizard-Vicuna language model, using Schwartz's theory of basic human values and a structured evaluation through a dialogue dataset. With this setup, we compare a baseline prompt to one explicitly conditioned on values, and show that value steering is possible even without altering the model or dynamically optimising prompts.

提示工程价值观对齐大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。