arXiv:2505.04260cs.HCcs.AI2025-05

让用户直接调节偏好强度,实现更自然的聊天机器人个性化。

Steerable Chatbots: Exploring Personalization Control Interfaces via LLM Activation Steering

  • 通过线性因子控制大模型输出中的偏好表达强度。
  • 用户研究显示调节方式比提示词更贴合真实偏好。
  • 适合关注个性化控制权与透明度的研究者和开发者。

个性化大语言模型响应通常需要用户用提示词描述偏好,这在冷启动阶段既费力又难以表达。我们提出一种新范式——可调节聊天机器人:不依赖用户描述,而是让用户直接通过线性因子操控偏好。方法基于激活调控(activation steering),利用标量控制偏好在模型输出中体现的强弱。我们首先验证了激活调控在计算上的可行性,随后探索如何将该因子暴露给用户。我们设计了三种界面原型,分别在用户主导/系统驱动、静态/自适应两个维度上变化。一项包含14名用户的受控实验表明,在冷启动个性化任务中,调节方式相比纯提示词更能匹配用户实际偏好,同时揭示了用户对控制权、持久性和透明度的多样化需求。

原文摘要 · Abstract (English)

Personalizing LLM responses typically requires users to articulate their preferences through prompting, which can be burdensome at cold start and difficult to articulate in natural language. We introduce an alternative paradigm, steerable chatbots: rather than asking users to describe what they want, let them directly manipulate it via a linear factor. We implement this through activation steering, leveraging a linear scalar to control how strongly a preference is expressed in the LLM's output. We first assess the computational viability of activation steering as a method to control granular preference expression, then we explore how the factor can be exposed to users. We prototype three activation steering interface designs that vary on the axes of agency (user-led vs. system-driven) and fluidity (static vs. adaptive). A within-subjects user study (n=14) in cold-start personalization tasks shows the potential for steerable chatbots to align better with underlying user preferences than prompting alone, while revealing heterogeneous values around control, persistence, and transparency in LLM personalization.

个性化提示工程交互设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。