通过分离用户偏好信号,提升大模型个性化生成的准确性
SteerX: Disentangled Steering for LLM Personalization
- 基于因果推断识别真正反映偏好的文本片段
- 聚焦偏好驱动信息,生成更精准的激活调整向量
- 适用于需要高精度个性化的智能助手场景
大型语言模型(LLMs)近年来取得了显著进展,广泛应用于支持用户日常和工作需求的智能助手。构建此类助手的关键在于模型个性化,因为用户偏好和需求差异巨大。激活转向技术通过直接利用模型激活空间中代表用户偏好的方向来调整其行为,是一种成本效益高的对齐方法。然而,现有方法依赖全部历史数据计算转向向量,忽略了并非所有内容都反映真实用户偏好,从而削弱了个性化信号。为此,我们提出SteerX,一种解耦式转向方法,能够将偏好驱动成分与非偏好成分分离。基于因果推断理论,SteerX估计词级因果效应以识别偏好驱动的标记,将这些离散信号转化为连贯描述,并据此引导个性化生成。通过聚焦真正偏好驱动的信息,SteerX生成更准确的激活转向向量,提升个性化效果。在两个代表性转向基线方法及真实世界数据集上的实验表明,SteerX持续提升转向向量质量,为更有效的LLM个性化提供实用方案。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown remarkable success in recent years, enabling a wide range of applications, including intelligent assistants that support users' daily life and work. A critical factor in building such assistants is personalizing LLMs, as user preferences and needs vary widely. Activation steering, which directly leverages directions representing user preference in the LLM activation space to adjust its behavior, offers a cost-effective way to align the model's outputs with individual users. However, existing methods rely on all historical data to compute the steering vector, ignoring that not all content reflects true user preferences, which undermines the personalization signal. To address this, we propose SteerX, a disentangled steering method that isolates preference-driven components from preference-agnostic components. Grounded in causal inference theory, SteerX estimates token-level causal effects to identify preference-driven tokens, transforms these discrete signals into a coherent description, and then leverages them to steer personalized LLM generation. By focusing on the truly preference-driven information, SteerX produces more accurate activation steering vectors and enhances personalization. Experiments on two representative steering backbone methods across real-world datasets demonstrate that SteerX consistently enhances steering vector quality, offering a practical solution for more effective LLM personalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。