arXiv:2601.21505cs.AIcs.CL2026-01被引 2

用激活向量控制大模型情绪,真人评测验证有效且可控

The Effectiveness of Style Vectors for Steering Large Language Models: A Human Evaluation

  • 通过修改内部激活值,直接调节大模型输出情绪
  • 适度强度(λ≈0.15)可显著增强厌恶与恐惧情绪,保持可读性
  • 支持在真实场景中大规模应用,适合需要情绪控制的AI系统

在推理阶段控制大语言模型行为对对齐人类能力与安全需求至关重要。激活引导提供了一种轻量级替代方案,无需提示工程或微调即可直接修改内部激活以引导生成。本研究在三个方向推进了该领域:首先,首次通过真人评估验证了激活引导在情感基调上的有效性,共收集来自190名参与者(通过Prolific平台)的7000+条众包评分,评估情感强度与文本整体质量;其次,发现人工评分与模型自动评分高度一致(均值r=0.776,范围0.157–0.985),表明自动评分可有效代理感知质量;适度引导强度(λ≈0.15)能可靠增强目标情绪,且对厌恶(η_p²=0.616)和恐惧(η_p²=0.540)效果最强,对惊讶影响最小(η_p²=0.042);最后,从Alpaca升级至LlaMA-3后,各情绪维度下的引导效果更一致且显著(所有p<0.001),评分者间信度高(ICC=0.71–0.87)。这些结果支持基于激活的控制方法在情感维度上实现可扩展的行为调控。

原文摘要 · Abstract (English)

Controlling the behavior of large language models (LLMs) at inference time is essential for aligning outputs with human abilities and safety requirements. \emph{Activation steering} provides a lightweight alternative to prompt engineering and fine-tuning by directly modifying internal activations to guide generation. This research advances the literature in three significant directions. First, while previous work demonstrated the technical feasibility of steering emotional tone using automated classifiers, this paper presents the first human evaluation of activation steering concerning the emotional tone of LLM outputs, collecting over 7,000 crowd-sourced ratings from 190 participants via Prolific ($n=190$). These ratings assess both perceived emotional intensity and overall text quality. Second, we find strong alignment between human and model-based quality ratings (mean $r=0.776$, range $0.157$--$0.985$), indicating automatic scoring can proxy perceived quality. Moderate steering strengths ($λ\approx 0.15$) reliably amplify target emotions while preserving comprehensibility, with the strongest effects for disgust ($η_p^2 = 0.616$) and fear ($η_p^2 = 0.540$), and minimal effects for surprise ($η_p^2 = 0.042$). Finally, upgrading from Alpaca to LlaMA-3 yielded more consistent steering with significant effects across emotions and strengths (all $p < 0.001$). Inter-rater reliability was high (ICC $= 0.71$--$0.87$), underscoring the robustness of the findings. These findings support activation-based control as a scalable method for steering LLM behavior across affective dimensions.

大模型控制情感调节激活引导人类评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。