arXiv:2601.10809cs.CL2026-01被引 1

让对话模型更简洁反而显得不专业,风格特征间存在隐藏副作用。

A Concise Agent is Less Expert: Revealing Side Effects of Using Style Features on Conversational Agents

  • 通过控制实验研究不同风格提示的相互影响
  • 提示简洁会显著降低用户感知的专业性
  • 适合关注对话模型风格控制风险的研究者

风格特征如友好、有用或简洁广泛用于引导大语言模型对话代理行为,但其未预期的副作用仍不清楚。本文首次系统研究跨风格的副作用,从ACL Anthology中梳理127篇论文,识别出12个常用风格特征。通过在任务导向和开放域场景下生成受控合成对话,并采用双模型评分框架量化单一风格提示对其他风格的影响。结果发现:提示简洁会显著降低感知专业性,风格特征并非独立,而是深度耦合。为此,我们构建了CASSE数据集以记录这些复杂交互。进一步评估基于提示与激活调控的缓解策略,发现虽可部分恢复被抑制特质,但常损害目标风格。该研究挑战了大模型风格控制的可靠性,呼吁采用多目标、更严谨的方法进行安全精准的风格引导。

原文摘要 · Abstract (English)

Style features such as friendly, helpful, or concise are widely used in prompts to steer the behavior of Large Language Model (LLM) conversational agents, yet their unintended side effects remain poorly understood. In this work, we present the first systematic study of cross-feature stylistic side effects. We conduct a comprehensive survey of 127 conversational agent papers from ACL Anthology and identify 12 frequently used style features. Using controlled, synthetic dialogues across task-oriented and open domain settings, we quantify how prompting for one style feature causally affects others via a pairwise LLM as a Judge evaluation framework. Our results reveal consistent and structured side effects, such as prompting for conciseness significantly reduces perceived expertise. They demonstrate that style features are deeply entangled rather than orthogonal. To support future research, we introduce CASSE (Conversational Agent Stylistic Side Effects), a dataset capturing these complex interactions. We further evaluate prompt based and activation steering based mitigation strategies and find that while they can partially restore suppressed traits, they often degrade the primary intended style. These findings challenge the assumption of faithful style control in LLMs and highlight the need for multi-objective and more principled approaches to safe, targeted stylistic steering in conversational agents.

风格控制大模型对话系统副作用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。