arXiv:2508.04530cs.CL2025-08被引 1

分离风格与事实表征,实现生成内容既风格鲜明又保持真实。

Balancing Stylization and Truth via Disentangled Representation Steering

  • 通过正交去耦分解,分离模型中风格与事实的潜在空间。
  • 在多语言多风格下显著降低风格化导致的事实性下降。
  • 适合需要精细控制生成风格且不牺牲准确性的场景。

通过表征编辑生成大语言模型的风格化输出,是实现细粒度输出控制的有前景方法。然而,存在固有权衡:施加鲜明风格常导致真实性下降。现有表征编辑方法盲目注入风格信号,忽视其对真实性的副作用,频繁污染模型的核心真实性表征,造成答案正确率降低。我们称此现象为风格化引发的真实性崩溃。我们将其归因于特定关键注意力头中风格与真实方向的潜在耦合,提出StyliTruth机制,在模型表征空间中通过正交去耦过程分离风格相关与事实相关子空间。该分解使风格与事实可在各自子空间中独立控制,最小化相互干扰。通过设计每个子空间内的自适应、分词级引导向量,动态精确调控生成过程,同时维持风格保真度与真实性。我们在多种风格和语言上验证了该方法。大量实验与分析表明,StyliTruth显著减少了风格化引发的真实性崩溃,在平衡风格遵循与真实性方面优于现有推理时干预方法。

原文摘要 · Abstract (English)

Generating stylized large language model (LLM) responses via representation editing is a promising way for fine-grained output control. However, there exists an inherent trade-off: imposing a distinctive style often degrades truthfulness. Existing representation editing methods, by naively injecting style signals, overlook this collateral impact and frequently contaminate the model's core truthfulness representations, resulting in reduced answer correctness. We term this phenomenon stylization-induced truthfulness collapse. We attribute this issue to latent coupling between style and truth directions in certain key attention heads, and propose StyliTruth, a mechanism that preserves stylization while keeping truthfulness intact. StyliTruth separates the style-relevant and truth-relevant subspaces in the model's representation space via an orthogonal deflation process. This decomposition enables independent control of style and truth in their own subspaces, minimizing interference. By designing adaptive, token-level steering vectors within each subspace, we dynamically and precisely control the generation process to maintain both stylistic fidelity and truthfulness. We validate our method on multiple styles and languages. Extensive experiments and analyses show that StyliTruth significantly reduces stylization-induced truthfulness collapse and outperforms existing inference-time intervention methods in balancing style adherence with truthfulness.

风格化生成真实性表征解耦大模型控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。