arXiv:2601.01060cs.CL2026-01Conference of the …被引 1

让大模型学会精准控制文本风格强弱,无需成对训练数据。

Unsupervised Text Style Transfer for Controllable Intensity

  • 先用合成数据微调,再用强化学习优化风格强度区分度。
  • 在两个基准上显著提升风格转移效果,连相近强度也能分清。
  • 适合需要精细调节文本语气的场景,如文案生成、情感控制。

无监督文本风格迁移(UTST)旨在不依赖平行语料的情况下实现文本风格转换。相较于极性对立风格的迁移,可控强度的风格迁移更具挑战性,因其风格特征在不同强度层级间差异细微。针对缺乏平行数据及相邻强度难以区分的问题,我们提出SFT-then-PPO范式对大语言模型进行微调:首先利用合成平行数据进行监督微调;随后通过近端策略优化(PPO)进一步训练,设计了兼顾全局与局部风格特征的奖励函数以区分层次化强度。在两个UTST基准上的实验表明,该方法能有效提升基于大模型主干的性能,各项评估指标均有改善。即使在强度相近时,生成文本仍能表现出明显的风格差异。

原文摘要 · Abstract (English)

Unsupervised Text Style Transfer (UTST) aims to build a system to transfer the stylistic properties of a given text without parallel text pairs. Compared with text transfer between style polarities, UTST for controllable intensity is more challenging due to the subtle differences in stylistic features across different intensity levels. Faced with the challenges posed by the lack of parallel data and the indistinguishability between adjacent intensity levels, we propose a SFT-then-PPO paradigm to fine-tune an LLM. We first fine-tune the LLM with synthesized parallel data. Then, we further train the LLM with PPO, where the rewards are elaborately designed for distinguishing the stylistic intensity in hierarchical levels. Both the global and local stylistic features are considered to formulate the reward functions. The experiments on two UTST benchmarks showcase that both rewards have their advantages and applying them to LLM fine-tuning can effectively improve the performance of an LLM backbone based on various evaluation metrics. Even for close levels of intensity, we can still observe the noticeable stylistic difference between the generated text.

风格迁移大模型强化学习文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。