纠正正式语体数据集的标注偏差,提升模型生成真正正式文本的能力。
Casual as an Anchor: Resolving Supervision Misalignment in Formality Transfer Dataset

- 将语体视为连续谱而非二元对立,引入'随意'作为中间状态明确标注信号。
- 新数据集3LF使GPT-4.1-nano在非正式到正式转换上F1从0.06提升至0.88。
- 揭示现有基准因标注设计缺陷导致模型生成伪正式文本,适合可控文本生成研究者。
正式语体转换常被视作非正式与正式之间的对称双向任务。我们指出,现有基准如GYAFC的标注存在设计缺陷:人类重写采用二元标签,反映的是风格相对变化而非绝对正式性定义。这导致模型学习生成满足标签但并非真正正式的语言。通过基于人类对正式性的理解重新评估基准标签,发现显著偏差,并导致各类模型在非正式到正式转换中持续失败。为此,我们将语体重构为三阶连续谱——非正式、随意、正式,其中‘随意’作为明确的中间状态以澄清监督信号。基于此框架,我们构建了3LF数据集,提供跨三级的平行标注。在3LF上训练可显著减少非正式到正式的失败,提升与人类感知的一致性。例如,GPT-4.1-nano在该方向的F1从0.06提升至0.88,尽管3LF远小于GYAFC。进一步证明这些提升无法仅通过上下文学习实现,并分析了由歧义引发的错误和语义扭曲。结果表明,监督设计深刻影响风格对齐,强调在可控文本生成中需注重对齐意识的数据集构建。
原文摘要 · Abstract (English)
Formality transfer is commonly framed as a symmetric bidirectional task between informal and formal registers. We argue that this framing conceals a supervision design flaw in existing benchmarks such as GYAFC: binary human rewrites encode relative stylistic shifts rather than absolute human notions of formality. Consequently, models learn to generate pseudo-formal outputs that satisfy benchmark labels while failing to produce genuinely formal language. We quantify this misalignment by re-evaluating benchmark formal labels under a human-aligned definition of formality, revealing substantial discrepancies that propagate to consistent informal-to-formal failures across model families. To address this issue, we reconceptualize formality transfer as a graded dimension rather than a binary attribute. We introduce a three-level spectrum: informal, casual, and formal, where casual serves as an explicit intermediate state that clarifies supervision signals. Based on this framework, we introduce 3LF, a dataset providing parallel supervision across all three levels. Training on 3LF substantially reduces informal-to-formal failures and improves alignment with human perception. For example, GPT-4.1-nano improves from 0.06 to 0.88 F1 in the informal-to-formal direction despite 3LF being significantly smaller than GYAFC. We further demonstrate that these gains cannot be reproduced through in-context learning alone and provide qualitative analyses of ambiguity-driven errors and meaning distortions. Overall, our findings demonstrate how supervision design shapes stylistic alignment and highlight the importance of alignment-aware benchmark construction in controllable text generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。