用大模型生成更吸引人又不失真的新闻标题,避免夸张误导。
LLM-guided headline rewriting for clickability enhancement without clickbait
- 通过双引导模型控制标题生成,平衡吸引力与真实性。
- 在真实新闻语料上训练,生成的标题保持语义忠实且减少夸张。
- 可调节引导权重,灵活生成从普通到吸引人的各类标题。
提升读者参与度的同时保持信息准确性,是新闻媒体可控文本生成的核心挑战。将标题优化等同于点击诱饵(clickbait)常导致夸大或误导性表述,损害编辑公信力。本文将点击诱饵视为合理吸引点过度放大的极端结果,提出一种基于大语言模型(LLM)的可控标题重写框架,采用生成时控制的FUDGE范式。通过两个辅助引导模型:(1)点击诱饵评分模型,负向引导抑制过度风格化;(2)吸引属性模型,正向引导增强目标吸引力。两者均基于精心筛选的真实新闻语料中的中性标题训练。同时,通过在可控激活预设吸引策略下使用LLM重写原始标题,合成点击诱饵变体。推理时调整引导权重,系统可在中性改写到更具吸引力但符合编辑标准的标题间连续生成。该框架为研究吸引力、语义保留与点击诱饵规避之间的权衡提供原则性方法,支持新闻场景中负责任的LLM标题优化。
原文摘要 · Abstract (English)
Enhancing reader engagement while preserving informational fidelity is a central challenge in controllable text generation for news media. Optimizing news headlines for reader engagement is often conflated with clickbait, resulting in exaggerated or misleading phrasing that undermines editorial trust. We frame clickbait not as a separate stylistic category, but as an extreme outcome of disproportionate amplification of otherwise legitimate engagement cues. Based on this view, we formulate headline rewriting as a controllable generation problem, where specific engagement-oriented linguistic attributes are selectively strengthened under explicit constraints on semantic faithfulness and proportional emphasis. We present a guided headline rewriting framework built on a large language model (LLM) that uses the Future Discriminators for Generation (FUDGE) paradigm for inference-time control. The LLM is steered by two auxiliary guide models: (1) a clickbait scoring model that provides negative guidance to suppress excessive stylistic amplification, and (2) an engagement-attribute model that provides positive guidance aligned with target clickability objectives. Both guides are trained on neutral headlines drawn from a curated real-world news corpus. At the same time, clickbait variants are generated synthetically by rewriting these original headlines using an LLM under controlled activation of predefined engagement tactics. By adjusting guidance weights at inference time, the system generates headlines along a continuum from neutral paraphrases to more engaging yet editorially acceptable formulations. The proposed framework provides a principled approach for studying the trade-off between attractiveness, semantic preservation, and clickbait avoidance, and supports responsible LLM-based headline optimization in journalistic settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。