用多个小模型提升文本生成质量,更全面识别错误模式。
Multi-Amateur Contrastive Decoding for Text Generation
- 用多个小模型对比大模型,捕捉更多生成错误
- 在新闻、百科、叙事任务中均优于传统方法
- 无需训练即可控制风格和内容,适合实际应用
对比解码(CD)通过利用大型专家模型与小型业余模型输出概率的差异,在不需额外训练的情况下提升开放文本生成的连贯性和流畅性。然而,其依赖单一业余模型,难以覆盖重复、幻觉、风格漂移等多重生成缺陷。本文提出多业余对比解码(MACD),通过集成多个业余模型,更全面地表征不良生成模式。MACD采用平均与共识惩罚机制融合对比信号,并将合理性约束扩展至多业余场景。此外,该框架可通过引入具有特定风格或内容偏见的业余模型实现可控生成。在新闻、百科、叙事等多个领域实验表明,MACD在流畅性、连贯性、多样性和适应性方面持续优于传统解码方法及原始CD,且无需额外训练或微调。
原文摘要 · Abstract (English)
Contrastive Decoding (CD) has emerged as an effective inference-time strategy for enhancing open-ended text generation by exploiting the divergence in output probabilities between a large expert language model and a smaller amateur model. Although CD improves coherence and fluency, its dependence on a single amateur restricts its capacity to capture the diverse and multifaceted failure modes of language generation, such as repetition, hallucination, and stylistic drift. This paper proposes Multi-Amateur Contrastive Decoding (MACD), a generalization of the CD framework that employs an ensemble of amateur models to more comprehensively characterize undesirable generation patterns. MACD integrates contrastive signals through both averaging and consensus penalization mechanisms and extends the plausibility constraint to operate effectively in the multi-amateur setting. Furthermore, the framework enables controllable generation by incorporating amateurs with targeted stylistic or content biases. Experimental results across multiple domains, such as news, encyclopedic, and narrative, demonstrate that MACD consistently surpasses conventional decoding methods and the original CD approach in terms of fluency, coherence, diversity, and adaptability, all without requiring additional training or fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。