动态调整解码焦点,让大模型生成更准确且多样的文本。
Odysseus Navigates the Sirens' Song: Dynamic Focus Decoding for Factual and Diverse Open-Ended Text Generation
- 根据层间分布差异动态调节解码重点
- 在7个数据集上同时提升事实性与多样性
- 无需额外数据或模型,可无缝集成现有方法
大型语言模型在开放领域文本生成中日益需要兼顾事实准确性与多样性。然而,现有随机解码方法难以平衡这两者。本文提出动态聚焦解码(DFD),一种即插即用的随机解码新方法,无需额外数据、知识或模型即可解决这一权衡问题。DFD基于各层间分布差异自适应调整解码焦点,利用大模型中事实知识的模块化与分层特性,在知识密集型步骤提升事实性,在依赖度较低步骤促进多样性。该方法可轻松集成至现有解码策略中,计算开销极小。在七个数据集上的大量实验表明,DFD显著提升性能,为开放领域文本生成提供了一种高效可扩展的解决方案。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly required to generate text that is both factually accurate and diverse across various open-ended applications. However, current stochastic decoding methods struggle to balance such objectives. We introduce Dynamic Focus Decoding (DFD), a novel plug-and-play stochastic approach that resolves this trade-off without requiring additional data, knowledge, or models. DFD adaptively adjusts the decoding focus based on distributional differences across layers, leveraging the modular and hierarchical nature of factual knowledge within LLMs. This dynamic adjustment improves factuality in knowledge-intensive decoding steps and promotes diversity in less knowledge-reliant steps. DFD can be easily integrated with existing decoding methods, enhancing both factuality and diversity with minimal computational overhead. Extensive experiments across seven datasets demonstrate that DFD significantly improves performance, providing a scalable and efficient solution for open-ended text generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。