通过协作模型动态选择,提升生成多样性与质量
Optimizing Diversity and Quality through Base-Aligned Model Collaboration
- 在生成时动态融合基础模型与对齐模型,按词元级决策来源
- 联合提升21.3%的多样性和质量,优于现有方法
- 无需重训练或复杂解码,适合开放生成任务
对齐显著提升了大语言模型的输出质量,却牺牲了多样性,尤其在开放式生成任务中导致输出高度相似。我们提出基线对齐模型协作(BACo),一种推理时的词元级模型协作框架,通过不确定性与内容信号动态结合基础模型与其对齐版本,优化多样性与质量。利用路由策略,在每个词元层面决定由哪个模型解码。先前促进多样性的方法常以牺牲质量为代价,或需昂贵解码与后训练。相比之下,BACo在单次推理中实现高质量与高多样性,且具备强可控性。我们设计了一类有效路由策略,并在三个开放式生成任务上,基于13项指标进行评估。BACo始终超越现有先进推理方法。最优路由下,其在多样性与质量上实现21.3%的联合提升,人类评估亦支持该结果。总体表明,基础模型与对齐模型的协作是优化多样性-质量权衡的有效且可控机制。
原文摘要 · Abstract (English)
Alignment has greatly improved large language models (LLMs)' output quality at the cost of diversity, yielding highly similar outputs across generations, especially in open-ended generation tasks. We propose Base-Aligned Model Collaboration (BACo), an inference-time token-level model collaboration framework that dynamically combines a base LLM with its aligned counterpart to optimize diversity and quality. Using uncertainty and content-based signals, BACo employs routing strategies to determine, at each token, which model to decode from. Prior diversity-promoting methods often improve diversity at the expense of quality or require expensive decoding or post-training. In contrast, BACo achieves both high diversity and quality post hoc within a single pass, while offering strong controllability. We introduce a family of effective routing strategies and evaluate them across three open-ended generation tasks with 13 diversity and quality metrics. BACo consistently surpasses state-of-the-art inference-time baselines. With our best router, BACo achieves a 21.3% joint improvement in diversity and quality, which is further supported by human evaluations. Overall, our results demonstrate that collaboration between base and aligned models provides an effective and controllable mechanism for optimizing the diversity-quality trade-off.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。