改进文本生成的鲁棒性,通过双向通道建模提升解码质量。
Noisy-Channel Minimum Bayes Risk Decoding
- 将MBR解码拆分为四个交互组件,引入双向信息流。
- 不同评估指标下各通道贡献差异显著,任务间保持一致。
- 可针对不同任务和指标定制权重,提升生成效果。
最小贝叶斯风险(MBR)解码通过在采样伪参考上选择最大化期望效用的假设,相比最大后验(MAP)解码能产生更鲁棒、更高质量的文本。然而,现有设计存在不一致:假设选择基于给定伪参考计算期望效用,而常用评估指标如BLEU和COMET具有方向性不对称性。因此需同时考虑假设到参考与参考到假设的双向影响。本文提出一种噪声信道分解的MBR解码方法,自然融合双向效应。将MBR解码分解为四个相互作用的组成部分:假设到参考似然、参考到假设似然、假设先验和参考先验。该分解统一解释了现有MBR变体,并通过分离各通道贡献实现度量和任务特异性可解释性。全面分析表明,各通道贡献在不同度量下呈现不同特征,但在不同任务间保持一致,提示适当加权可能优于原始MBR解码。
原文摘要 · Abstract (English)
Minimum Bayes Risk (MBR) decoding yields more robust and higher-quality text generation than maximum a posteriori (MAP) decoding by selecting hypotheses that maximize expected utility over sampled pseudo-references. However, there exists a discrepancy in the design: hypothesis selection calculates expected utility scores conditioned on given pseudo-references, while commonly used evaluation metrics, e.g., BLEU and COMET, are asymmetric. Therefore, it is important to consider both hypothesis-to-reference and reference-to-hypothesis directional effects. In this study, we introduce a noisy channel decomposition of MBR decoding that naturally incorporates bidirectional effects to account for these asymmetries. We decompose MBR decoding into four interacting components: hypothesis-to-reference likelihood, reference-to-hypothesis likelihood, hypothesis prior, and reference prior. This decomposition provides a unified interpretation of existing MBR variants and enables metric- and task-specific interpretability by isolating the contribution of each channel. Our comprehensive analysis reveals that channel-wise contributions exhibit distinct characteristics across metrics while remaining consistent across tasks, and suggests that appropriate channel weighting may lead to improvements over original MBR decoding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。