证明了MBR解码能以根号n速度逼近最优解,解释其为何表现优秀。
Theoretical Guarantees for Minimum Bayes Risk Decoding
- 基于参考假设集大小n,分析MBR解码的收敛速率
- 在|Y|≫n的语言空间中,仍以O(n⁻¹/²)收敛至最优解
- 相比MAP解码,MBR收敛更快,适合追求高精度生成任务
最小贝叶斯风险(MBR)解码通过最大化潜在人类分布的期望效用值来优化输出选择。尽管已有研究通过实证证明MBR的有效性,但对其理论机制的分析仍较少。本文在一定假设下证明:当参考假设集大小为n时,MBR解码以概率趋近于1的速度收敛至最优解,收敛速率为O(n⁻¹/²),即使语言空间|Y|远大于n。该结果为先前多项实证研究中观察到的优异性能提供了理论解释。此外,本文还给出了最大后验(MAP)解码的性能差距,并与MBR解码进行比较,结果表明在某些情况下MBR收敛速度优于MAP解码。
原文摘要 · Abstract (English)
Minimum Bayes Risk (MBR) decoding optimizes output selection by maximizing the expected utility value of an underlying human distribution. While prior work has shown the effectiveness of MBR decoding through empirical evaluation, few studies have analytically investigated why the method is effective. As a result of our analysis, we show that, given the size $n$ of the reference hypothesis set used in computation, MBR decoding approaches the optimal solution with high probability at a rate of $O\left(n^{-\frac{1}{2}}\right)$, under certain assumptions, even though the language space $Y$ is significantly larger $|Y|\gg n$. This result helps to theoretically explain the strong performance observed in several prior empirical studies on MBR decoding. In addition, we provide the performance gap for maximum-a-posteriori (MAP) decoding and compare it to MBR decoding. The result of this paper indicates that MBR decoding tends to converge to the optimal solution faster than MAP decoding in several cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。