提出可计算评估框架,让掩码扩散模型的改进更可信。
CaRE Compute-aware Remasking Evaluation Protocol for Masked Diffusion Language Models

- 统一实际计算量(NFE),控制随机性与多指标报告
- 发现温度影响最大,计算量匹配后排名多次反转
- 揭示高熵重掩码与随机解掩码存在性能权衡
掩码扩散语言模型(MDLMs)发展迅速,但评估标准未能同步跟进。尽管MDLM已媲美自回归模型,近期七篇重掩码论文在不同步数、指标和采样温度下评估,导致策略对比不可靠,难以判断提升是算法进步还是评估偏差。本文提出CaRE——一个计算感知的评估协议,通过标准化实际函数评估次数(NFE)、强制多指标报告并显式控制随机性,对LLaDA-8B-Base与Dream-7B-Base在OpenWebText和LM1B上、4种随机性水平与3种步数预算下共7种重掩码策略进行审计。结果表明:(i) 温度解释了大部分MAUVE方差;(ii) 计算量匹配后多个策略排名逆转;(iii) 有意识重掩码与随机解掩码存在冲突,在256步、解掩温度0.25时,高熵重掩码使MAUVE降低0.296(p=0.020)。覆盖12个开源MDLM(参数量150M至8B)的CaRE排行榜显示该趋势跨架构与规模成立。研究证明当前评估易将计算与随机性选择误认为算法优势。我们公开协议、实现与排行榜,确保未来研究可复现且可比。
原文摘要 · Abstract (English)
Masked diffusion language models (MDLMs) are advancing rapidly, yet the evaluation standards needed to reliably interpret their progress have not kept pace. Despite MDLMs becoming competitive with autoregressive language models, seven recent remasking papers evaluate under incompatible settings, varying nominal step counts, metrics, and sampling temperatures without jointly controlling these factors, rendering their strategy rankings largely incomparable and leaving open whether reported gains reflect algorithmic improvements or evaluation artifacts. We present CaRE, a compute-aware evaluation framework that audits MDLM remasking strategies by standardizing actual number of function evaluations (NFE), enforcing multi-metric reporting, and explicitly controlling stochasticity. Applied to 7 remasking strategies across LLaDA-8B-Base and Dream-7B-Base at 4 stochasticity levels and 3 step budgets on OpenWebText and LM1B, CaRE reveals that: (i) temperature explains the majority of MAUVE variance, (ii) compute-matched comparisons reverse several published strategy rankings, and (iii) informed remasking and stochastic unmasking are in tension, with high-entropy remasking reducing MAUVE by 0.296 at 256 steps at unmask_temp=0.25 (p=0.020). A CaRE leaderboard covering 12 open-weight MDLMs (150M to 8B parameters) shows that this interaction direction holds across architectures and scales. These findings demonstrate that current MDLM evaluations can systematically conflate algorithmic improvements with hidden choices of compute and stochasticity. We release the evaluation protocol, implementation, and leaderboard to ensure future remasking claims are reproducible and comparable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。