arXiv:2505.21941cs.CL2025-05被引 2

重复采样提升多语言文本生成质量,推理任务效果更显著

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation

  • 通过重复采样提升生成质量,结合验证器评估
  • 部分任务提升超35%,推理类任务依赖奖励验证器
  • 适合多语言生成与复杂推理场景研究者

通过重复采样进行推理时缩放在推理任务中表现良好,但在多语言生成中的效果仍待探索。我们在Aya Evaluation Suite和m-ArenaHard两个多语言基准上,使用基于困惑度和基于奖励的验证器评估该方法。结果表明,生成质量持续提升,某些情况下提升超过35%。对于开放式提示,基于困惑度的评分有效;而对于需要推理的任务(如数学、代码),仅基于奖励的验证器能提升性能。结果证明重复采样在多语言文本生成中具有更广泛适用性,并强调了针对任务选择合适验证器的重要性。

原文摘要 · Abstract (English)

Inference-time scaling via repeated sampling has shown promise in reasoning tasks, but its effectiveness in multilingual generation remains underexplored. We evaluate this approach using perplexity- and reward-based verifiers on two multilingual benchmarks: the Aya Evaluation Suite and m-ArenaHard. Our results show consistent quality improvements, with gains exceeding 35% in some cases. While perplexity-based scoring is effective for open-ended prompts, only reward-based verifiers improve performance on tasks requiring reasoning (e.g., math, code). Our results demonstrate the broader utility of repeated sampling for multilingual text generation and underscore the importance of selecting right verifiers for the task.

多语言生成重复采样推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。