arXiv:2512.22650cs.LGcs.AI2025-12

通过分阶段计算优化,提升无明确反馈任务的生成质量。

Scaling Unverifiable Rewards: A Case Study on Visual Insights

  • 分阶段分配算力,用专用判别器提前淘汰低质分支
  • 在固定算力下,洞察质量均值从61.64提升至65.86
  • 适合科学发现、故事生成等缺乏标准答案的任务

大型语言模型代理可通过测试时扩展(TTS)实现复杂推理的自动化,但许多现实任务涉及多阶段流程,其最终结果无法验证,也难以训练可靠的奖励模型,导致基于判别器的迭代优化易累积误差。本文提出选择性TTS,一种基于过程的优化框架,在多代理流程中跨阶段扩展推理,而非以往重复时间上的迭代优化。通过在各阶段分配计算资源,并利用针对流程特性的判别器早期剪枝低质量路径,该方法缓解了判别器漂移问题,稳定了优化过程。基于数据科学流程,我们构建了端到端多代理系统,用于生成给定数据集的可视化图表与报告,并设计了一个与人类专家对齐的可靠LLM判别器(Kendall's τ=0.55)。实验表明,在固定计算预算下,选择性TTS将洞察质量均值从61.64提升至65.86,同时降低方差。本研究为扩展无明确反馈的复杂开放任务(如科学发现、故事生成)提供了初步范式。

原文摘要 · Abstract (English)

Large Language Model (LLM) agents can increasingly automate complex reasoning through Test-Time Scaling (TTS), iterative refinement guided by reward signals. However, many real-world tasks involve multi-stage pipeline whose final outcomes lack verifiable rewards or sufficient data to train robust reward models, making judge-based refinement prone to accumulate error over stages. We propose Selective TTS, a process-based refinement framework that scales inference across different stages in multi-agent pipeline, instead of repeated refinement over time by prior work. By distributing compute across stages and pruning low-quality branches early using process-specific judges, Selective TTS mitigates the judge drift and stabilizes refinement. Grounded in the data science pipeline, we build an end-to-end multi-agent pipeline for generating visually insightful charts and report of given dataset, and design a reliable LLM-based judge model, aligned with human experts (Kendall's τ=0.55). Our proposed selective TTS then improves insight quality under a fixed compute budget, increasing mean scores from 61.64 to 65.86 while reducing variance. We hope our findings serve as the first step toward to scaling complex, open-ended tasks with unverifiable rewards, such as scientific discovery and story generation.

多智能体生成质量无验证奖励推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。