arXiv:2511.19489cs.SEcs.AI2025-11被引 3

用大模型当裁判,让进化算法在无标准答案场景下高效优化。

Evolution without an Oracle: Driving Effective Evolution with LLM Judges

  • 将模糊指令拆解为可验证子任务,降低大模型评分噪声。
  • 在DevAI和InfoBench上需求满足率提升至61.9%,超基线50%以上。
  • 适合无明确标准的开放域问题,如创意设计、复杂指令执行。

大型语言模型(LLM)与进化计算(EC)的结合推动了科学发现的新边界,但始终受限于对“真值”(Oracle)——即机器可计算的客观适应度函数——的依赖。本文提出:进化能否在完全主观的评价体系中蓬勃发展?我们引入MADE(多智能体分解进化)框架,通过“问题定义”机制,将模糊指令分解为具体、可验证的子要求,从而将高方差的LLM反馈转化为稳定精确的选择压力。实验结果表明,在DevAI和InfoBench等复杂基准上,MADE在软件需求满足率上从39.9%提升至61.9%,性能超越强基线超过50%,并在复杂指令遵循任务中达到95%的完美通过率。本工作验证了一种根本性范式转变:从优化“可计算指标”转向优化“可描述品质”,从而为无真实标签的开放领域开启了进化优化的大门。

原文摘要 · Abstract (English)

The integration of Large Language Models (LLMs) with Evolutionary Computation (EC) has unlocked new frontiers in scientific discovery but remains shackled by a fundamental constraint: the reliance on an Oracle--an objective, machine-computable fitness function. This paper breaks this barrier by asking: Can evolution thrive in a purely subjective landscape governed solely by LLM judges? We introduce MADE (Multi-Agent Decomposed Evolution), a framework that tames the inherent noise of subjective evaluation through "Problem Specification." By decomposing vague instructions into specific, verifiable sub-requirements, MADE transforms high-variance LLM feedback into stable, precise selection pressure. The results are transformative: across complex benchmarks like DevAI and InfoBench, MADE outperforms strong baselines by over 50% in software requirement satisfaction (39.9% to 61.9%) and achieves a 95% perfect pass rate on complex instruction following. This work validates a fundamental paradigm shift: moving from optimizing "computable metrics" to "describable qualities," thereby unlocking evolutionary optimization for the vast open-ended domains where no ground truth exists.

进化计算大模型评估主观优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。