多智能体辩论能提升数据清洗准确率,但可能引发混淆导致效果下降。
When Helping Hurts and How to Fix It: Multi-Agent Debate for Data Cleaning
- 通过辩论机制让不同角色协作纠错,但需避免反馈误导。
- 辩论使生成质量下降1.6至15.5个百分点,但错误检测提升27.4个百分点。
- 使用独立验证者与代码执行能力可显著提升生成效果,适合高精度任务。
在三个基准、四种模型及超过6000个任务-条件组合中,我们发现多智能体辩论的效果会反转:它在所有四类模型上均降低生成质量(-1.6至-15.5个百分点),原因在于批判引发的困惑(CIC)和生成器对虚构批评的盲目接受;然而,其错误检测能力却提升27.4个百分点(F1值),效应量d=1.0。我们推导出辩论有效性的条件:当修正错误输出的概率(基于可修复性加权的验证机会)高于破坏正确输出的概率时,辩论才有效。因子实验表明,对抗性分离至关重要:使用相同工具的自我验证无效,而配备代码执行能力与证据门控生成的独立批评者首次在生成任务上显著超越单智能体(+5.3个百分点,p<0.05)。该条件可准确预测全部九种任务类型,并在七个领域的19项已发表比较中实现零误报泛化。
原文摘要 · Abstract (English)
When does multi-agent debate help data cleaning, and when does it hurt? Across three benchmarks, four model families, and over 6,000 task-condition pairs, we find debate's effect reverses sign: it degrades generation across all four models (-1.6 to -15.5pp) through critique-induced confusion (CIC), hallucinated Critic feedback that the Generator accepts uncritically, yet improves error detection (+27.4pp F1, d=1.0). We derive a debate benefit condition: debate helps when the probability of rescuing a wrong output (Critic verification odds weighted by fixability) exceeds the probability of destroying a correct one. A factorial experiment proves adversarial separation is essential: self-verification with identical tools fails, while a separate Critic with code-execution grounding and evidence-gated generation produces the first debate configuration to significantly exceed single-agent on a generative task (+5.3pp, p<0.05). The condition correctly predicts all nine task types and generalizes with zero false positives across 19 published comparisons in seven domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。