arXiv:2510.05909cs.AI2025-10被引 1

用说服力优化让大模型推理更通用,效果优于追求真理的常规方法。

Optimizing for Persuasion Improves LLM Generalization: Evidence from Quality-Diversity Evolution of Debate Strategies

  • 通过比赛式辩论演化多样说服策略,单模型即可实现群体多样性优势。
  • 在多个模型规模下,说服优化使训练与测试差距缩小13.94%。
  • 适合关注模型泛化能力、推理鲁棒性的研究者与开发者。

大型语言模型(LLMs)若以输出真相为目标,常出现过拟合,导致推理脆弱且难以泛化。尽管基于说服的优化在辩论场景中展现出潜力,却未系统对比主流的真理导向方法。本文提出DebateQD——一种极简的质量-多样性(QD)进化算法,通过锦标赛式辩论,让两个LLM对决,第三方裁判评分,从而演化出理性、权威、情感诉求等多类辩论策略。不同于需多模型群体的方法,本方法仅用单一模型架构,通过提示词设计维持对手多样性,兼顾可实验性与群体优化优势。我们通过固定辩论协议、仅更换评估目标,明确分离优化目标的影响:说服型奖励鼓励说动裁判,不问真假;真理型奖励强调协作正确性。在包含7B、32B、72B参数量模型及多个数据集大小的QuALITY基准测试中,说服优化策略相较真理优化,训练-测试泛化差距最大缩小13.94%,同时保持或超越其测试表现。这是首个受控证据表明,竞争性说服压力比合作求真更能促进可迁移的推理能力,为提升大模型泛化提供新路径。

原文摘要 · Abstract (English)

Large Language Models (LLMs) optimized to output truthful answers often overfit, producing brittle reasoning that fails to generalize. While persuasion-based optimization has shown promise in debate settings, it has not been systematically compared against mainstream truth-based approaches. We introduce DebateQD, a minimal Quality-Diversity (QD) evolutionary algorithm that evolves diverse debate strategies across different categories (rationality, authority, emotional appeal, etc.) through tournament-style competitions where two LLMs debate while a third judges. Unlike previously proposed methods that require a population of LLMs, our approach maintains diversity of opponents through prompt-based strategies within a single LLM architecture, making it more accessible for experiments while preserving the key benefits of population-based optimization. In contrast to prior work, we explicitly isolate the role of the optimization objective by fixing the debate protocol and swapping only the fitness function: persuasion rewards strategies that convince the judge irrespective of truth, whereas truth rewards collaborative correctness. Across three model scales (7B, 32B, 72B parameters) and multiple dataset sizes from the QuALITY benchmark, persuasion-optimized strategies achieve up to 13.94% smaller train-test generalization gaps, while matching or exceeding truth optimization's test performance. These results provide the first controlled evidence that competitive pressure to persuade, rather than seek the truth collaboratively, fosters more transferable reasoning skills, offering a promising path for improving LLM generalization.

大模型泛化能力辩论优化QD算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。