用多智能体竞赛机制优化大模型回答,提升准确性和可靠性。
ART: Adaptive Response Tuning Framework -- A Multi-Agent Tournament-Based Approach to LLM Response Optimization
- 多模型竞逐+评审核心,通过锦标赛排名筛选最优答案
- 综合表现优于单模型,整体质量提升8.4%,评分收敛R²超0.96
- 适合需要高可信度输出的生产级应用,支持灵活配置
大语言模型在自然语言理解与生成方面展现出强大能力,但单一模型输出常存在不一致、幻觉及领域差异问题。本文提出ART(自适应响应调优框架),采用锦标赛式ELO评分与多智能体推理,使多个大模型智能体在结构化赛程中竞争、批判与协作,生成共识性回答,显著优于单模型输出。框架引入可配置赛制参数、动态智能体选择及多种共识融合策略。实验表明,相比基线方法,响应准确率、连贯性与可靠性均显著提升。ART为高要求应用场景提供可扩展、可落地的解决方案,在整体质量指标上提升8.4%,且ELO评分收敛的R²值超过0.96。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable capabilities in natural language understanding and generation. However, single-model responses often exhibit inconsistencies, hallucinations, and varying quality across different query domains. This paper presents ART (Adaptive Response Tuning), a novel framework that employs tournament-style ELO ranking and multi-agent reasoning to systematically optimize LLM outputs. By enabling multiple LLM agents to compete, critique, and collaborate through structured tournament workflows, ART produces consensus responses that outperform individual model outputs. Our framework introduces configurable tournament parameters, dynamic agent selection, and multiple consensus fusion strategies. Experimental evaluations demonstrate significant improvements in response accuracy, coherence, and reliability compared to baseline single-model approaches. The ART framework provides a scalable, production-ready solution for applications requiring high-quality, vetted LLM responses, achieving an 8.4% improvement in overall quality metrics and R^2 values exceeding 0.96 in ELO rating convergence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。