用真实应用中的用户反馈实时评估大模型性能,更贴近实际使用。
Inclusion Arena: An Open Platform for Evaluating Large Foundation Models with Real-World Apps
- 将模型对比嵌入真实应用交互,直接收集用户反馈进行评估。
- 采用改进的布拉德利-特里模型,支持新模型快速评级与高效比较。
- 适合关注模型真实表现、追求实用部署的研究者和开发者。
大型语言模型(LLMs)和多模态大语言模型(MLLMs)已展现出接近人类水平的综合能力。尽管已有诸多基准测试(如MMLU)和排行榜(如Chatbot Arena),但多数依赖静态数据集或众包通用提示,难以反映模型在真实应用场景中的表现。为此,我们推出Inclusion Arena——一个基于真实应用中用户反馈的实时排行榜。平台将成对模型比较融入自然用户交互,确保评估结果贴近实际使用场景。为实现稳健排名,采用增强版布拉德利-特里模型,引入两项关键创新:(1) 放置匹配(Placement Matches),用于新模型快速获得初始评分;(2) 近邻采样(Proximity Sampling),优先比较能力相近的模型以最大化信息收益并提升评分稳定性。大量实证分析与模拟表明,Inclusion Arena能生成可靠稳定的排名,数据传递性高于普通众包数据集,且显著降低恶意操纵风险。通过构建基础模型与真实应用间的开放协作生态,该平台旨在加速面向实际、以用户为中心的模型优化进程。平台已公开访问:https://www.tbox.cn/about/model-ranking。
原文摘要 · Abstract (English)
Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) have ushered in a new era of AI capabilities, demonstrating near-human-level performance across diverse scenarios. While numerous benchmarks (e.g., MMLU) and leaderboards (e.g., Chatbot Arena) have been proposed to help evolve the development of LLMs and MLLMs, most rely on static datasets or crowdsourced general-domain prompts, often falling short of reflecting performance in real-world applications. To bridge this critical gap, we present Inclusion Arena, a live leaderboard that ranks models based on human feedback collected directly from AI-powered applications. Our platform integrates pairwise model comparisons into natural user interactions, ensuring evaluations reflect practical usage scenarios. For robust model ranking, we employ the Bradley-Terry model augmented with two key innovations: (1) Placement Matches, a cold-start mechanism to quickly estimate initial ratings for newly integrated models, and (2) Proximity Sampling, an intelligent comparison strategy that prioritizes battles between models of similar capabilities to maximize information gain and enhance rating stability. Extensive empirical analyses and simulations demonstrate that Inclusion Arena yields reliable and stable rankings, exhibits higher data transitivity compared to general crowdsourced datasets, and significantly mitigates the risk of malicious manipulation. By fostering an open alliance between foundation models and real-world applications, Inclusion Arena aims to accelerate the development of LLMs and MLLMs truly optimized for practical, user-centric deployments. The platform is publicly accessible at https://www.tbox.cn/about/model-ranking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。