用区块链构建去中心化评估框架,解决大模型评测的不一致问题
InfiCoEvalChain: A Blockchain-Based Decentralized Framework for Collaborative LLM Evaluation
- 基于区块链让全球用户作为独立验证者参与评测
- 同一模型十次运行的标准差从1.67降至0.28
- 适合关注评测可信度和公平性的研究者与开发者
大语言模型的快速演进对评测可靠性提出更高要求,但现有集中式评测存在透明度低、过拟合和硬件差异导致的波动问题。实证分析显示:在HumanEval上对单一模型重复十次测试的标准差(1.67)甚至超过前10名模型间的性能差距(0.91),使当前排名统计上极不可靠。为此,我们提出一种基于区块链的去中心化评估框架,通过异构计算节点大规模基准测试实现硬件与参数多样性。该框架利用区块链协议激励全球贡献者作为独立验证者,通过稳健的奖励机制保障评测完整性,防止恶意参与。多方共识与多样化推理环境共同形成“去中心化背书”,将评测从“集中式黑箱”转变为更稳定、更具代表性的度量方式。实验表明,该框架将同一模型十次运行的标准差降至0.28,显著优于传统方法,提升模型排名的统计置信度。平台已完整实现,即将开源。
原文摘要 · Abstract (English)
The rapid advancement of large language models (LLMs) demands increasingly reliable evaluation, yet current centralized evaluation suffers from opacity, overfitting, and hardware-induced variance. Our empirical analysis reveals an alarming inconsistency in existing evaluations: the standard deviation across ten repeated runs of a single model on HumanEval (1.67) actually exceeds the performance gap among the top-10 models on the official leaderboard (0.91), rendering current rankings statistically precarious. To mitigate these instabilities, we propose a decentralized evaluation framework that enables hardware and parameter diversity through large-scale benchmarking across heterogeneous compute nodes. By leveraging the blockchain-based protocol, the framework incentivizes global contributors to act as independent validators, using a robust reward system to ensure evaluation integrity and discourage dishonest participation. This collective verification transforms evaluation from a "centralized black box" into a "decentralized endorsement" where multi-party consensus and diverse inference environments yield a more stable, representative metric. Experimental results demonstrate that the decentralized evaluation framework reduces the standard deviation across ten runs on the same model to 0.28. This significant improvement over conventional frameworks ensures higher statistical confidence in model rankings. We have completely implemented this platform and will soon release it to the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。