用区块链验证大模型评测,防止厂商造假和偏见影响结果
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation

- 用区块链+哈希机制让评测过程不可篡改、可追溯
- 发现披露模型身份后得分变化显著,尤其在敏感政治题上
- 适合关注评测公正性的研究者、平台方和监管机构
LLM评测对机构声誉和市场信心至关重要,但当前依赖厂商自报数据,易出现隐瞒模型改动、训练数据污染和选择性报告等问题。为减少人工评审负担,已有研究采用模型自评,但发现评测模型可能因身份偏见而评分不公。本文通过七种验证模型(GPT-OSS 120B、Llama 3.3 70B、GLM 5.1、Qwen3 32B、DeepSeek V4 Pro、Mistral Large3、Sarvam M)对三个主模型在58个事实、推理、政治及偏好类问题上的匿名与身份披露回答进行评估。结果显示:身份披露使事实题得分略有上升,压力推理任务中等影响,地理政治敏感话题则产生大幅变化。其中GLM5.1得分提升7.00分(p=0.0249),Llama 3.3 70B提升1.56分(p=0.00)。本文进一步提出基于以太坊兼容链的承诺-揭示协议:第一阶段,评测者提前提交分数哈希与随机盐;第二阶段,身份与原始分数上链公开验证。该机制建立防篡改审计链,实现盲评与事后声明分离,显著降低独立研究者与排行榜运营者的验证成本。
原文摘要 · Abstract (English)
LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims that DeepSeek R1 outperformed OpenAI's o1 contributed to market panic on January 27, 2025, when Nvidia lost USD589 billion in market value. Yet vendor benchmarks often depend on an honor system. Academic reassessments and independent leaderboards have found undisclosed changes to proprietary models, contaminated training data, and selective reporting. LLM-as-a-judge methods scale evaluation by reducing human review. Studies, however, suggest that judges may show identity-aware bias, scoring an answer according to its source model rather than its quality. This bias has not been fully measured or corrected across politically sensitive, reasoning-intensive, and preference-based tasks. We examine this problem using seven verifier models: GPT-OSS 120B, Llama 3.3 70B, GLM 5.1, Qwen3 32B, DeepSeek V4 Pro, Mistral Large3, and Sarvam M. They score anonymous and identity-disclosed responses from three primary models on 58 factual, reasoning, political, and preference-based questions. Identity disclosure slightly raises scores for factual questions, moderately affects stress-reasoning tasks, and causes large changes for geopolitically sensitive topics. Notable results include GLM5.1 (+7.00 points, p = 0.0249) and Llama 3.3 70B (+1.56 points, p = 0.00). We also introduce a blockchain-based commit-reveal protocol using Autonomous Economic Agents on an Ethereum-compatible ledger. In Phase 1, each judge records a one-way hash of its score and a secret salt before candidate identities are revealed. In Phase 2, the identity and raw score are disclosed and verified on-chain. This creates a tamper-evident audit trail that separates blind evaluation from post-hoc claims and reduces the verification burden on independent researchers and leaderboard operators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。