arXiv:2605.08518cs.AI2026-05被引 2

揭秘工业多智能体竞赛的评分真相:隐藏评测与实际表现大不同

Results and Retrospective Analysis of the CODS 2025 AssetOpsBench Challenge

论文配图:Results and Retrospective Analysis of the CODS 2025 AssetOpsBench Challenge
图 1 · 摘自论文原文
  • 通过300份提交日志和149个注册团队数据,还原竞赛真实表现
  • 公开排行榜最高仅72.73%,隐藏测试中部分低分系统达63.64%
  • 关键评价项权重极低,成功靠防御机制而非新架构

我们回顾了CODS 2025 extit{AssetOpsBench} 挑战赛,该赛事基于 extit{AssetOps} 构建,聚焦工业级多智能体编排的隐私感知问题。结合最终排名、300次提交日志、149支注册队伍、最佳提交记录、组织者获奖报告、配套系统论文及经验证的规划赛道源码树,揭示五大发现:其一,公开规划榜饱和于72.73%,更优提示无法突破;其二,隐藏评测中规划得分相关性为0.69,执行得分相关性为-0.13,多个公开执行得分45.45%的系统在私有集达63.64%;其三, extit{tmatch} 项在官方综合评分中贡献不足0.05分,重新缩放将调换前两名;其四,虽以账户为单位运营,实则149支队伍仅24支获公开分,11支完整排名,52.3%的去重注册含多个用户名;其五,成功策略集中于增强护栏机制,如响应选择、污染清理、降级处理与上下文控制,而非新型智能体结构。这些结果揭示评估机制的激励方向,推动构建更敏感的复合指标、能力分级诊断与版本化成果发布。

原文摘要 · Abstract (English)

Competition retrospectives are useful when they explain what a leaderboard measured, how hidden evaluation changed conclusions, and which design patterns were rewarded. We revisit the CODS 2025 \assetopslive{} challenge, a privacy-aware Codabench competition on industrial multi-agent orchestration built on \assetops{}. We combine final rank sheets, a 300-submission server log, 149-team registrations, best-submission exports, the organizer winners report, the companion \assetopslive{} system paper, and verified planning-track source trees. Five results stand out. First, the public planning leaderboard saturates at 72.73\%, and richer prompts do not improve that peak. Second, hidden evaluation changes the story: public and private scores correlate moderately in planning ($r{=}0.69$) but negatively in execution ($r{=}{-}0.13$), with several 45.45\% public execution systems reaching 63.64\% on the hidden set. Third, the \tmatch{} term is numerically almost inert in the official composite -- combined on a 0--1 scale with 0--100 percentage scores, it contributes at most 0.05 points per track, and rescaling would swap the top two teams. Fourth, the competition is operationally account-based but substantively team-based: 149 registered teams reduce to 24 with non-zero public scores and 11 fully ranked, while 52.3\% of deduplicated registrations list multiple usernames. Fifth, successful execution methods mostly improve guardrails -- response selection, contamination cleanup, fallback, and context control -- rather than novel agent architectures. These findings identify which behaviors the evaluation rewarded, and motivate scale-aware composites, skill-level diagnostics, and versioned artifact release.

多智能体评测分析竞赛复盘隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。