arXiv:2510.12692cs.HCcs.AI2025-10AAAI

AI可媲美专家匹配评委,效率提升数倍。

Who is a Better Matchmaker? Human vs. Algorithmic Judge Assignment in a High-Stakes Startup Competition

  • 用混合语义相似度算法自动匹配参赛项目与评委。
  • AI匹配质量与人工相当(平均分3.90 vs 3.94,p=0.40)。
  • 适合需要高效高质评审的大型创业竞赛场景。

人工智能在复杂决策中的应用日益广泛,但其在需语义理解与专业领域知识情境下的表现仍不明确。本文以哈佛大学校长创新挑战赛为例,该赛事为学生及校友初创企业提供超50万美元奖金,评委匹配质量至关重要。我们开发了基于混合词法-语义相似度的集成算法HLSE,并在比赛中部署。通过盲评对309组评委-项目配对的质量进行评估,使用曼-惠特尼U检验发现,算法与人工匹配无显著差异(AUC=0.48,p=0.40),平均评分分别为3.90和3.94(满分5分)。此外,原本需一周的人工匹配工作,算法可在数小时内完成。结果表明,HLSE实现了专家级匹配质量,同时具备更强的可扩展性与效率,凸显了AI在高风险评审任务中辅助甚至替代人工决策的巨大潜力。

原文摘要 · Abstract (English)

There is growing interest in applying artificial intelligence (AI) to automate and support complex decision-making tasks. However, it remains unclear how algorithms compare to human judgment in contexts requiring semantic understanding and domain expertise. We examine this in the context of the judge assignment problem, matching submissions to suitably qualified judges. Specifically, we tackled this problem at the Harvard President's Innovation Challenge, the university's premier venture competition awarding over \$500,000 to student and alumni startups. This represents a real-world environment where high-quality judge assignment is essential. We developed an AI-based judge-assignment algorithm, Hybrid Lexical-Semantic Similarity Ensemble (HLSE), and deployed it at the competition. We then evaluated its performance against human expert assignments using blinded match-quality scores from judges on $309$ judge-venture pairs. Using a Mann-Whitney U statistic based test, we found no statistically significant difference in assignment quality between the two approaches ($AUC=0.48, p=0.40$); on average, algorithmic matches are rated $3.90$ and manual matches $3.94$ on a 5-point scale, where 5 indicates an excellent match. Furthermore, manual assignments that previously required a full week could be automated in several hours by the algorithm during deployment. These results demonstrate that HLSE achieves human-expert-level matching quality while offering greater scalability and efficiency, underscoring the potential of AI-driven solutions to support and enhance human decision-making for judge assignment in high-stakes settings.

智能匹配创业竞赛算法决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。