arXiv:2512.04921cs.AIcs.CL2025-12被引 1

首个评估AI完成日常消费任务能力的基准,发现顶级模型仍有明显短板

The AI Consumer Index (ACE)

  • 构建涵盖购物、餐饮、游戏、家装的400项测试集,动态验证回复是否基于真实网页信息
  • 顶尖模型最高得分仅56.1%,购物类任务不足50%,普遍存在价格等关键信息幻觉
  • 适合关注AI实用性、评测者及希望提升模型可信度的研究人员

我们推出首个人工智能消费者指数(ACE),用于评估前沿AI模型执行日常消费任务的能力。ACE包含400个隐藏测试用例,分为购物、餐饮、游戏和家装四类。80个案例已以CC-BY许可开源作为开发集。在排行榜中,我们采用新型评分方法,动态检查响应中相关部分是否基于检索到的网络来源。GPT 5(Thinking = High)表现最佳,得分为56.1%,其次为o3 Pro(Thinking = On)的55.2%和GPT 5.1(Thinking = High)的55.1%。各模型在不同领域表现差异明显,购物类任务中最高得分低于50%。研究发现,模型容易在价格等关键信息上产生幻觉。ACE揭示了当前最先进模型与实际消费者需求之间的显著差距。

原文摘要 · Abstract (English)

We introduce the first version of the AI Consumer Index (ACE), a benchmark for assessing whether frontier AI models can perform everyday consumer tasks. ACE contains a hidden heldout set of 400 test cases, split across four consumer activities: shopping, food, gaming, and DIY. We are also open sourcing 80 cases as a devset with a CC-BY license. For the ACE leaderboard we evaluated 10 frontier models (with websearch turned on) using a novel grading methodology that dynamically checks whether relevant parts of the response are grounded in the retrieved web sources. GPT 5 (Thinking = High) is the top-performing model, scoring 56.1%, followed by o3 Pro (Thinking = On) at 55.2% and GPT 5.1 (Thinking = High) at 55.1%. Model scores differ across domains, and in Shopping the top model scores under 50\%. We find that models are prone to hallucinating key information, such as prices. ACE shows a substantial gap between the performance of even the best models and consumers' AI needs.

AI评测消费任务幻觉检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。