arXiv:2603.20225cs.CYcs.AI2026-03

专家角色能提升大模型表现吗?实验证明:关键在评测设计,而非角色本身。

The Arrival of AGI? When Expert Personas Exceed Expert Benchmarks

  • 用最难题集强制真实推理,排除模式匹配干扰。
  • 专家角色使准确率达95%以上,消除基线错误。
  • 现有评测体系无法捕捉角色优势,需重建评估标准。

专家角色能否提升语言模型性能?沃顿生成式人工智能实验室报告称其无效,并通过社交媒体向数百万用户传播建议,呼吁从业者放弃Anthropic、Google和OpenAI推荐的技术。我们证明这一零结果在结构上是可预测的。五种核心机制在数据收集前就阻碍了检测:基线污染使起点接近天花板,系统提示层级压制实验干预,不可能的专家设定退化为通用能力,格式限制抑制推理过程,提供方排除削弱泛化性。通过控制实验修正这些缺陷,揭示了原设计所掩盖的真相。我们选取GPQA钻石级最难问题,防止基线模式匹配,迫使依赖真正的专家推理。在存在有效关键答案的题目中,专家角色实现接近天花板的准确率(95%以上),通过信心放大消除了所有基线错误。对模型差异的法医分析发现,一半最难的GPQA题目包含化学或逻辑上不可行的答案,模型的思维链(CoT)显示其推理远离这些不可能答案,导致准确化学知识被惩罚。这些发现重新诠释了原始的零结果。方法严谨的角色研究面临基准有效性限制带来的测量困境。回答角色问题需要当前领域尚未具备的评估基础设施。

原文摘要 · Abstract (English)

Do expert personas improve language model performance? The Wharton Generative AI Lab reports that they do not, broadcasting to millions via social media the recommendation that practitioners abandon a technique recommended by Anthropic, Google, and OpenAI. We demonstrate that this null finding was structurally predictable. Five core mechanisms precluded detection before data collection began: baseline contamination elevating the starting point to near-ceiling, system prompt hierarchy subordinating experimental manipulation, impossible expert specifications collapsing to generic competence, format constraints suppressing reasoning processes, and provider exclusion limiting generalizability. Controlled trials correcting these limitations reveal what the original design obscured. To test this, we selected the GPQA Diamond hardest questions to prevent baseline pattern matching, forcing reliance on genuine expert reasoning. On items with valid key answers, expert personas achieve ceiling accuracy. They eliminated all baseline errors through confidence amplification. Furthermore, forensic examination of model divergence identified that half of the hardest GPQA items contain chemically or logically indefensible answers. The model's CoT revealed reasoning away from impossible answers, yielding penalization for accurate chemistry. These findings recontextualize the original null results. Methodologically sound persona research faces measurement constraints imposed by benchmark validity limitations. Answering the persona question requires evaluation infrastructure the field does not yet possess.

大模型评估专家角色推理能力评测设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。