给AI设定专家人设,无法提升其回答难题的准确性。
Prompting Science Report 4: Playing Pretend: Expert Personas Don't Improve Factual Accuracy
- 测试不同人设对模型答题的影响,包括领域内/外专家和低知识人设。
- 在两个高端测评集上,人设基本未提升准确率,部分反而降低。
- 适合关注AI输出质量而非语气风格的研究者和应用者。
本报告是系列短篇技术报告中的第四篇,旨在帮助商业、教育与政策决策者理解与AI协作的技术细节。研究探讨为模型分配人设是否能提升其在高难度客观选择题上的表现。我们测试了领域内专家人设(如物理问题配物理专家)、跨领域专家人设(如法律问题配物理专家)以及低知识人设(如普通人、儿童、幼儿),评估六种模型在GPQA Diamond(Rein等,2024)和MMLU-Pro(Wang等,2024)两个基准上的表现。结果表明:领域内专家人设对性能无显著影响(仅Gemini 2.0 Flash例外);跨领域专家人设影响微弱;低知识人设普遍损害准确率。整体而言,人设提示未优于无提示基线,且专家人设未展现一致优势,跨领域人设有时还导致性能下降。该结论仅针对事实准确性,人设可能在改变输出语气等方面仍有价值。
原文摘要 · Abstract (English)
This is the fourth in a series of short reports that help business, education, and policy leaders understand the technical details of working with AI through rigorous testing. Here, we ask whether assigning personas to models improves performance on difficult objective multiple-choice questions. We study both domain-specific expert personas and low-knowledge personas, evaluating six models on GPQA Diamond (Rein et al. 2024) and MMLU-Pro (Wang et al. 2024), graduate-level questions spanning science, engineering, and law. We tested three approaches: -In-Domain Experts: Assigning the model an expert persona ("you are a physics expert") matched to the problem type (physics problems) had no significant impact on performance (with the exception of the Gemini 2.0 Flash model). -Off-Domain Experts (Domain-Mismatched): Assigning the model an expert persona ("you are a physics expert") not matched to the problem type (law problems) resulted in marginal differences. -Low-Knowledge Personas: We assigned the model negative capability personas (layperson, young child, toddler), which were generally harmful to benchmark accuracy. Across both benchmarks, persona prompts generally did not improve accuracy relative to a no-persona baseline. Expert personas showed no consistent benefit across models, with few exceptions. Domain-mismatched expert personas sometimes degraded performance. Low-knowledge personas often reduced accuracy. These results are about the accuracy of answers only; personas may serve other purposes (such as altering the tone of outputs), beyond improving factual performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。