arXiv:2504.06460cs.CL2025-04被引 2

测试大模型能否模拟表现相反的角色,发现其能力严重不足。

Can LLMs Simulate Personas with Reversed Performance? A Systematic Investigation for Counterfactual Instruction Following in Math Reasoning Context

  • 构建首个评估反事实指令遵循能力的数学推理基准数据集
  • 主流大模型在模拟低水平角色时表现普遍差,连o1也难达标
  • 同时模拟能力与种族属性会进一步恶化表现,适合教育仿真研究者

大型语言模型(LLMs)现被广泛用于虚拟环境中模拟人物角色,依赖其指令遵循能力。然而我们发现,即使是先进的大模型也无法有效模拟表现相反的人物(如教育场景中低能力的学生角色),这限制了仿真环境的多样性与实用性。本文以数学推理为典型场景,提出首个评估大模型模拟反向表现角色能力的基准数据集,该能力被称为“反事实指令遵循”。我们在该任务上评估了开源与闭源大模型,结果表明,包括OpenAI o1推理模型在内的所有模型均难以正确遵循反事实指令来模拟低水平角色。当同时需模拟角色的能力水平与种族属性时,表现进一步恶化。这些结果揭示了反事实指令遵循的挑战,凸显了未来研究的必要性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are now increasingly widely used to simulate personas in virtual environments, leveraging their instruction-following capability. However, we discovered that even state-of-the-art LLMs cannot simulate personas with reversed performance (e.g., student personas with low proficiency in educational settings), which impairs the simulation diversity and limits the practical applications of the simulated environments. In this work, using mathematical reasoning as a representative scenario, we propose the first benchmark dataset for evaluating LLMs on simulating personas with reversed performance, a capability that we dub "counterfactual instruction following". We evaluate both open-weight and closed-source LLMs on this task and find that LLMs, including the OpenAI o1 reasoning model, all struggle to follow counterfactual instructions for simulating reversedly performing personas. Intersectionally simulating both the performance level and the race population of a persona worsens the effect even further. These results highlight the challenges of counterfactual instruction following and the need for further research.

大模型角色模拟反事实推理数学推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。