arXiv:2510.06800cs.CLcs.AI2025-10被引 2

构建可任意定制的角色扮演评测基准,支持多智能体协作自动生成测试场景。

FURINA: A Fully Customizable Role-Playing Benchmark via Scalable Multi-Agent Collaboration Pipeline

  • 通过多智能体协作流水线自动生成可扩展的角色扮演测试集。
  • 实测显示推理能力越强的模型越易产生幻觉,存在性能与可靠性权衡。
  • 适合研究角色扮演、评估大模型对话质量或设计自适应评测体系的团队。

随着大语言模型在角色扮演任务上的发展,现有评测基准因范围狭窄、交互范式过时及适应性差而迅速过时。为此,我们提出FURINA-Builder,一种新型多智能体协作流水线,可全自动构建任意规模的可定制角色扮演评测基准。该系统模拟测试角色与其他角色之间的对话,从精心构建的角色-场景池中选取参与者,并由大模型裁判选择细粒度评估维度,将测试角色回复调整为最终测试语句。基于此流程,我们构建了FURINA-Bench,一个涵盖真实与合成角色的综合性角色扮演评测基准,每项角色均采用维度特定评估标准。人工评估与初步可区分性分析验证了设计的有效性。对前沿大模型的广泛评估发现,o3和DeepSeek-R1分别在英文与中文角色扮演任务中表现最佳。所有模型中,真实角色持续优于合成角色,且推理能力进一步放大这一差距。有趣的是,模型规模并未单调减少幻觉。更关键的是,对于推理型模型,我们发现新现象:推理虽提升角色扮演表现,但同时增加幻觉风险。这一权衡扩展为所有模型的更广帕累托前沿,揭示了角色扮演性能与可靠性的根本挑战。这些发现证明了FURINA-Builder的有效性以及FURINA-Bench带来的评测难度。

原文摘要 · Abstract (English)

As large language models (LLMs) advance in role-playing (RP) tasks, existing benchmarks quickly become obsolete due to their narrow scope, outdated interaction paradigms, and limited adaptability across diverse application scenarios. To address this gap, we introduce FURINA-Builder, a novel multi-agent collaboration pipeline that automatically constructs fully customizable RP benchmarks at any scale. It enables evaluation of arbitrary characters across diverse scenarios and prompt formats, as the first benchmark builder in RP area for adaptable assessment. FURINA-Builder simulates dialogues between a test character and other characters drawn from a well-constructed character-scene pool, while an LLM judge selects fine-grained evaluation dimensions and adjusts the test character's responses into final test utterances. Using this pipeline, we build FURINA-Bench, a new comprehensive role-playing benchmark featuring both established and synthesized test characters, each assessed with dimension-specific evaluation criteria. Human evaluation and preliminary separability analysis justify our pipeline and benchmark design. We conduct extensive evaluations of cutting-edge LLMs and find that o3 and DeepSeek-R1 achieve the best performance on English and Chinese RP tasks, respectively. Across all models, established characters consistently outperform synthesized ones, with reasoning capabilities further amplifying this disparity. Interestingly, we observe that model scale does not monotonically reduce hallucinations. More critically, for reasoning LLMs, we uncover a novel trade-off: reasoning improves RP performance but simultaneously increases RP hallucinations. This trade-off extends to a broader Pareto frontier between RP performance and reliability for all LLMs. These findings demonstrate the effectiveness of FURINA-Builder and the challenge posed by FURINA-Bench.

角色扮演评测基准多智能体大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。