评测大模型在跨宇宙英雄角色扮演中的一致性与道德判断能力
Beyond One World: Benchmarking Super Heros in Role-Playing Across Multiversal Contexts
- 构建30位英雄90个版本的跨宇宙角色扮演基准
- 发现强模型在思维与行动上难以同时保持准确与一致
- 提出思考-行动匹配度新指标,评估模型可信度
大型语言模型(LLMs)被越来越多地用作角色扮演代理,但其忠实且一致地呈现特定版本角色(如漫画与电影宇宙中的超级英雄)的能力仍待深入探索。漫威与DC等超英正典提供了丰富测试场景:同一角色在数十年叙事中拥有不同背景、价值观与道德准则。为此,我们提出Beyond One World基准,涵盖30位标志性英雄及90个正典特定版本。该基准包含两项任务:(i) 正典事件,测试对关键人生阶段的事实记忆;(ii) 道德困境,考察模型在伦理挑战下的反应。我们采用框架分离内部推理(思考)与外部决策(行动),并提出思考-行动匹配度(Think-Act Matching)指标,作为模型可信度的代理。在推理与非推理模型上的实验揭示三方面发现:(1) 思维链提示可提升弱模型的叙述连贯性,但可能降低强模型的正典准确性;(2) 同一角色跨版本泛化仍是重大障碍;(3) 模型常在思考或行动中表现优异,但很少两者兼备。Beyond One World揭示了多宇宙一致性与推理对齐的关键缺陷,为角色扮演类大模型提供挑战性评估标准。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used as role-playing agents, yet their capacity to faithfully and consistently portray version-specific characters -- for example, superheroes across comic and cinematic universes -- remains underexplored. Superhero canons such as Marvel and DC provide a rich testbed: decades of storytelling yield multiple incarnations of the same character with distinct histories, values, and moral codes. To study this problem, we introduce Beyond One World, a benchmark for character-grounded roleplay spanning 30 iconic heroes and 90 canon-specific versions. The benchmark comprises two tasks: (i) Canon Events, which probes factual recall of pivotal life stages, and (ii) Moral Dilemmas, which confronts models with ethically charged scenarios. We score responses for canonical accuracy and reasoning fidelity under a framework that separates internal deliberation ("thinking") from outward decisions ("acting"). We further propose Think-Act Matching, a metric that quantifies alignment between reasons and actions and serves as a proxy for model trustworthiness. Experiments across reasoning- and non-reasoning-oriented models yield three findings: (1) chain-of-thought prompting improves narrative coherence in weaker models but can reduce canonical accuracy in stronger ones; (2) cross-version generalization within a character remains a major obstacle; and (3) models often excel at either thinking or acting, but rarely both. Beyond One World exposes critical gaps in multiversal consistency and reasoning alignment, offering a challenging evaluation for role-playing LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。