提出新评估框架,精准区分角色与情绪一致性表现
MERRY: Semantically Decoupled Evaluation of Multimodal Emotional and Role Consistencies of Role-Playing Agents
- 分离情绪与角色评估维度,避免评价混淆
- 实测显示真实数据训练提升情绪一致性,合成数据则降低
- 适合评估多模态角色扮演模型的开发者与研究者
多模态角色扮演智能体(MRPAs)因其能提供更沉浸式的多模态情感交互而受到关注。然而,现有研究仍依赖纯文本基准评估文本响应,将多模态表达评估完全交给生成指标,导致语义评估与模态生成纠缠不清,误差归因模糊,且过度依赖人工判断。为此,我们提出MERRY——一种语义解耦的多模态情感与角色一致性评估框架,包含5个情绪一致性(EC)和3个角色一致性(RC)细化指标。关键创新在于将传统主观评分转化为双向证据查找任务,显著提升大模型作为评判者的共识度。基于MERRY的实证分析揭示:(1)在合成数据上训练会降低情绪一致性,而真实数据训练可提升;(2)现有模型普遍存在情绪模板化与简化问题,在细粒度负面情绪上存在正向偏差和性能瓶颈;(3)简单提示方法虽增强弱模型,却限制强模型表现,而简单微调则导致角色泛化能力差。代码与数据集已公开。
原文摘要 · Abstract (English)
Multimodal Role-Playing Agents (MRPAs) are attracting increasing attention due to their ability to deliver more immersive multimodal emotional interactions. However, existing studies still rely on pure textual benchmarks to evaluate the text responses of MRPAs, while delegating the assessment of their multimodal expressions solely to modality-synthesis metrics. This evaluation paradigm, on the one hand, entangles semantic assessment with modality generation, leading to ambiguous error attribution, and on the other hand remains constrained by the heavy reliance on human judgment. To this end, we propose MERRY, a semantically decoupled evaluation framework for assessing Multimodal Emotional and Role consistencies of Role-playing agents. This framework introduce five refined metrics for EC and three for RC. Notably, we transform the traditional subjective scoring approach into a novel bidirectional-evidence-finding task, significantly improving the human agreement of LLM-as-Judge evaluations. Based on MERRY, we conduct extensive evaluations. Our empirical results primarily reveal that: (1) Training on synthetic datasets tends to reduce emotional consistency, whereas training on real-world datasets improves it; (2) Existing models suffer from emotional templatization and simplification, exhibiting positive-bias and performance bottleneck in fine-grained negative emotions; (3) Simple prompting method strengthens the weak models but constrains the strong ones, while simple fine-tuning method suffers from poor role generalization. Codes and dataset are available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。