arXiv:2505.17223cs.CV2025-05被引 14

生成多样且同步的自然人脸反应,提升人机交互真实感

REACT 2025: the Third Multiple Appropriate Facial Reaction Generation Challenge

  • 基于多模态数据构建动态反应生成模型
  • 首次发布含2856个会话的MAFRG数据集(MARS)
  • 适合研究对话系统与情感计算的开发者

在双人互动中,对同一说话行为可能有多种合适的人脸反应。继成功举办REACT 2023和REACT 2024挑战赛后,我们推出REACT 2025挑战,旨在推动机器学习模型在生成多种合适、多样、真实且同步的人类听众面部反应方面的发展与评估。该反应基于输入刺激(即对应说话者表达的音视频行为)。作为挑战核心,我们提供了首个自然、大规模的多模态多反应生成数据集MARS,记录了137组人-人双人互动,共2856个互动会话,涵盖五个不同主题。本文还介绍了挑战规则,并报告了基准模型在两个子任务(离线与在线反应生成)上的性能表现。基准代码已公开于https://github.com/reactmultimodalchallenge/baseline_react2025。

原文摘要 · Abstract (English)

In dyadic interactions, a broad spectrum of human facial reactions might be appropriate for responding to each human speaker behaviour. Following the successful organisation of the REACT 2023 and REACT 2024 challenges, we are proposing the REACT 2025 challenge encouraging the development and benchmarking of Machine Learning (ML) models that can be used to generate multiple appropriate, diverse, realistic and synchronised human-style facial reactions expressed by human listeners in response to an input stimulus (i.e., audio-visual behaviours expressed by their corresponding speakers). As a key of the challenge, we provide challenge participants with the first natural and large-scale multi-modal MAFRG dataset (called MARS) recording 137 human-human dyadic interactions containing a total of 2856 interaction sessions covering five different topics. In addition, this paper also presents the challenge guidelines and the performance of our baselines on the two proposed sub-challenges: Offline MAFRG and Online MAFRG, respectively. The challenge baseline code is publicly available at https://github.com/reactmultimodalchallenge/baseline_react2025

人脸生成多模态对话系统数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。