让虚拟助手读懂表情和语气,实现有温度的对话交互。
AIVA: An AI-based Virtual Companion for Emotion-aware Interaction
- 用跨模态融合网络分析语音表情等非语言信号。
- 通过情感提示生成共情回复,提升交互自然度。
- 适合心理陪伴、社交机器人等情感交互场景。
大型语言模型(LLMs)虽提升了人机对话能力,但仅处理文本,无法理解表情、语调等非语言情绪信号,限制了沉浸式共情交互。本文提出 ours,一个基于AI的虚拟伴侣,通过多模态情感感知网络(MSPN)融合语音、表情等信号,利用交叉模态融合变压器与监督对比学习提取情感线索。同时设计情感感知提示工程生成共情回应,并集成语音合成(TTS)与动态虚拟形象模块,实现具象化情感表达。该框架支持情感驱动的人机交互,在陪伴机器人、社会照护、心理健康及以人为本的AI领域具有应用前景。
原文摘要 · Abstract (English)
Recent advances in Large Language Models (LLMs) have significantly improved natural language understanding and generation, enhancing Human-Computer Interaction (HCI). However, LLMs are limited to unimodal text processing and lack the ability to interpret emotional cues from non-verbal signals, hindering more immersive and empathetic interactions. This work explores integrating multimodal sentiment perception into LLMs to create emotion-aware agents. We propose \ours, an AI-based virtual companion that captures multimodal sentiment cues, enabling emotionally aligned and animated HCI. \ours introduces a Multimodal Sentiment Perception Network (MSPN) using a cross-modal fusion transformer and supervised contrastive learning to provide emotional cues. Additionally, we develop an emotion-aware prompt engineering strategy for generating empathetic responses and integrate a Text-to-Speech (TTS) system and animated avatar module for expressive interactions. \ours provides a framework for emotion-aware agents with applications in companion robotics, social care, mental health, and human-centered AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。