让AI通过表情识别理解用户情绪,提升对话自然度
Empathic Prompting: Non-Verbal Context Integration for Multimodal LLM Conversations
- 用面部识别捕捉用户情绪,自动融入对话提示中
- 5人测试显示情绪信息有效融入回复,对话更流畅
- 适合医疗教育等需关注情绪的交互场景
我们提出情感提示(Empathic Prompting)框架,通过集成商用面部表情识别服务,在多模态人机交互中引入隐含的非语言情绪线索,并将其作为上下文信号嵌入提示过程。与传统多模态界面不同,该方法无需用户主动操作,能无感增强文本输入的情感信息,实现对话连贯性与自然度的对齐。系统架构模块化且可扩展,支持接入更多非语言模态。基于本地部署的DeepSeek模型实现,初步服务与可用性评估(N=5)表明,非语言输入被一致整合到连贯的LLM输出中,参与者普遍认可对话流畅性。本工作为聊天机器人在医疗、教育等情绪敏感领域提供新思路。
原文摘要 · Abstract (English)
We present Empathic Prompting, a novel framework for multimodal human-AI interaction that enriches Large Language Model (LLM) conversations with implicit non-verbal context. The system integrates a commercial facial expression recognition service to capture users' emotional cues and embeds them as contextual signals during prompting. Unlike traditional multimodal interfaces, empathic prompting requires no explicit user control; instead, it unobtrusively augments textual input with affective information for conversational and smoothness alignment. The architecture is modular and scalable, allowing integration of additional non-verbal modules. We describe the system design, implemented through a locally deployed DeepSeek instance, and report a preliminary service and usability evaluation (N=5). Results show consistent integration of non-verbal input into coherent LLM outputs, with participants highlighting conversational fluidity. Beyond this proof of concept, empathic prompting points to applications in chatbot-mediated communication, particularly in domains like healthcare or education, where users' emotional signals are critical yet often opaque in verbal exchanges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。