arXiv:2604.15336cs.HCcs.AI2026-04

用面部表情提升AI导师共情力,不需重训练就能见效。

Facial-Expression-Aware Prompting for Empathetic LLM Tutoring

论文配图:Facial-Expression-Aware Prompting for Empathetic LLM Tutoring
图 1 · 摘自论文原文
  • 通过提取面部动作单元(AU)生成文本描述或选峰值帧,让模型理解表情。
  • 在960次对话中,基于AU的方法显著提升共情响应能力。
  • 适合做智能教育、情感计算方向的研究者和开发者参考。

大型语言模型(LLMs)正推动更具能力的对话式教学代理发展,但有效教学需要感知学习者的认知与情感状态,而不仅限于文本。面部表情能即时反映困惑、挫败或投入状态,但在基于LLM的教学系统中仍被忽视。本文研究了在提示层集成面部表情信号是否能提升共情响应能力,且无需端到端重训练。我们构建了一个可扩展的模拟教学环境,其中学生代理从大规模未标注的人类面部表情视频数据集中表现多样化的面部行为,并对比四种导师变体:纯文本基线、使用随机面部帧的多模态基线,以及两种基于动作单元估计模型(AUM)的方法——一种注入文本化AU描述,另一种选择峰值表达帧进行视觉定位。在涵盖三个导师主干模型(GPT-5.1、Claude Opus 4.5 和 Gemini 2.5 Pro)的960次多轮对话中,由五位人类评分员与一个全面的AI评估器进行配对比较。结果显示,所有主干模型下,基于AU的条件化均显著提升对表情的共情响应;而基于AUM的峰值帧选择优于随机帧输入。文本化AU抽象与峰值帧视觉注入表现出模型依赖优势。控制分析表明,该提升未以牺牲教学清晰度或对文本线索的响应为代价。总体而言,轻量级、结构化的面部表情表示可显著增强基于LLM的教学系统的共情能力,且开销极小。未来工作应使用真实学习互动中采集的学生面部数据验证这些方法。

原文摘要 · Abstract (English)

Large language models (LLMs) enable increasingly capable tutoring-style conversational agents, yet effective tutoring requires sensitivity to learners' affective and cognitive states beyond text alone. Facial expressions provide immediate and practical cues of confusion, frustration, or engagement, but remain underexplored in LLM-driven tutoring. We investigate whether facial-expression-aware signals can improve empathetic tutoring responses through prompt-level integration, without end-to-end retraining. We build a scalable simulated tutoring environment where a student agent exhibits diverse facial behaviors from a large unlabeled human facial expression video dataset, and compare four tutor variants: a text-only LLM baseline, a multimodal baseline using a random facial frame, and two Action Unit estimation model (AUM)-based methods that either inject textual AU descriptions or select a peak-expression frame for visual grounding. Across 960 multi-turn conversations spanning three tutor backbones (GPT-5.1, Claude Opus 4.5, and Gemini 2.5 Pro), we evaluate targeted pairwise comparisons with five human raters and an exhaustive AI evaluator. AU-based conditioning consistently improves empathetic responsiveness to facial expressions across all tutor backbones, while AUM-guided peak-frame selection outperforms random-frame visual input. Textual AU abstraction and peak-frame visual injection show model-dependent advantages. Control analyses show that this improvement does not come at the expense of worse pedagogical clarity or responsiveness to textual cues. Overall, our results show that lightweight, structured facial expression representations can meaningfully enhance empathy in LLM-based tutoring systems with minimal overhead. Future work should evaluate these methods using real student facial data collected during authentic learning interactions.

AI导师共情生成面部识别多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。