用AI帮工科生提升演讲能力,自动分析语言与肢体动作。
Enhancing Public Speaking Skills in Engineering Students Through AI
- 融合语音、视觉和情感检测,多模态评估演讲表现。
- 生成反馈与专家评价中度一致,Gemini Pro表现最佳。
- 适合需要反复练习演讲的工科学生和教育者使用。
本研究针对工程专业学生沟通能力不足的问题,提出一种基于人工智能的演讲能力评估模型。该模型整合语音分析、计算机视觉与情感识别,从语言(音调、音量、语速、语调)、非语言(面部表情、手势、姿态)及表达一致性三个维度进行综合评估。不同于以往分项评估的系统,本模型通过多模态融合实现个性化、可扩展的反馈。初步测试显示,AI生成反馈与专家评价中度一致。在对比的多个大语言模型中,Gemini Pro表现最优,与人工标注者一致性最高。该系统摆脱对人工评审的依赖,支持反复练习,帮助学生自然协调语言与身体语言,提升专业沟通效果。
原文摘要 · Abstract (English)
This research-to-practice full paper was inspired by the persistent challenge in effective communication among engineering students. Public speaking is a necessary skill for future engineers as they have to communicate technical knowledge with diverse stakeholders. While universities offer courses or workshops, they are unable to offer sustained and personalized training to students. Providing comprehensive feedback on both verbal and non-verbal aspects of public speaking is time-intensive, making consistent and individualized assessment impractical. This study integrates research on verbal and non-verbal cues in public speaking to develop an AI-driven assessment model for engineering students. Our approach combines speech analysis, computer vision, and sentiment detection into a multi-modal AI system that provides assessment and feedback. The model evaluates (1) verbal communication (pitch, loudness, pacing, intonation), (2) non-verbal communication (facial expressions, gestures, posture), and (3) expressive coherence, a novel integration ensuring alignment between speech and body language. Unlike previous systems that assess these aspects separately, our model fuses multiple modalities to deliver personalized, scalable feedback. Preliminary testing demonstrated that our AI-generated feedback was moderately aligned with expert evaluations. Among the state-of-the-art AI models evaluated, all of which were Large Language Models (LLMs), including Gemini and OpenAI models, Gemini Pro emerged as the best-performing, showing the strongest agreement with human annotators. By eliminating reliance on human evaluators, this AI-driven public speaking trainer enables repeated practice, helping students naturally align their speech with body language and emotion, crucial for impactful and professional communication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。