构建了12小时多模态学生演讲数据集,支持智能评估与学习分析。
A Multimodal Dataset of Student Oral Presentations with Sensors and Evaluation Data
- 采集65名学生50场演讲的视听、生理与交互数据
- 平均演讲时长9分48秒,问答环节6分28秒,含评分与标注
- 适合教育技术、自动化评估与行为分析研究者使用
口语表达能力是高等教育的关键组成部分,但涵盖多模态真实表现的综合性数据集仍很稀缺。为填补这一空白,我们提出了SOPHIAS(基于传感器的学生演讲全景洞察与分析),一个包含12小时多模态数据的公开数据集,涵盖马德里自治大学65名本科生与硕士生的50场演讲,其中46场为个人演讲(平均时长9分48秒,标准差33秒),随后为平均6分28秒(标准差3分02秒)的问答环节;4场为小组演讲(平均时长14分11秒,标准差1分40秒),问答平均8分22秒(标准差1分21秒)。数据集整合了8路时间对齐的传感器流:高清网络摄像头、环境与摄像头音频、眼动追踪眼镜、智能手表生理传感器,以及答题器、键盘和鼠标交互。同时包含课件、教师、同伴与自评的评分表,以及时间戳标注。所有演讲均在真实课堂环境中进行,保留了学生的真实行为、互动与生理反应。SOPHIAS支持探究多模态行为与生理信号与演讲表现的关系,促进同伴评估研究,并为自动化反馈与多模态学习分析工具提供基准。数据通过Science Data Bank受控获取,需签署数据使用协议(DUA),代码发布于GitHub。
原文摘要 · Abstract (English)
Oral presentation skills are a critical component of higher education, yet comprehensive datasets capturing real-world student performance across multiple modalities remain scarce. To address this gap, we present SOPHIAS (Student Oral Presentation monitoring for Holistic Insights & Analytics using Sensors), a 12-hour multimodal dataset containing recordings of 50 oral presentations delivered by 65 undergraduate and master's students at the Universidad Autonoma de Madrid, comprising 46 individual presentations with a mean presentation duration of 9 min 48 s (SD = 33 s) followed by a mean Q&A duration of 6 min 28 s (SD = 3 min 02 s), and 4 group presentations with a mean presentation duration of 14 min 11 s (SD = 1 min 40 s) followed by a mean Q&A duration of 8 min 22 s (SD = 1 min 21 s). SOPHIAS integrates eight timestamped sensor streams from high-definition webcams, ambient and webcam audio, eye-tracking glasses, smartwatch physiological sensors, and clicker, keyboard and mouse interactions. In addition, the dataset includes slides and rubric-based evaluations from teachers, peers, and self-assessments, along with timestamped contextual annotations. The dataset captures presentations conducted in real classroom settings, preserving authentic student behaviors, interactions, and physiological responses. SOPHIAS enables the exploration of relationships between multimodal behavioral and physiological signals and presentation performance, supports the study of peer-assessment and provides a benchmark for developing automated feedback and Multimodal Learning Analytics tools. The dataset is available for research, including both academic and legitimate commercial research and development, under controlled access through Science Data Bank, subject to approval of a Data Usage Agreement (DUA), with code provided through GitHub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。