首个标注儿童非语言不确定信号的数据集,用于训练实时预测模型。
Learning Multimodal Cues of Children's Uncertainty
- 联合心理学家构建儿童不确定性的多模态标注数据集
- 发现不确定信号与任务难度及表现相关,可被模型有效识别
- 提出新多模态模型,实时视频中预测不确定优于基线
理解不确定性对达成共同认知至关重要,尤其在人机协作中。本文首次与发展与认知心理学家合作,构建了一个标注儿童非语言不确定信号的数据集。通过分析该数据,研究了不确定性在不同任务难度下的作用及其与表现的关系。进一步提出一种多模态机器学习模型,能够基于实时视频片段预测参与者是否处于不确定状态,性能优于基线多模态变换器模型。该工作为人类-人类与人类-AI的认知协调研究提供支持,对手势理解与生成具有广泛意义。匿名化数据与代码将在完成必要同意书和数据表后公开。
原文摘要 · Abstract (English)
Understanding uncertainty plays a critical role in achieving common ground (Clark et al.,1983). This is especially important for multimodal AI systems that collaborate with users to solve a problem or guide the user through a challenging concept. In this work, for the first time, we present a dataset annotated in collaboration with developmental and cognitive psychologists for the purpose of studying nonverbal cues of uncertainty. We then present an analysis of the data, studying different roles of uncertainty and its relationship with task difficulty and performance. Lastly, we present a multimodal machine learning model that can predict uncertainty given a real-time video clip of a participant, which we find improves upon a baseline multimodal transformer model. This work informs research on cognitive coordination between human-human and human-AI and has broad implications for gesture understanding and generation. The anonymized version of our data and code will be publicly available upon the completion of the required consent forms and data sheets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。