用语言和眼神线索实时预测听者理解状态,提升教学交互效果。
Predicting States of Understanding in Explanatory Interactions Using Cognitive Load-Related Linguistic Cues
- 通过说话人话语的意外性与句法复杂度、听者注视变化来捕捉认知负荷
- 三类语言线索结合文本特征,对四种理解状态分类准确率显著提升
- 适合教育技术、人机交互领域研究者参考
我们研究对话中说话人与听者的语言特征如何用于在解释性互动中实时预测听者的理解状态。具体分析三类与认知负荷相关的语言线索:说话人话语的信息量(以意外性衡量)和句法复杂度,以及听者互动时的注视行为变化。基于面对面桌游解释的MUNDEX语料库进行统计分析发现,这些线索与听者理解水平相关。听者通过回溯视频自标注其状态(完全理解、部分理解、不理解、误解)。后续分类实验使用两种现成分类器及一个微调的德语BERT多模态分类器,结果表明四类理解状态可被有效预测,且加入三类语言线索后性能进一步提升。
原文摘要 · Abstract (English)
We investigate how verbal and nonverbal linguistic features, exhibited by speakers and listeners in dialogue, can contribute to predicting the listener's state of understanding in explanatory interactions on a moment-by-moment basis. Specifically, we examine three linguistic cues related to cognitive load and hypothesised to correlate with listener understanding: the information value (operationalised with surprisal) and syntactic complexity of the speaker's utterances, and the variation in the listener's interactive gaze behaviour. Based on statistical analyses of the MUNDEX corpus of face-to-face dialogic board game explanations, we find that individual cues vary with the listener's level of understanding. Listener states ('Understanding', 'Partial Understanding', 'Non-Understanding' and 'Misunderstanding') were self-annotated by the listeners using a retrospective video-recall method. The results of a subsequent classification experiment, involving two off-the-shelf classifiers and a fine-tuned German BERT-based multimodal classifier, demonstrate that prediction of these four states of understanding is generally possible and improves when the three linguistic cues are considered alongside textual features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。