用图模型分析对话逻辑连贯性,提升英语口语考试自动评分精度
Automated Speaking Assessment of Conversation Tests with Novel Graph-based Modeling on Spoken Response Coherence
- 构建分层图模型捕捉对话中语义与话语关系
- 在NICT-JLE数据集上显著优于多个基线模型
- 适合研究语言评估与自然语言理解的学者
自动口语评估在对话测试(ASAC)中旨在评估第二语言学习者在与对话伙伴互动时的整体口语能力。尽管已有方法在特定数据集上表现良好,但对对话内部逻辑连贯性的建模仍显不足。为此,我们提出一种分层图模型,融合对话响应间的宏观交互(如话语关系)与细微语义信息(如语义词和说话意图),并结合上下文信息进行最终预测。在NICT-JLE基准数据集上的大量实验表明,该方法在多种评估指标上均显著优于多个强基线模型,凸显了在自动评分中考虑对话连贯性的重要性。
原文摘要 · Abstract (English)
Automated speaking assessment in conversation tests (ASAC) aims to evaluate the overall speaking proficiency of an L2 (second-language) speaker in a setting where an interlocutor interacts with one or more candidates. Although prior ASAC approaches have shown promising performance on their respective datasets, there is still a dearth of research specifically focused on incorporating the coherence of the logical flow within a conversation into the grading model. To address this critical challenge, we propose a hierarchical graph model that aptly incorporates both broad inter-response interactions (e.g., discourse relations) and nuanced semantic information (e.g., semantic words and speaker intents), which is subsequently fused with contextual information for the final prediction. Extensive experimental results on the NICT-JLE benchmark dataset suggest that our proposed modeling approach can yield considerable improvements in prediction accuracy with respect to various assessment metrics, as compared to some strong baselines. This also sheds light on the importance of investigating coherence-related facets of spoken responses in ASAC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。