构建教育场景下多模态互动自动分析的方法框架
Rapport du Projet de Recherche TRAIMA
- 基于对话分析定义解释性互动的三段式结构
- 通过30小时课堂数据验证转录的可变性与解释性
- 为智能教育研究提供可机器学习的标注标准
TRAIMA项目(2019年3月至2020年6月)探索教育场景中多模态互动的自动化处理。当前口语、副语言及非语言数据的分析依赖人工,耗时且难扩展。项目聚焦法语作为外语(FLE)和第一语言(FLM)教学中的解释性与协作性互动序列,将其视为融合言语、语调、手势、姿态、目光与空间位置的多模态现象。项目提出解释性话语的三段式结构(起始、核心解释、收束),并系统梳理现有转录规范(ICOR、Mondada、GARS、VALIBEL、Ferré),揭示其在课堂数据应用中的优劣。基于约30小时的INTER-EXPLIC语料库与EXPLIC-LEXIC语料库,对比分析人工转录结果,证实转录存在固有变异性和解释依赖性。研究特别关注教师动作(动觉与近距资源)与声调特征的功能作用。项目依托TechnéLAB平台(多摄像头视频、同步音频、眼动追踪、数字交互痕迹)实现多模态数据采集,作为研究基础设施与自动化工具测试环境。最终,项目不提供完整自动化系统,而是建立适用于机器学习的标注框架,强调理论明晰性与研究者反思,为教育学、话语分析、多模态与人工智能交叉研究奠定方法基础。
原文摘要 · Abstract (English)
The TRAIMA project (TRaitement Automatique des Interactions Multimodales en Apprentissage), conducted between March 2019 and June 2020, investigates the potential of automatic processing of multimodal interactions in educational settings. The project addresses a central methodological challenge in educational and interactional research: the analysis of verbal, paraverbal, and non-verbal data is currently carried out manually, making it extremely time-consuming and difficult to scale. TRAIMA explores how machine learning approaches could contribute to the categorisation and classification of such interactions. The project focuses specifically on explanatory and collaborative sequences occurring in classroom interactions, particularly in French as a Foreign Language (FLE) and French as a First Language (FLM) contexts. These sequences are analysed as inherently multimodal phenomena, combining spoken language with prosody, gestures, posture, gaze, and spatial positioning. A key theoretical contribution of the project is the precise linguistic and interactional definition of explanatory discourse as a tripartite sequence (opening, explanatory core, closure), drawing on discourse analysis and interactional linguistics. A substantial part of the research is devoted to the methodological foundations of transcription, which constitute a critical bottleneck for any form of automation. The report provides a detailed state of the art of existing transcription conventions (ICOR, Mondada, GARS, VALIBEL, Ferr{é}), highlighting their respective strengths and limitations when applied to multimodal classroom data. Through comparative analyses of manually transcribed sequences, the project demonstrates the inevitable variability and interpretative dimension of transcription practices, depending on theoretical positioning and analytical goals. Empirical work is based on several corpora, notably the INTER-EXPLIC corpus (approximately 30 hours of classroom interaction) and the EXPLIC-LEXIC corpus, which serve both as testing grounds for manual annotation and as reference datasets for future automation. Particular attention is paid to teacher gestures (kin{é}sic and proxemic resources), prosodic features, and their functional role in meaning construction and learner comprehension. The project also highlights the strategic role of the Techn{é}LAB platform, which provides advanced multimodal data capture (multi-camera video, synchronized audio, eye-tracking, digital interaction traces) and constitutes both a research infrastructure and a test environment for the development of automated tools. In conclusion, TRAIMA does not aim to deliver a fully operational automated system, but rather to establish a rigorous methodological framework for the automatic processing of multimodal pedagogical interactions. The project identifies transcription conventions, annotation categories, and analytical units that are compatible with machine learning approaches, while emphasizing the need for theoretical explicitness and researcher reflexivity. TRAIMA thus lays the groundwork for future interdisciplinary research at the intersection of didactics, discourse analysis, multimodality, and artificial intelligence in education.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。