构建首个用于研究二语习得者习语认知负荷的眼动数据集。
Assessing Cognitive Effort in L2 Idiomatic Processing: An Eye-Tracking Dataset

- 基于眼动追踪记录葡萄牙语母语者英语学习者的阅读行为。
- 发现语言水平越高,回视次数越少,认知负担越轻。
- 适合研究语言模型与人类理解习语方式的差异。
本文开发并验证了一个眼动追踪数据集,用于研究第二语言(L2)学习者处理习语表达的认知机制。母语者通常直接提取习语的隐喻意义,而二语学习者常采用字面先行策略,导致可测量的认知代价。该数据集通过眼动指标捕捉了葡萄牙语母语的英语学习者在所有CEFR水平(A1-C2)下的认知成本。尽管使用的是入门级60 Hz硬件(Tobii Pro Spark),研究证明该采样率足以检测阅读中的注视和回视等宏观认知事件。初步分析显示,语言水平与回视次数呈显著负相关。该数据集已整合进MIA(Modeling Idiomaticity in Human and Artificial Language Processing)计划,为评估人类认知模型及大语言模型与人类习语理解一致性的基准提供支持。
原文摘要 · Abstract (English)
This paper presents the development and validation of an eye-tracking dataset designed to investigate how second-language (L2) learners process idiomatic expressions. While native speakers often rely on direct retrieval of figurative meanings, L2 speakers frequently adopt a literal-first approach, which incurs measurable cognitive costs. This resource captures these costs through ocular metrics recorded from Portuguese L1 speakers of English across all CEFR proficiency levels (A1-C2). Although the study uses entry-level 60 Hz hardware (Tobii Pro Spark), we demonstrate that this sampling rate provides sufficient data density to detect macro-cognitive events such as fixations and regressions in reading. Preliminary analysis validates the dataset by revealing a strong inverse correlation between language proficiency and regressive eye movements. Integrated into the MIA (Modeling Idiomaticity in Human and Artificial Language Processing) initiative, this dataset serves as a cognitively grounded benchmark for evaluating both human processing models and the alignment of large language models with human-like figurative understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。