arXiv:2607.11213cs.CL2026-07

AI让语言测试的分数解释变得复杂,需重新设计测试以应对现实中的智能辅助。

When the Target Domain Changes: AI-Mediated Construct Drift in High-Stakes English Language AssessmenW

  • 提出‘有界AI辅助’测试设计,统一控制AI使用边界与记录
  • 指出当前测试仍基于无人辅助模型,与真实学术场景脱节
  • 适合关注语言测评有效性的教育研究者和考试机构

高风险英语能力测试通常将无辅助表现视为学术英语能力的证据。但随着生成式AI在目标语言场景中日益普及,从无辅助测试表现推断学术沟通准备度的合理性逐渐减弱。本文将这一问题定位为分数解释的有效性危机,而非单纯的评分、反馈或安全问题。综述现有文献发现,多数研究仅将AI视为评估基础设施,极少深入探讨其对构念效度和推论依据的影响。本文定义‘AI中介的构念漂移’为:当目标领域因AI介入而改变沟通能力需求,但测试构念仍锚定于无辅助表现时产生的不匹配。提出‘有界AI中介’作为有效性导向的设计原则:所有考生在标准化条件下使用同一受控的AI助手,设定明确帮助边界,记录交互过程,并区分理解支持与答案生成任务。论文主张,在用于推断AI辅助学术交流能力时,应缩小分数解释范围并增加补充说明。

原文摘要 · Abstract (English)

High-stakes English proficiency tests treat standardized, unaided performance as evidence for score interpretations about academic English proficiency. This interpretation remains meaningful, but as target language use domains increasingly involve generative AI, the extrapolation from unaided test performance to academic communicative readiness becomes less self-evident. This conceptual validity argument reframes AI as a score-interpretation problem in high-stakes language testing, not only an operational issue of scoring, feedback, security, or misconduct. Synthesizing current literature in three uneven layers, the paper shows that most work treats AI as assessment infrastructure, while far less theorizes its implications for construct validity and extrapolation warrants. It defines AI-mediated construct drift as the misalignment that arises when communicative abilities required in the target domain change through AI mediation while test constructs remain anchored to an unaided-performance model. It proposes bounded AI mediation as a validity-oriented design principle: a standardized condition in which all test takers access the same institutionally controlled AI assistant, with predefined assistance boundaries, logged interactions, and tasks that distinguish comprehension support from answer generation. The paper argues that score interpretations should be narrowed and supplemented when used to support claims about AI-mediated academic communication.

语言测试AI中介构念效度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。