arXiv:2607.11981cs.CLcs.AI2026-07

提出条件化信度框架,评估自动作文评分在不同作答条件下的可靠性差异。

Evaluating Nonuniform Dependability Across Response Conditions: A Conditional Generalizability Framework Illustrated in Automated Essay Scoring

  • 将评分配置视为可测量条件的总体,而非偶然选择。
  • 在熵分层下,信度在高熵组下降至0.84,显示决策需求差异。
  • 适用于需要精准评估评分系统稳定性的教育测评研究者。

整体信度估计可能掩盖不同作答条件下测量负担的异质性,单一G-或D研究可能误判设计在特定子群体中的适用性。本研究提出一种包含三个部分的条件化一般化框架:首先,将自动评分配置(编码器架构与评分头组合)视为固定流程内的可接受测量条件总体,而非偶然建模选择;其次,通过分析D研究预测与有限评分池的实证配置扫描比较,获得两个设计适切性估计量,其一致性或分歧可用于诊断实际配置总体;第三,将证据基于熵定义的作答分层进行条件化,将熵作为操作性分层变量而非写作质量的构念主张。尽管近期广义理论扩展关注生成式题目变体,本框架解决的是评分端的类似问题——由AI中介的评分配置。以限时第二语言写作的自动作文评分为例,整体设计具有较高依赖性(Phi ≈ 0.76)。在熵分层重估后,依赖性仍保持高位但适度且稳健下降(Phi = 0.88, 0.87, 0.84),呈现梯度趋势,表明不同分层需不同交叉条件,最高熵组要求最复杂的交叉设计。该框架提供了一种可迁移的评估非均匀依赖性的工作流。

原文摘要 · Abstract (English)

Aggregate reliability estimates can obscure heterogeneity in measurement-design burden across response conditions, so a single G- or D-study may mischaracterize a design's adequacy for particular strata. This study introduces a conditional generalizability framework with three components. First, automated scoring configurations -- the encoder architectures and scoring-head families admissible within a fixed pipeline -- are treated as a universe of admissible measurement conditions rather than incidental modeling choices. Second, analytical D-study projections are compared with empirical configuration sweeps over a finite scoring pool, yielding two estimands of design adequacy whose agreement or divergence diagnoses the realized configuration universe. Third, evidence is conditioned on entropy-defined response strata, treating entropy as an operational stratification variable, not a construct claim about writing quality. Whereas recent generalizability-theory extensions address AI-generated item variants on the response side, this framework addresses the analogous scoring-side problem: AI-mediated scoring configurations. Demonstrated with automated essay scoring of timed L2 writing, the realized design was dependable in aggregate (Phi approx 0.76). Re-estimated within entropy strata, dependability stayed high but declined modestly and robustly (Phi = 0.88, 0.87, 0.84) -- a gradient implying different decision-study requirements, the highest-entropy stratum requiring the most crossed conditions. The framework offers a portable workflow for evaluating nonuniform dependability.

自动评分信度分析教育测评

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。