AI生成题目时,不同表达方式会影响最终被专家评审的内容。
What Reaches Expert Review? Representation, Structural Screening, and Candidate-Form Dependence in AI-Assisted Item Development
- 用语义表示和结构筛选决定哪些题目进入专家评审
- 同一组题目在不同配置下仅保留6/40个相同题目
- 适合心理测量设计者和AI评估系统开发者参考
在AI辅助题目生成与专家评审之间,存在一个计算评估器,其决策常被视为技术前置步骤。然而,表示方式、结构简化和选择策略决定了心理学家最终收到哪些题目和证据。在两项关联的模拟研究中,基于32,000个精选的大五人格题目,我们从语义表示经结构评估到候选题型构建,追踪了固定来源群体的演化过程。语义空间的整体一致性掩盖了关键局部差异:相同表述获得不同的构念证据,不同题目被保留,意图属性甚至在群体对应性提升时消失。这些敏感性在不同生成源群体间也存在差异。在最终评审边界,两种准入政策均填满所有内容单元,但呈现不同措辞。在不同嵌入配置下,包容性主形式间仅共享6/40个题目,反映出表示变化对结构证据和排序的全链条影响。全局摘要与完整形式的表面稳定性,掩盖了真正进入心理测量师视野的内容的不稳定性。因此,计算评估器并非生成与专家之间的中立基础设施,而是可审视、可重构的测量设计组成部分。
原文摘要 · Abstract (English)
Between AI-assisted item generation and expert review sits a computational evaluator whose decisions are usually treated as technical preliminaries. Yet representation, structural reduction, and selection policy determine which items and evidence psychometricians ever receive. Across two linked in-silico studies of 32,000 selected Big Five items, we followed fixed source populations from semantic representation through structural evaluation and candidate-form construction. Broad agreement in semantic geometry concealed consequential local differences: identical wording acquired different construct evidence, different items survived, and intended attributes could disappear even as community correspondence improved. These sensitivities also differed across generated source populations. At the final review boundary, both eligibility policies filled every content cell in every evaluable form, yet they presented different wording. Across embedding configurations, inclusive primary forms shared a median of only 6 of 40 items, reflecting the total downstream consequence of changing representation across structural evidence and ranking. The apparent stability of global summaries and complete forms therefore concealed instability in the content reaching psychometricians. The computational evaluator is not neutral infrastructure between generation and expertise; it is an inspectable and revisable part of measurement design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。