arXiv:2510.20584cs.CLcs.AI2025-10中稿 · the Journal of Edu…

用ChatGPT自动编码沟通数据,表现跨性别种族一致。

Automated Coding of Communication Data Using ChatGPT: Consistency Across Subgroups

  • 基于编码标准直接指令ChatGPT进行文本分类
  • 在三种协作任务中对不同性别/族裔群体表现一致
  • 支持其用于大规模沟通协作评估

大规模评估沟通与协作依赖于人工将沟通数据按框架分类,耗时费力。已有研究证明,通过直接提供编码标准,ChatGPT可实现与人工评判相当的准确率。但其在不同人口学群体(如性别、种族)间是否具有一致性尚不明确。为填补这一空白,我们引入三种检查方法,借鉴自动化评分领域的既有框架,评估基于大语言模型的编码在子群体间的稳定性。利用典型协作问题解决编码框架及三类协作任务数据,我们检验了ChatGPT在性别与种族/族裔群体中的编码表现。结果表明,其编码一致性与人类评分者相当,验证了该技术在大规模协作与沟通评估中的可行性。

原文摘要 · Abstract (English)

Assessing communication and collaboration at scale depends on a labor-intensive task of coding communication data into categories according to different frameworks. Prior research has established that ChatGPT can be directly instructed with coding rubrics to code the communication data and achieves accuracy comparable to human raters. However, whether the coding from ChatGPT or similar AI technology perform consistently across different demographic groups, such as gender and race, remains unclear. To address this gap, we introduce three checks for evaluating subgroup consistency in LLM-based coding by adapting an existing framework from the automated scoring literature. Using a typical collaborative problem-solving coding framework and data from three types of collaborative tasks, we examine ChatGPT-based coding performance across gender and racial/ethnic groups. Our results show that ChatGPT-based coding perform consistently in the same way as human raters across gender or racial/ethnic groups, demonstrating the possibility of its use in large-scale assessments of collaboration and communication.

AI编码大语言模型公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。