arXiv:2512.00009cs.HCcs.AI2025-12被引 1

用AI助手辅助定性研究,与人类标注一致性达0.71,可识别错误并纠正偏见。

Development and Benchmarking of a Blended Human-AI Qualitative Research Assistant

  • 开发交互式AI系统Muse,支持主题识别与数据标注
  • 与人类标注者一致性κ=0.71,达到良好可靠水平
  • 能发现分析错误并修正人类偏见,适合研究团队使用

定性研究依赖对文本数据的反复互动以构建意义。传统人工方法易受编码疲劳和解释偏差影响,难以应对大规模复杂数据集。尽管计算方法曾遭质疑,因难以复现人类分析的细微差别与上下文感知能力,但大语言模型为自动化定性分析提供了新可能。为评估其优劣并建立研究者信任,需在人类标注数据集上严格测试。本文对Muse——一个交互式AI辅助定性研究系统进行基准测试,该系统支持主题识别与数据标注。结果显示,Muse与人类标注者间具有一致性,对明确界定的代码,科恩kappa值κ=0.71。同时,通过深入误差分析,识别出系统失效模式,指导未来改进,并验证了其纠正人类偏见的能力。

原文摘要 · Abstract (English)

Qualitative research emphasizes constructing meaning through iterative engagement with textual data. Traditionally this human-driven process requires navigating coder fatigue and interpretative drift, thus posing challenges when scaling analysis to larger, more complex datasets. Computational approaches to augment qualitative research have been met with skepticism, partly due to their inability to replicate the nuance, context-awareness, and sophistication of human analysis. Large language models, however, present new opportunities to automate aspects of qualitative analysis while upholding rigor and research quality in important ways. To assess their benefits and limitations - and build trust among qualitative researchers - these approaches must be rigorously benchmarked against human-generated datasets. In this work, we benchmark Muse, an interactive, AI-powered qualitative research system that allows researchers to identify themes and annotate datasets, finding an inter-rater reliability between Muse and humans of Cohen's $κ$ = 0.71 for well-specified codes. We also conduct robust error analysis to identify failure mode, guide future improvements, and demonstrate the capacity to correct for human bias.

定性分析AI助手大模型标注一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。