arXiv:2604.19309cs.HCcs.AI2026-04

用AI实时检测定性编码中的理解漂移,提升研究可信度。

Co-Refine: AI-Powered Tool Supporting Qualitative Analysis

论文配图:Co-Refine: AI-Powered Tool Supporting Qualitative Analysis
图 1 · 摘自论文原文
  • 三阶段审计流程:先算数学一致性,再用大模型校准,最后生成反馈
  • 在真实数据上实现98.3%的漂移检测准确率,延迟低于0.5秒
  • 适合需要高信度定性分析的研究者,尤其适合大数据量场景

定性编码依赖研究者对文本数据打标签,但随着分析进行,代码理解常发生随时间变化的偏移(时序漂移),影响分析可信度。现有计算机辅助定性数据分析工具虽能管理数据,却缺乏实时检测漂移的工作流。我们提出Co-Refine,一个由AI增强的定性编码平台,可在不打断研究流程的情况下持续提供基于证据的编码一致性反馈。系统采用三阶段审计管道:第一阶段通过确定性嵌入指标计算数学一致性;第二阶段将大语言模型的判断约束在±0.15的确定性分数范围内;第三阶段从历史模式中提炼代码定义,形成深度反馈循环。实验表明,确定性评分可有效约束大模型输出,生成可靠、实时的审计信号,显著提升定性分析的稳定性与可信度。

原文摘要 · Abstract (English)

Qualitative coding relies on a researcher's application of codes to textual data. As coding proceeds across large datasets, interpretations of codes often shift (temporal drift), reducing the credibility of the analysis. Existing Computer-Assisted Qualitative Data Analysis (CAQDAS) tools provide support for data management but offer no workflow for real-time detection of these drifts. We present Co-Refine, an AI-augmented qualitative coding platform that delivers continuous, grounded feedback on coding consistency without disrupting the researcher's workflow. The system employs a three-stage audit pipeline: Stage 1 computes deterministic embedding-based metrics for mathematical consistency; Stage 2 grounds LLM verdicts within $\pm0.15$ of the deterministic scores; and Stage 3 produces code definitions from previous patterns to create a deepening feedback loop. Co-Refine demonstrates that deterministic scoring can effectively constrain LLM outputs to produce reliable, real-time audit signals for qualitative analysis.

定性分析AI辅助编码一致性LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。