arXiv:2605.24600cs.AI2026-05中稿 · EMNLP被引 1

用多智能体模拟人类质性分析中的同行评议,提升大模型分析的可信度。

Agent-as-Peer-Debriefer: A Multi-Agent Framework with Perspective-Based Refinement for Qualitative Analysis

论文配图:Agent-as-Peer-Debriefer: A Multi-Agent Framework with Perspective-Based Refinement for Qualitative Analysis
图 1 · 摘自论文原文
  • 设计三个不同视角的智能体,对编码结果进行批判性修正。
  • 在三个数据集上测试,结果与人工标注代码的语义相似度显著提升。
  • 不同分析视角带来可调控的权衡,适合需要多维审校的研究场景。

大型语言模型(LLMs)在质性数据分析(QDA)中应用日益广泛,但其输出常缺乏人类分析的深度与细微差别。我们指出这一差距源于人类质性分析中缺失的一种可信实践:同行评议——分析师向无关第三方寻求反馈并据此修正编码。为将此实践融入大模型辅助的质性分析,我们提出「Agent-as-Peer-Debriefer」多智能体框架,在关键编码步骤中嵌入同行评议机制。该框架中,层级编码智能体按标准流程生成代码、子主题和主题,并附带自我解释与反思笔记;随后将其输出分享给三个采用不同分析视角(理论驱动、数据驱动、应用导向)的同行评议智能体,分别通过保留、重命名、重新分配、合并或拆分代码来优化结果。这些视角源自跨领域、跨数据集通用的人类质性分析方法。我们在两个领域的三个数据集上,使用三种LLM进行评估,以语义相似度衡量与人工标注代码的一致性。所有设置下,基于视角的同行评议修正均优于单一模型基线,且消融实验表明收益不仅来自额外修正。三种视角产生明显不同的权衡,表明视角选择是可控制且有意义的设计决策。总体而言,通过显式视角模拟同行评议,是提升大模型辅助质性分析可信度的可行路径。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used for qualitative data analysis (QDA), yet their outputs often miss the depth and nuance of human analysis. We argue this gap reflects a missing credibility practice from human QDA: peer debriefing, in which an analyst seeks feedback from a disinterested peer and uses it to refine their coding. To bring this practice into LLM-assisted QDA, we propose Agent-as-Peer-Debriefer, a multi-agent QDA framework that builds peer debriefing into key coding steps. In our framework, a Hierarchical Coding Agent follows the standard QDA process to generate codes, sub-themes, and themes, along with self-explanations and reflection memos. It then shares these outputs with three Peer-Debriefing Agents, each applying a distinct analytical perspective (Theory-Driven, Data-Driven, or Applied) and refining the codes by keeping, renaming, reassigning, merging, or splitting them. These perspectives are drawn from established human QDA practices that generalize across domains and datasets. To evaluate the framework, we test it on three datasets across two domains with three LLMs, measuring semantic similarity to human-annotated codes. Across all settings, perspective-based, peer-debriefing refinement aligns more closely with human codes than a single-LLM baseline, and an ablation further shows the gain is not merely from additional refinement. The three perspectives also produce distinct trade-offs, showing that the choice of perspective is a meaningful and controllable design decision. More broadly, these findings suggest that simulating peer debriefing with explicit perspectives is a promising route to more credible LLM-assisted QDA.

质性分析多智能体同行评议大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。