研究温度与角色对大模型协作编码的影响,发现未显著提升准确率
Temperature and Persona Shape LLM Agent Consensus With Minimal Accuracy Gains in Qualitative Coding
- 设计多角色大模型协作系统模拟人类编码流程
- 温度影响共识达成速度,角色多样性延缓共识但未提准
- 适合关注人机协作机制与代码本优化的研究者
大语言模型(LLMs)为大规模定性研究带来新可能,包括教育数据的标注与编码。尽管基于LLM的多智能体系统(MAS)可模拟人类编码工作流,但其相较于单个模型在编码上的优势尚不明确。为此,我们通过实验研究了组件智能体的角色与温度如何影响对话片段的共识构建与编码准确率。使用包含8个代码且具有良好评分者间信度的成熟代码本,对6个开源大模型(参数量3亿至32亿)在18种配置下进行测试,分析超过7.7万次编码决策,参照在线数学辅导会话的人类标注金标准数据集。温度显著影响所有6个模型达成共识的时间与方式;采用多种角色(中立、强势、共情)的系统在4个模型上比统一角色更延迟共识;其中3个模型中,高温削弱了角色多样性对共识的影响。然而,温度与角色组合均未带来稳定的编码准确率提升,多数情况下单个模型表现不低于或优于多智能体共识。对多智能体协作过程及编码分歧的定性分析,或有助于改进代码本设计与人机协同编码。
原文摘要 · Abstract (English)
Large Language Models (LLMs) enable new possibilities for qualitative research at scale, including annotation and qualitative coding of educational data. While LLM-based multi-agent systems (MAS) can emulate human coding workflows, their benefits over single LLM agents for coding remain poorly understood. To that end, we conducted an experimental study of how persona and temperature of component agents of a MAS shapes consensus-building and coding accuracy for dialog segments. LLMs were prompted to code these segments deductively using a mature codebook with 8 codes and high inter-rater reliability derived from prior research. Our open-source MAS mirrors deductive human coding through structured agent discussion and consensus arbitration. Using six open-source LLMs (with 3 to 32 billion parameters) and 18 experimental configurations, we analyze over 77,000 coding decisions against a gold-standard dataset of human-annotated transcripts from online math tutoring sessions facilitated by educational software. Temperature significantly impacted whether and when consensus was reached across all six LLMs. MAS with multiple personas (including neutral, assertive, or empathetic) significantly delayed consensus in four out of six LLMs compared to uniform personas. In three of those LLMs, higher temperatures significantly diminished the effects of multiple personas on consensus. However, neither temperature nor persona pairing led to robust improvements in coding accuracy. Single agents matched or outperformed MAS consensus in most conditions. Qualitative analysis of MAS collaboration and coding disagreement may, however, improve codebook design and human-AI coding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。