arXiv:2608.22566cs.CL2026-08中稿 · ICQE 2026

用对话分析诊断并优化多智能体大模型的推理过程。

From Diagnosis to Redesign: Using Quantitative Ethnography to Improve Multi-Agent LLM Reasoning

论文配图:From Diagnosis to Redesign: Using Quantitative Ethnography to Improve Multi-Agent LLM Reasoning
图 1 · 摘自论文原文
  • 通过话语分析识别正确与错误推理的交互模式。
  • 改进提示后评分准确率从27.78%提升至40.28%。
  • 适合关注AI推理可解释性与系统优化的研究者。

多智能体大语言模型系统通过分工协作提升推理能力,但多智能体的存在并不保证推理连贯或结果符合任务目标。本文提出一种定量民族志(QE)方法,基于智能体交互产生的对话进行诊断与系统重设计。以自动作文评分为例,采用知识网络分析(ENA)建模五智能体辩论系统,比较产生正确与错误评分决策的对话差异。结果显示,初始系统中,正确决策表现为基于评分标准的论证、一致性和扩展性;而错误决策则呈现较长的主张-质疑-回应循环,且与评分标准关联较弱。据此优化各智能体提示词,使系统精确评分准确率从27.78%提升至40.28%,错误辩论的对话模式也向正确模式趋近,两者几乎难以区分。研究证明,定量民族志能构建从诊断到重设计的闭环,追踪交互模式与系统性能的关系,指导提示词优化,并验证其对结果与互动模式的影响。

原文摘要 · Abstract (English)

Multi-agent large language model (LLM) systems are designed to improve reasoning by decomposing tasks across multiple agents with specialized functions, but the presence of multiple agents does not inherently guarantee coherent reasoning or outputs that align with task objectives. This paper introduces a quantitative ethnographic (QE) approach for diagnosing and redesigning multi-agent LLM systems based on the discourse produced through agent interactions. We test this approach using automated essay scoring as an example context, applying Epistemic Network Analysis (ENA) to model a five-agent multi-agent debate system and examine differences between debates that produced correct versus incorrect scoring decisions. Results show that, in the initial system, correct scoring decisions were characterized by rubric-grounded justification, agreement, and elaboration. Incorrect scoring decisions, in contrast, were characterized by extended proposition-challenge-response exchanges that were less consistently tied to rubric criteria. We then used the findings to revise the agents' prompts. The revised system improved exact scoring accuracy from 27.78% to 40.28% and shifted the discourse of incorrect debates toward the rubric-grounded pattern of correct ones, making the two nearly indistinguishable. Based on these results, we argue that QE can support a diagnostic-to-redesign loop for AI reasoning by tracing how patterns of agent interaction relate to system performance, informing prompt redesign, and evaluating whether those redesigns change both outcomes and interaction patterns.

多智能体推理优化定量分析提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。