多模型协作提升复杂问题回答可靠性
Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models
- 用多个大模型协同生成和回答高难度统计问题
- Claude和GPT-4生成问题更清晰,一致性更高
- 适合研究AI协作推理与可信问答系统的人
我们提出一个协作框架,让多个大语言模型(包括GPT-4-0125-preview、Meta-LLaMA-3-70B-Instruct、Claude-3-Opus和Gemini-1.5-Flash)在缺乏明确真实答案的情况下,共同回答复杂的博士级统计问题。通过卡方检验、Fleiss' Kappa和置信区间分析,量化模型间的一致性与评分者间一致率,评估回答精确度与问题质量。结果显示,Claude与GPT-4生成的问题结构更优、歧义更少,表现为更窄的置信区间和更高的与生成模型的一致性;而Gemini与LLaMA在问题构建上变异更大,可靠性较低。研究表明,大模型间的协作可显著提升回答可靠性,并为优化人工智能驱动的协同推理系统提供重要参考。
原文摘要 · Abstract (English)
We propose a collaborative framework in which multiple large language models -- including GPT-4-0125-preview, Meta-LLaMA-3-70B-Instruct, Claude-3-Opus, and Gemini-1.5-Flash -- generate and answer complex, PhD-level statistical questions when definitive ground truth is unavailable. Our study examines how inter-model consensus improves both response reliability and identifies the quality of the generated questions. Employing chi-square tests, Fleiss' Kappa, and confidence interval analysis, we quantify consensus rates and inter-rater agreement to assess both response precision and question quality. Key results indicate that Claude and GPT-4 produce well-structured, less ambiguous questions with a higher inter-rater agreement, as shown by narrower confidence intervals and greater alignment with question-generating models. In contrast, Gemini and LLaMA exhibit greater variability and lower reliability in question formulation. These findings demonstrate that collaborative interactions among large language models enhance response reliability and provide valuable insights for optimizing AI-driven collaborative reasoning systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。