多大模型协作验证答案,无需标准答案也能判断问题和回答质量。
Collective Reasoning Among LLMs: A Framework for Answer Validation Without Ground Truth
- 多个大模型协作生成并互相验证复杂概率题答案。
- Claude和Gemini的题目更清晰,答案一致性更高,置信区间更窄。
- 适合研究AI推理评估、无真值场景下的模型协同方法的人参考。
我们提出一种新方法,让GPT-4-0125-preview、Meta-LLAMA-3-70B-Instruct、Claude-3-Opus和Gemini-1.5-Flash等先进大语言模型协作,解决博士级概率问题,不依赖任何单一正确答案。研究聚焦于不同模型间的一致性如何反映输出可靠性及问题质量。通过卡方检验、Fleiss' Kappa系数和置信区间计算等统计方法,评估答案精度与题干清晰度。分析显示,Claude和Gemini的问题表述更连贯、歧义更少,表现为更紧的置信区间和更高的模型间一致性;而LLAMA则表现出更宽的置信区间和更低的共识度,说明其问题表述更具变异性与不一致。结果表明,多模型协同不仅能提升答案可信度,还为无真值情况下提供了一种数据驱动的问题质量评估与优化机制。本研究为增强异构大模型间的协同推理能力提供了可操作的洞见。
原文摘要 · Abstract (English)
We introduce a new approach in which several advanced large language models-specifically GPT-4-0125-preview, Meta-LLAMA-3-70B-Instruct, Claude-3-Opus, and Gemini-1.5-Flash-collaborate to both produce and answer intricate, doctoral-level probability problems without relying on any single "correct" reference. Rather than depending on an established ground truth, our investigation focuses on how agreement among diverse models can signal the reliability of their outputs and, by extension, reflect the overall quality of the generated questions. To measure this inter-model alignment, we apply a suite of statistical evaluations, including chi-square tests, Fleiss' Kappa coefficients, and confidence interval calculations, thereby capturing both precision in answers and clarity in question phrasing. Our analysis reveals that Claude and Gemini tend to frame questions more coherently and unambiguously, which is evidenced by their tighter confidence intervals and greater concordance with responding agents. In contrast, LLAMA exhibits wider confidence bands and a lower level of agreement, indicating more variability and reduced consistency in its question formulations. These observations support the notion that a multi-model collaborative strategy not only improves answer dependability but also offers an effective, data-driven mechanism for evaluating and refining question quality when no definitive solution exists. Ultimately, this work delivers actionable insights into enhancing AI-guided reasoning processes through coordinated interactions among heterogeneous language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。