研究多个大模型协作问答时,如何避免幻觉传播并提升准确率。
Collaborative QA using Interacting LLMs. Impact of Network Structure, Node Capability and Distributed Data
- 用平均场动力学建模模型间信息传播,结合经济随机效用理论分析行为。
- 100个开源模型实验表明,网络结构与节点能力显著影响结果准确性。
- 适合关注多模型协作、可信AI的开发者和研究人员参考。
本文研究分布式文档下多个交互式大语言模型协作问答(CQA)的性能,以逼近真实答案。由于大模型在缺乏直接证据时常产生幻觉,且这些错误在模型网络中会扩散,导致原本准确的模型也出现幻觉。为此,我们结合网络科学中的平均场动力学(MFD)与经济学中的随机效用模型,构建了一个可生成的分析框架。模型将每个大模型的状态设为潜在的真伪值,并通过可解析的平均场模型描述有向网络中信息的传播过程。基于随机效用模型设定动态概率,推导出固定点存在的充分条件,并分析不同激励(如推理时计算资源)对固定点行为的影响。我们在包含100个开源大模型的网络上,针对数据异质性、节点能力、网络结构及提示框架敏感性,在多个半合成数据集上进行实验与分析。
原文摘要 · Abstract (English)
In this paper, we model and analyze how a network of interacting LLMs performs collaborative question-answering (CQA) in order to estimate a ground truth given a distributed set of documents. This problem is interesting because LLMs often hallucinate when direct evidence to answer a question is lacking, and these effects become more pronounced in a network of interacting LLMs. The hallucination spreads, causing previously accurate LLMs to hallucinate. We study interacting LLMs and their hallucination by combining novel ideas of mean-field dynamics (MFD) from network science and the randomized utility model from economics to construct a useful generative model. We model the LLM with a latent state that indicates if it is truthful or not with respect to the ground truth, and extend a tractable analytical model considering an MFD to model the diffusion of information in a directed network of LLMs. To specify the probabilities that govern the dynamics of the MFD, we propose a randomized utility model. For a network of LLMs, where each LLM has two possible latent states, we posit sufficient conditions for the existence and uniqueness of a fixed point and analyze the behavior of the fixed point in terms of the incentive (e.g., test-time compute) given to individual LLMs. We experimentally study and analyze the behavior of a network of $100$ open-source LLMs with respect to data heterogeneity, node capability, network structure, and sensitivity to framing on multiple semi-synthetic datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。