arXiv:2502.03711cs.CLcs.AI2025-02被引 2

用自动化众包方式测试大模型回答的鲁棒性,发现其在扰动下仍保持稳定。

MultiQ&A: An Analysis in Measuring Robustness via Automated Crowdsourcing of Question Perturbations and Answers

  • 通过独立LLM代理大规模生成问题扰动和对应答案
  • 分析190万次问题扰动和230万条回答,量化幻觉与不一致
  • 适合评估企业级大模型部署时的可信度与一致性

大型语言模型(LLMs)在生成回答时常出现幻觉,阻碍其在机构中的应用。为此,我们提出MultiQ&A,一种系统性评估LLM回答鲁棒性与一致性的方法。该方法通过独立的LLM代理大规模众包生成问题扰动及其相应答案。实验共分析190万次问题扰动和230万条回答。结果表明,集成式LLM如gpt-3.5-turbo在扰动下仍具有相对鲁棒性和一致性。MultiQ&A为回答生成过程提供了清晰视图,有效检测分歧与变异性,可作为机构采纳大模型的潜在框架,实现对置信度、一致性及幻觉程度的量化评估。

原文摘要 · Abstract (English)

One critical challenge in the institutional adoption journey of Large Language Models (LLMs) stems from their propensity to hallucinate in generated responses. To address this, we propose MultiQ&A, a systematic approach for evaluating the robustness and consistency of LLM-generated answers. We demonstrate MultiQ&A's ability to crowdsource question perturbations and their respective answers through independent LLM agents at scale. Our experiments culminated in the examination of 1.9 million question perturbations and 2.3 million answers. Furthermore, MultiQ&A shows that ensembled LLMs, such as gpt-3.5-turbo, remain relatively robust and consistent under perturbations. MultiQ&A provides clarity in the response generation space, offering an effective method for inspecting disagreements and variability. Therefore, our system offers a potential framework for institutional LLM adoption with the ability to measure confidence, consistency, and the quantification of hallucinations.

大模型评估鲁棒性测试幻觉检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。