arXiv:2411.00492cs.CL2024-11EMNLP被引 23

用多专家模拟提升大模型回答的可信度与安全性

Multi-expert Prompting Improves Reliability, Safety, and Usefulness of Large Language Models

  • 模拟多个专家独立思考后合并选择最优答案
  • 在多个评测中比基线高8.69%的可信度表现
  • 无需手动调参,适合各类实际应用场景

我们提出Multi-expert Prompting,一种对ExpertPrompting(Xu等,2023)的改进方法,旨在提升大语言模型(LLM)生成质量。该方法通过模拟多位专家分别思考、聚合结果,并在个体与聚合结果中择优,完成指令理解与生成。整个过程基于名义小组技术(Nominal Group Technique, NGT)设计的七个子任务,在单链思维中实现。实验表明,该方法显著优于ExpertPrompting及同类基线,在真实性、事实性、信息量和实用性方面均有提升,同时降低毒性与伤害性。尤其在可信度上超越最佳基线8.69%(以ChatGPT为参照)。该方法高效、可解释,且高度适应不同场景,无需人工构建提示词。

原文摘要 · Abstract (English)

We present Multi-expert Prompting, a novel enhancement of ExpertPrompting (Xu et al., 2023), designed to improve the large language model (LLM) generation. Specifically, it guides an LLM to fulfill an input instruction by simulating multiple experts, aggregating their responses, and selecting the best among individual and aggregated responses. This process is performed in a single chain of thoughts through our seven carefully designed subtasks derived from the Nominal Group Technique (Ven and Delbecq, 1974), a well-established decision-making framework. Our evaluations demonstrate that Multi-expert Prompting significantly outperforms ExpertPrompting and comparable baselines in enhancing the truthfulness, factuality, informativeness, and usefulness of responses while reducing toxicity and hurtfulness. It further achieves state-of-the-art truthfulness by outperforming the best baseline by 8.69% with ChatGPT. Multi-expert Prompting is efficient, explainable, and highly adaptable to diverse scenarios, eliminating the need for manual prompt construction.

大模型提示工程可信生成多专家模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。