arXiv:2507.12370cs.CLcs.HC2025-07中稿 · the 2025 SICE Fest…被引 2

用多模型辩论提升大模型对模糊请求的识别能力

Beyond Single Models: Enhancing LLM Detection of Ambiguity in Requests through Debate

  • 让多个大模型通过辩论协作,共同判断用户请求是否模糊
  • 采用辩论机制后,Mistral-7B模型达成76.7%的成功率,尤其擅长处理复杂模糊问题
  • 适合研究大模型协同推理与交互系统鲁棒性的开发者参考

大型语言模型(LLMs)在理解与生成人类语言方面展现出显著能力,推动了与复杂系统更自然的交互。然而,它们在处理用户请求中的模糊性方面仍面临挑战。本文提出并评估了一种多智能体辩论框架,旨在超越单一模型的能力,增强对模糊性的检测与解决。该框架整合了三种LLM架构(Llama3-8B、Gemma2-9B和Mistral-7B变体)及一个包含多样模糊场景的数据集。实验表明,辩论框架显著提升了Llama3-8B和Mistral-7B变体的表现,其中由Mistral-7B主导的辩论实现76.7%的成功率,在处理复杂模糊问题和快速达成共识方面尤为有效。尽管不同模型对协作策略响应各异,但结果证实该框架是提升LLM能力的有力手段。本研究为构建更具鲁棒性与自适应性的语言理解系统提供了重要启示,展示了结构化辩论如何提升交互系统的清晰度。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated significant capabilities in understanding and generating human language, contributing to more natural interactions with complex systems. However, they face challenges such as ambiguity in user requests processed by LLMs. To address these challenges, this paper introduces and evaluates a multi-agent debate framework designed to enhance detection and resolution capabilities beyond single models. The framework consists of three LLM architectures (Llama3-8B, Gemma2-9B, and Mistral-7B variants) and a dataset with diverse ambiguities. The debate framework markedly enhanced the performance of Llama3-8B and Mistral-7B variants over their individual baselines, with Mistral-7B-led debates achieving a notable 76.7% success rate and proving particularly effective for complex ambiguities and efficient consensus. While acknowledging varying model responses to collaborative strategies, these findings underscore the debate framework's value as a targeted method for augmenting LLM capabilities. This work offers important insights for developing more robust and adaptive language understanding systems by showing how structured debates can lead to improved clarity in interactive systems.

大模型多智能体模糊识别辩论框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。