简单提示法在论辩型大模型中比复杂方法更有效量化不确定性
Evaluating Uncertainty Quantification Methods in Argumentative Large Language Models
- 用直接提示法评估论辩大模型的不确定性
- 简单提示法在论点验证任务中表现显著优于复杂方法
- 适合关注模型可信度与可解释性的研究者
大语言模型(LLM)的不确定性量化(UQ)研究对保障该技术可靠性日益重要。本文探讨将LLM UQ方法融入论辩型大语言模型(ArgLLMs)——一种基于计算论辩的可解释决策框架,其中UQ起关键作用。通过在论点验证任务上测试不同UQ方法的表现,评估其有效性。实验过程本身成为一种新颖的评估方式,尤其适用于复杂且可能引发争议的陈述。结果表明,尽管结构简单,直接提示法在ArgLLMs中表现出色,显著优于更复杂的策略。
原文摘要 · Abstract (English)
Research in uncertainty quantification (UQ) for large language models (LLMs) is increasingly important towards guaranteeing the reliability of this groundbreaking technology. We explore the integration of LLM UQ methods in argumentative LLMs (ArgLLMs), an explainable LLM framework for decision-making based on computational argumentation in which UQ plays a critical role. We conduct experiments to evaluate ArgLLMs' performance on claim verification tasks when using different LLM UQ methods, inherently performing an assessment of the UQ methods' effectiveness. Moreover, the experimental procedure itself is a novel way of evaluating the effectiveness of UQ methods, especially when intricate and potentially contentious statements are present. Our results demonstrate that, despite its simplicity, direct prompting is an effective UQ strategy in ArgLLMs, outperforming considerably more complex approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。