让大模型的不确定性评估更贴近真实用户决策需求
From Calibration to Collaboration: LLM Uncertainty Quantification Should Be More Human-Centered
- 从用户视角重构评估方法,关注实际任务中的可靠性
- 识别三类阻碍有效评估的常见问题:场景失真、忽略认知不确定性等
- 适合关注人机协作与可信AI的研究者和工程师
大型语言模型在现实世界中日益承担辅助角色,但其可靠性仍存疑。不确定性量化(UQ)被视为提升人-模型协作的关键,使用户知晓何时应信任模型输出。通过对40种LLM UQ方法的分析,我们发现当前实践存在三大问题:1)在生态效度低的基准上评估;2)仅关注认知不确定性(epistemic uncertainty);3)优化与下游实用价值无关的指标。针对每项问题,我们提出以用户为中心的具体改进方向与研究路径。我们主张,研究社区不应在非代表性的任务上盲目优化指标,而应转向更以人为本的不确定性量化范式。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly assisting users in the real world, yet their reliability remains a concern. Uncertainty quantification (UQ) has been heralded as a tool to enhance human-LLM collaboration by enabling users to know when to trust LLM predictions. We argue that current practices for uncertainty quantification in LLMs are not optimal for developing useful UQ for human users making decisions in real-world tasks. Through an analysis of 40 LLM UQ methods, we identify three prevalent practices hindering the community's progress toward its goal of benefiting downstream users: 1) evaluating on benchmarks with low ecological validity; 2) considering only epistemic uncertainty; and 3) optimizing metrics that are not necessarily indicative of downstream utility. For each issue, we propose concrete user-centric practices and research directions that LLM UQ researchers should consider. Instead of hill-climbing on unrepresentative tasks using imperfect metrics, we argue that the community should adopt a more human-centered approach to LLM uncertainty quantification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。