arXiv:2510.04048cs.AI2025-10被引 3

通过可变投票阈值提升大模型回答可信度

Increasing LLM response trustworthiness using voting ensembles

  • 设计可让模型在把握不足时弃权的投票集成机制
  • 在数学与医疗笔记问答中,可信度显著提升
  • 适合对准确性要求高但允许部分问题不回答的场景

尽管大语言模型取得巨大进展,但仍缺乏便捷可靠的不确定性量化方法,使其在高风险应用中难以信赖。一种简单提升准确性的方法是选择多个响应的众数,即集成。本文拓展常规集成方法,提出可变投票阈值的集成策略。我们建立问答理论框架,证明当主导答案未达阈值时允许集成‘弃权’,可显著提升剩余答案的可信度。基于此框架,我们在数学题求解与临床笔记问答两个领域开展实验,结果表明:使用高度严格的投票集成,可在响应率与准确率小幅下降的前提下,实现回答可信度的大幅提高。因此,该方法特别适用于医疗、数据标注等需高确定性但不要求每题都有答案的应用场景。

原文摘要 · Abstract (English)

Despite huge advances, LLMs still lack convenient and reliable methods to quantify the uncertainty in their responses, making them difficult to trust in high-stakes applications. One of the simplest approaches to eliciting more accurate answers is to select the mode of many responses, a technique known as ensembling. In this work, we expand on typical ensembling approaches by looking at ensembles with a variable voting threshold. We introduce a theoretical framework for question answering and show that, by permitting ensembles to "abstain" from providing an answer when the dominant response falls short of the threshold, it is possible to dramatically increase the trustworthiness of the remaining answers. From this framework, we derive theoretical results as well as report experimental results on two problem domains: arithmetic problem solving and clinical-note question-answering. In both domains, we observe that large gains in answer trustworthiness can be achieved using highly restrictive voting ensembles, while incurring relatively modest reductions in response yield and accuracy. Due to this quality, voting ensembles may be particularly useful in applications - such as healthcare and data annotation - that require a high degree of certainty but which may not require that every question receive an automated answer.

大模型可信度集成医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。