arXiv:2503.23687cs.CLcs.LG2025-03被引 1

让大模型通过多语言共识判断何时该沉默,提升回答可靠性。

MKA: Leveraging Cross-Lingual Consensus for Model Abstention

  • 利用多语言知识构建自信度校准管道,引导模型在不确定时拒绝回答。
  • 在孟加拉语上准确率提升71.2%,英语也达15.5%提升。
  • 适合关注模型可信性与跨语言鲁棒性的研究者使用。

大型语言模型在任务表现提升的同时,其可靠性仍存疑。要推动广泛应用,关键在于其能否保持事实准确性,或在无法保证时正确校准自身信心。本文提出一种基于多语言共识的模型拒答机制,通过多语言管道校准模型置信度,使其在不确定时选择不回答。我们对多个多语言模型在不同语言下进行测试,发现该管道在多数情况下有效。结果表明,在孟加拉语上的准确率相比基线提升了71.2%,即使是高资源语言英语也实现了15.5%的提升。这些结果暗示了进一步优化的可能性。

原文摘要 · Abstract (English)

Reliability of LLMs is questionable even as they get better at more tasks. A wider adoption of LLMs is contingent on whether they are usably factual. And if they are not, on whether they can properly calibrate their confidence in their responses. This work focuses on utilizing the multilingual knowledge of an LLM to inform its decision to abstain or answer when prompted. We develop a multilingual pipeline to calibrate the model's confidence and let it abstain when uncertain. We run several multilingual models through the pipeline to profile them across different languages. We find that the performance of the pipeline varies by model and language, but that in general they benefit from it. This is evidenced by the accuracy improvement of $71.2\%$ for Bengali over a baseline performance without the pipeline. Even a high-resource language like English sees a $15.5\%$ improvement. These results hint at possible further improvements.

大模型拒答机制多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。