arXiv:2506.14067cs.LG2025-06被引 2

让AI在不确定时选择不回答,还能在用户只给点赞或点踩的情况下学得更好。

Online Conformal Abstention for Factuality Control Under Adversarial Bandit Feedback

  • 通过反馈解封技术从有限反馈中挖掘学习信号。
  • 在非平稳对抗环境下,错误率风险控制达到O(√T)。
  • 适合需要高可靠性的大模型问答系统使用。

随着交互式生成系统在现实应用中日益普及,其生成不可靠或虚假回答的倾向引发严重担忧。合规弃答通过仅在有把握时作答来缓解此风险。然而,实际部署通常仅提供部分用户反馈(如点赞/点踩),且运行环境常为非平稳或对抗性环境,现有有效学习方法匮乏。为此,我们提出ExAUL,一种针对对抗性和部分反馈的在线合规弃答学习框架。技术上,我们引入(i)新颖的转换引理,将任意贝叶斯算法的遗憾转化为错误发现率(FDR)界;(ii)反馈解封策略,利用合规弃答结构从部分反馈中提取额外学习信号。我们证明ExAUL实现遗憾界O(√(T ln|H|)),进而带来O(√T)的FDR风险控制,即便仅接收部分反馈,仍可媲美全信息设置的可控性。尽管适用于通用生成任务,我们在多样非平稳与对抗性场景下的问答任务中验证了ExAUL在保障大语言模型(LLM)可靠性方面的有效性。结果表明,ExAUL能稳健控制FDR,同时保持较高的回答覆盖率。

原文摘要 · Abstract (English)

As interactive generative systems are increasingly deployed in real-world applications, their tendency to generate unreliable or false responses raises serious concerns. Conformal abstention mitigates this risk by ensuring that the system answers only when confident. However, real-world deployments typically provide only partial user feedback (e.g., thumbs up/down) on the selected response and often operate in non-stationary or adversarial environments, for which effective learning methods are largely missing. To bridge this gap, we propose ExAUL, a novel online learning framework for conformal abstention with adversarial and partial feedback. Technically, we introduce (i) a novel conversion lemma}that translates the regret of any bandit algorithm into an FDR bound, and (ii) feedback unlocking, a strategy that exploits the structure of conformal abstention to extract additional learning signals from partial feedback. We prove that ExAUL achieves a regret bound of $O(\sqrt{T \ln |{H}|})$, which translates into an ${O}(\sqrt{T})$ bound on FDR risk control, matching the controllability of full-information settings despite receiving only partial feedback. While applicable to general generative tasks, we demonstrate the efficacy of ExAUL for ensuring the reliability of Large Language Models (LLMs) through empirical validation on question-answering tasks across diverse non-stationary and adversarial settings. Our results demonstrate that ExAUL robustly controls the FDR while maintaining competitive answering coverage.

大模型可靠性在线学习对抗反馈合规弃答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。