理论上证明生成模型无法完全避免幻觉,必须引入先验知识才能学习不产生幻觉的模型。
No Free Lunch: Fundamental Limits of Learning Non-Hallucinating Generative Models
- 从学习理论出发,证明仅靠训练数据无法实现无幻觉生成。
- 即使数据全真、假设空间极小,也无法避免幻觉,统计上不可能。
- 需引入与事实一致的归纳偏置,适合研究生成模型可靠性的人读。
生成模型在多个领域展现出生成高质量输出的能力,但持续面临‘幻觉’问题——即模型生成看似合理却无效的内容。尽管已有经验性策略尝试缓解此问题,但缺乏严格的理论理解。本文从学习理论角度分析非幻觉生成模型的可学习性。结果表明,仅依赖训练数据时,即便假设空间大小为2且训练集完全真实,非幻觉学习在统计上仍不可行。为克服此限制,我们证明将与真实事实对齐的归纳偏置引入学习过程至关重要。通过将事实集合限制在有限VC维的概念类中,提出系统性方法,并在多种学习范式下验证其有效性。尽管结论以概念为主,但为解决生成模型幻觉问题提供了首个原理性框架。
原文摘要 · Abstract (English)
Generative models have shown impressive capabilities in synthesizing high-quality outputs across various domains. However, a persistent challenge is the occurrence of "hallucinations", where the model produces outputs that are plausible but invalid. While empirical strategies have been explored to mitigate this issue, a rigorous theoretical understanding remains elusive. In this paper, we develop a theoretical framework to analyze the learnability of non-hallucinating generative models from a learning-theoretic perspective. Our results reveal that non-hallucinating learning is statistically impossible when relying solely on the training dataset, even for a hypothesis class of size two and when the entire training set is truthful. To overcome these limitations, we show that incorporating inductive biases aligned with the actual facts into the learning process is essential. We provide a systematic approach to achieve this by restricting the facts set to a concept class of finite VC-dimension and demonstrate its effectiveness under various learning paradigms. Although our findings are primarily conceptual, they represent a first step towards a principled approach to addressing hallucinations in learning generative models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。