让生成模型少出错,通过少量验证查询提升输出有效性。
Distribution Learning with Valid Outputs Beyond the Worst-Case
- 在数据分布属于模型类且最小化对数损失时,有效性的样本需求弱依赖于严格程度。
- 当有效区域属于VC类时,少量验证查询即可满足多数情况的有效性要求。
- 适合关注生成模型可靠性与验证机制效率的研究者。
生成模型有时会产生无效输出,如带有生成伪影的图像或不自然的声音。有效性约束的分布学习通过要求学习到的分布在其无效空间中的质量极小来解决此问题——这并非标准损失最小化能保证的。为此,该框架允许学习者使用“有效性查询”来检验单个样本是否有效。以往研究采用最坏情况分析,表明正确学习需要指数级有效性查询,并提出一种虽具广泛适用性但仅需多项式数量查询的非正规算法。本文首次尝试刻画有效性保障更易实现的场景:当数据分布位于模型类别中且对数损失被最小化时,确保有效性的样本需求对有效性要求的依赖极弱;此外,若有效性区域属于VC类,则有限数量的有效性查询通常已足够。
原文摘要 · Abstract (English)
Generative models at times produce "invalid" outputs, such as images with generation artifacts and unnatural sounds. Validity-constrained distribution learning attempts to address this problem by requiring that the learned distribution have a provably small fraction of its mass in invalid parts of space -- something which standard loss minimization does not always ensure. To this end, a learner in this model can guide the learning via "validity queries", which allow it to ascertain the validity of individual examples. Prior work on this problem takes a worst-case stance, showing that proper learning requires an exponential number of validity queries, and demonstrating an improper algorithm which -- while generating guarantees in a wide-range of settings -- makes an atypical polynomial number of validity queries. In this work, we take a first step towards characterizing regimes where guaranteeing validity is easier than in the worst-case. We show that when the data distribution lies in the model class and the log-loss is minimized, the number of samples required to ensure validity has a weak dependence on the validity requirement. Additionally, we show that when the validity region belongs to a VC-class, a limited number of validity queries are often sufficient.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。