arXiv:2510.02571cs.CVcs.AI2025-10被引 6

首次为视频生成模型提供不确定性量化方法,提升生成可靠性。

How Confident are Video Models? Empowering Video Models to Express their Uncertainty

  • 基于隐空间建模,分解生成不确定性的两类来源。
  • 提出S-QUBED方法,计算出与任务准确率负相关的可信度评分。
  • 构建首个视频模型不确定性评估数据集,支持基准测试。

生成式视频模型展现出强大的文本到视频能力,推动其在诸多实际应用中的广泛采用。然而,与大语言模型类似,视频生成模型易产生幻觉,即在事实错误的情况下仍生成看似合理的视频。尽管大语言模型的不确定性量化(UQ)已有深入研究,但视频模型尚无相应的UQ方法,引发重大安全顾虑。本文首次提出视频模型不确定性量化框架,包含:(i) 基于稳健秩相关估计的校准评估指标,无需严格建模假设;(ii) 一种黑箱视频模型不确定性量化方法(S-QUBED),通过隐空间建模,严谨分解预测不确定性为偶然性与认知性成分;(iii) 用于视频模型校准基准测试的UQ数据集。通过在隐空间中条件化生成任务,分离了由任务描述模糊引起的不确定性与知识不足导致的不确定性。在多个基准视频数据集上的实验表明,S-QUBED能输出与任务准确率负相关的校准总不确定性估计,并有效区分偶然性和认知性不确定性。

原文摘要 · Abstract (English)

Generative video models demonstrate impressive text-to-video capabilities, spurring widespread adoption in many real-world applications. However, like large language models (LLMs), video generation models tend to hallucinate, producing plausible videos even when they are factually wrong. Although uncertainty quantification (UQ) of LLMs has been extensively studied in prior work, no UQ method for video models exists, raising critical safety concerns. To our knowledge, this paper represents the first work towards quantifying the uncertainty of video models. We present a framework for uncertainty quantification of generative video models, consisting of: (i) a metric for evaluating the calibration of video models based on robust rank correlation estimation with no stringent modeling assumptions; (ii) a black-box UQ method for video models (termed S-QUBED), which leverages latent modeling to rigorously decompose predictive uncertainty into its aleatoric and epistemic components; and (iii) a UQ dataset to facilitate benchmarking calibration in video models. By conditioning the generation task in the latent space, we disentangle uncertainty arising due to vague task specifications from that arising from lack of knowledge. Through extensive experiments on benchmark video datasets, we demonstrate that S-QUBED computes calibrated total uncertainty estimates that are negatively correlated with the task accuracy and effectively computes the aleatoric and epistemic constituents.

视频生成不确定性生成模型校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。