首次量化现代艺术生成视频的不确定性结构,揭示多种子结果背后的语义差异。
Toward Uncertainty Quantification in Modern Art

- 提出无源盲与有参考感知的13种不确定性估计器,分析多种子生成结果的分布特征。
- 在1000个视频上实现拓扑分类准确率98%,异常值识别AUROC达1.00,远超传统标量方法。
- 适合关注生成艺术可信度、多模态解释力的研究者和创作者使用。
当同一幅现代艺术作品在不同随机种子下被文本到视频模型生成时,会产出明显不同的影片,每种结果对应一种解读。由于现代艺术本就具有意图模糊性,这种差异是有效信号而非噪声。然而现有不确定性量化(UQ)仅将多种子生成结果压缩为单一离散度数值,无法区分紧凑解释、主导读法加异常值、多重模式或整体不稳定性,也无法判断结果中是否仍包含忠实原作的呈现。本文首次研究现代艺术动画生成中的不确定性结构,提出可复用的盲源多种子不确定性识别协议:包含7种无源盲估计算法和6种参考感知算法;构建分布特征谱(鲁棒跨度、异常值影响、显式拓扑、多模态性、各向异性、留一种子影响、参考覆盖度);进行分布模型消融实验(vMF、Kent、ACG、Student t、核函数、混合模型);设计8个识别问题与作品级统计协议。构建首个数据集:250幅现代艺术作品描述,由Wan2.1 14B模型在4种子下生成(共1000个视频),艺术作品未参与训练。作为诊断工具,该协议在拓扑分类中达到平衡准确率0.98(随机基准0.25),异常值定位的AUROC达1.00,而传统标量方法仅为0.35;能可靠地将高不确定性作品分为含参考覆盖(n=97)与缺失参考覆盖(n=56)两类,仅需3个种子且跨编码器稳定有效。
原文摘要 · Abstract (English)
Asked to animate the same modern artwork under different random seeds, a text to video model returns visibly different films, one reading per seed. Because modern art is ambiguous by intent, this disagreement is signal, not noise. Yet prevailing uncertainty quantification (UQ) collapses a set of generations to a dispersion scalar that says how much the seeds differ but not how: it cannot tell a compact interpretation from a dominant reading plus an outlier, two competing modes, or diffuse instability, nor whether the set still contains a rendering faithful to the original. We present the first study of the structure of generative uncertainty for modern art animation, and a reusable protocol for identifying source blind multiseed uncertainty: a suite of seven source blind and six reference aware estimators; a distributional profile (robust spread, outlier influence, explicit topology, multimodality, anisotropy, leave one seed influence, reference coverage); a distribution model ablation (vMF, Kent, ACG, Student t, kernel, mixture); eight identification questions; and an artwork level statistical protocol. We build the first corpus: 250 modern artwork captions rendered by Wan2.1 14B under four seeds (1000 videos) across 4 encoders, artworks withheld from generation. As a diagnostic the protocol succeeds: it classifies seed set topology at balanced accuracy 0.98 (chance 0.25), isolates the outlier configuration at AUROC 1.00 where a scalar reaches only 0.35, and splits high uncertainty artworks into reference covering (n=97) and reference missing (n=56) diversity, reliably from three seeds and across encoders.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。