提出隐空间频域有效性机制,实现快速精准视频频谱编辑。
Latent-Frequency Validity: Fast Spectral Editing with Screened Video-VAE Transfer Operators
- 设计可验证的隐空间频域编辑路径,动态选择简单或混合操作。
- 99%编辑结果通过严格验证,423个操作中仅146个需复杂通道混合。
- 速度比传统滤波重编码快3倍,适用于多种视频生成模型。
在视频变分自编码器(video-VAE)的隐空间直接进行频谱编辑,可无需解码-滤波-重编码流程,控制噪声、闪烁、平滑度和频率内容。然而,视频VAE可能将像素空间的频带分布至不同隐层通道,导致隐空间编辑破坏编码-解码循环稳定性。本文提出隐空间频域有效性(Latent-Frequency Validity, LFV),学习特定于VAE的频域响应,并仅在提升重建质量且不加剧循环漂移时启用。LFV采用从每频段对角校准器(C1)到全通道混合(CM)的验证选择路径,使跨通道容量成为可控资源。在覆盖六种频谱类别的544个VAE-编辑组合中,共生成423个低成本操作:277个由C1处理,146个(占发出操作的34.5%)需通道混合。在主120单元径向扫描中,100个操作中有99个通过源视频组留出评估。在五个额外滤波族中,全部323个操作均通过留出评估。完全冻结的OpenVid适配操作(含验证选择系数)在20个CogVideoX与HunyuanVideo生成域单元中无需微调即通过测试。所选响应延迟接近直接隐空间滤波,约为像素级滤波-重编码的1/3。结果揭示了不同的VAE工作区间,包括强通道耦合的CogVideoX响应及Open-Sora高带宽稳定边界。
原文摘要 · Abstract (English)
Direct spectral editing in video-VAE latents can control noise, flicker, smoothness, and frequency content without a decode--filter--reencode pass. However, video VAEs may redistribute pixel-space frequency bands across latent channels, and latent edits can disrupt VAE round-trip dynamics. We introduce \emph{latent-frequency validity} (LFV), which learns a compact VAE-specific spectral response and deploys it only when it improves decoded-target fidelity without worsening round-trip drift. LFV follows a validation-selected path from a diagonal per-frequency calibrator (C1) to full channel mixing (CM), making cross-channel capacity a controllable per-edit resource. Across 544 VAE--edit cells spanning six spectral families, LFV emits 423 cheap operators: 277 are handled by C1, while 146 (34.5\% of emitted operators) require channel mixing. On the primary 120-cell radial sweep, 99/100 emitted operators pass source-video-grouped held-out evaluation. Across five additional filter families, all 323 emitted operators pass held-out evaluation. Fully frozen OpenVid-fitted operators, including the validation-selected path coefficient, pass all 20 tested CogVideoX and HunyuanVideo generated-domain cells without adaptation. The selected response matches direct latent-filter latency and is about $3\times$ faster than pixel filter--reencode. The resulting maps reveal distinct VAE regimes, including strongly channel-coupled CogVideoX responses and a sharp Open-Sora high-band stability frontier.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。