提出MG-RWKV模型,高效定位视频伪造片段。
MG-RWKV: Multi-Grained Context-Aware RWKV for Temporal Forgery Localization

- 用双向RWKV捕捉时序上下文,复杂度仅O(T)。
- 多粒度专家路由自适应选择伪造时长,提升可解释性。
- 跨粒度一致性机制减少真实区域误报,适合伪造检测场景。
随着AI生成内容(AIGC)的发展,音视频内容的真实性面临严峻挑战。时间伪造定位(TFL)旨在精准识别未修剪序列中的篡改片段。现有方法受限于CNN的局部感受野或Transformer的二次复杂度,而新兴线性模型难以兼顾全局真实上下文压缩与局部突发伪造感知。为此,我们提出MG-RWKV,一种多粒度框架,利用RWKV的数据依赖状态演化实现全序列高效处理,复杂度为O(T)。核心创新包括:(1) 双向RWKV架构,在无二次开销下捕捉双向时序上下文;(2) 多粒度专家混合(MG-MoE),基于伪造持续时间动态路由显式时序感受野,显著提升决策可解释性;(3) 跨粒度一致性(CGC),通过分层尺度配对和空间边界感知加权对齐相邻特征金字塔层级,有效降低真实区域的误报率。在Lav-DF、TVIL和Psynd数据集上的大量实验表明,MG-RWKV以低计算成本达到领先性能。
原文摘要 · Abstract (English)
Driven by Artificial Intelligence-Generated Content (AIGC), the authenticity of audio-visual content is facing severe challenges. Temporal Forgery Localization (TFL) aims to precisely identify manipulated segments within untrimmed sequences. However, existing methods are limited by CNNs' local receptive fields or Transformers' quadratic complexity, while emerging linear models often struggle to balance global authentic context compression with local abrupt forgery perception. To address this, we propose MG-RWKV, a multi-granularity framework that leverages the data-dependent state evolution of RWKV to achieve efficient full-sequence processing with O(T) complexity. Our framework features three core innovations: (1) a Bidirectional RWKV architecture that captures bidirectional temporal contexts without quadratic overhead; (2) a Multi-Granularity Mixture of Experts (MG-MoE) that performs dynamic routing over explicit temporal receptive fields, adaptively selecting granularities based on forgery duration to significantly enhance decision interpretability; and (3) Cross-Granularity Consistency (CGC), which aligns adjacent feature pyramid levels through hierarchical scale-wise pairing and spatial boundary-aware weighting, effectively reducing false positives in authentic regions. Extensive experiments on Lav-DF, TVIL, and Psynd datasets demonstrate that MG-RWKV achieves state-of-the-art performance with low computational cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。