arXiv:2607.28375cs.AIcs.MM2026-07

用超图建模视频真伪检测中的细粒度跨模态关系,提升识别精度。

HyperClaim: Fine-Grained Cross-Modal Hypergraph Reasoning for Video Misinformation Detection

论文配图:HyperClaim: Fine-Grained Cross-Modal Hypergraph Reasoning for Video Misinformation Detection
图 1 · 摘自论文原文
  • 构建稀疏异构超图,融合查询词、证据词与帧的高阶交互
  • 在三个数据集上准确率达83.7%至87.3%,超越现有方法
  • 无需外部工具或生成解释,适合细粒度视频真伪分析任务

视频虚假信息检测常采用全局多模态融合或自由形式的多模态推理。这两种范式往往忽略由查询短语、上下文文本和短时帧段之间耦合互动产生的局部真实性线索。由于此类互动本质为高阶关系,成对图模型无法充分捕捉多向跨模态依赖,而超图则具备合适表达能力。本文提出HyperClaim,一种用于样本级真实性分类的判别性时间超图框架。以标题或基准提供的配对文本作为类声明查询,HyperClaim在查询词、证据词与采样帧间构建稀疏异构超图;通过置信度感知过滤与源预算机制形成紧凑的文本-帧及短时程证据单元;执行自适应软关联推理并结合残差文本-视频校准;最后通过差异感知读出聚合文本、视觉及超边状态。不依赖生成推理过程或外部工具调用,保留了全局融合易丢失的细粒度跨模态与时间结构。在FactGuard时间协议下,其在FakeSV、FakeTT和FakeVV数据集上的准确率分别为83.7%、82.0%和87.3%,优于强判别性与推理主导基线。学习到的关联与注意力权重进一步揭示了词与帧级别的结构。

原文摘要 · Abstract (English)

Video misinformation detection is often approached through global multimodal fusion or free-form multimodal reasoning. Both paradigms can under-represent localized authenticity cues that arise from coupled interactions among query phrases, contextual text, and short temporal spans of frames. Because such interactions are inherently higher-order, pairwise graph formulations are insufficient to capture multi-way cross-modal dependencies, whereas hypergraphs offer a suitable representation for these relations. We propose HyperClaim, a discriminative temporal hypergraph framework for sample-level authenticity classification. Using the title or benchmark-provided paired text as a claim-like query, HyperClaim constructs a sparse heterogeneous hypergraph over query tokens, evidence tokens, and sampled frames; applies confidence-aware filtering and source budgeting to form compact text-frame and short-range temporal evidence units; performs adaptive soft-incidence reasoning with residual text-video calibration; and aggregates textual, visual, and hyperedge states through a discrepancy-aware readout. Without relying on generated rationales or external tool calls, HyperClaim preserves fine-grained cross-modal and temporal structure that global fusion tends to flatten. Under the FactGuard temporal protocol, it achieves 83.7%, 82.0%, and 87.3% accuracy on FakeSV, FakeTT, and FakeVV, respectively, outperforming strong discriminative and reasoning-centric baselines. Learned incidence and attention weights further reveal token- and frame-level structure.

视频检测超图多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。