首个系统评估视频假新闻检测中感知、理解与推理能力的多模态基准
Perception, Understanding and Reasoning, A Multimodal Benchmark for Video Fake News Detection
- 构建10项任务,覆盖感知、理解、推理全流程
- 含36,240条人工标注问答,分15个维度评估
- 基于Qwen2.5VL-7B-Instruct微调,达当前最佳性能
多模态大语言模型(MLLMs)的兴起推动了视频假新闻检测(VFND)研究进展。现有基准多关注检测准确率,缺乏对检测全过程的细粒度评估。为此,我们提出面向过程的视频假新闻检测基准(POVFNDB),包含10项任务,系统评估MLLM在感知、理解与推理方面的能力。该基准涵盖36,240条人工标注的问答对,采用结构化或开放式格式,覆盖15个不同评估维度,刻画视频假新闻检测过程的多个层面。我们利用POVFNDB对专有及开源的MLLM进行综合评估,并通过自研的POVFND-CoT框架构建过程导向的思维链数据,微调Qwen2.5VL-7B-Instruct,建立强基线,在VFND任务上取得当前最优表现。
原文摘要 · Abstract (English)
The advent of multi-modal large language models (MLLMs) has greatly advanced research on video fake news detection (VFND) tasks. Existing benchmarks typically focus on the detection accuracy, while failing to provide fine-grained assessments for the entire detection process. To address these limitations, we introduce {POVFNDB (Process-oriented Video Fake News Detection Benchmark)}, a process-oriented benchmark comprising 10 tasks designed to systematically evaluate MLLMs' perception, understanding, and reasoning capabilities in VFND. This benchmark contains \textit{36,240} human-annotated question-answer (QA) in structured or open-ended formats, spanning 15 distinct evaluation dimensions that characterize different aspects of the video fake news detection process. Using POVFNDB, we conduct comprehensive evaluations on both proprietary and open-source MLLMs. Moreover, we establish a strong benchmark baseline by fine-tuning Qwen2.5VL-7B-Instruct on process-oriented chain-of-thought data constructed with our proposed POVFND-CoT framework, achieving state-of-the-art performance on VFND.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。