arXiv:2512.13281cs.CV2025-12被引 2

新基准测试AI生成的ASMR视频能否骗过人类和视觉模型。

VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans?

  • 基于ASMR感官体验设计细粒度音画感知评测
  • 顶尖视觉模型仍难识别9种生成方式的假视频
  • 适合评估生成模型真实感与感知模型检测能力

随着AI生成视频日益逼真,现有基准主要关注语义一致性和基础物理合理性,难以有效区分。为此,我们提出VideoASMR-Bench,一个基于自主感官愉悦反应(ASMR)视频的评测基准,强调细粒度音视频感知与沉浸体验。该基准包含1,500个来自社交媒体的真实高质量ASMR视频,以及由九种视频生成模型(VGMs)生成的2,235个合成视频。我们还开源可扩展的提示与参考图像套件,支持未来模型动态接入。此外,构建了视觉理解模型(VLM)与视频生成模型(VGM)之间的自动对抗评估框架:VGM试图生成以欺骗VLM的逼真假视频,而VLM则尽力识别。评测显示,即使顶尖的VLM如Gemini-3-Pro也难以可靠检测出这些假视频;当前前沿生成模型已能产出令VLM难以分辨的逼真ASMR视频,但人类仍可较容易识别。

原文摘要 · Abstract (English)

With AI-generated videos increasingly indistinguishable from reality, current benchmarks primarily focus on broad semantic alignment and basic physical consistency, offering limited discriminative power for evaluating them. To address this, we introduce VideoASMR-Bench, a benchmark based on Autonomous Sensory Meridian Response (ASMR) videos that emphasizes fine-grained audio-visual perception and sensory immersion. This benchmark aims to answer two key questions: (i) Are today's video understanding models (VLMs) sensitive enough to detect AI-generated ASMR videos by recognizing minor visual, physical, or auditory artifacts? (ii) Can today's video generation models (VGMs) produce convincing ASMR videos with immersive experiences? This benchmark comprises a diverse set of 1,500 high-quality real ASMR videos curated from social media, alongside 2,235 synthetic counterparts generated by nine VGMs. Additionally, we open-source an extensible suite of prompts and reference images, enabling the benchmark to scale dynamically with future video models. Moreover, we design an automatic understanding-generation evaluation framework between VGMs and VLMs, where VGMs aim to produce realistic fake videos to fool the VLMs, while the VLMs seek to detect them, forming an adversarial game between the two parties. Our evaluation on VideoASMR-Bench reveals that even state-of-the-art VLMs, such as Gemini-3-Pro, fail to reliably detect AI-generated ASMR videos. Meanwhile, current frontier video generation models can produce ASMR videos that are difficult for VLMs to distinguish from real ones, while humans can still identify them relatively easily.

视频生成视觉模型ASMR评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。