开源838段真实环境口型视频,支持高质量视频压缩与增强模型评测。
A Camera-Native Talking-Head Video Dataset for Various Computer Vision Tasks
- 采集799人443种摄像头场景的原生信号视频,无损保存
- 含120段分层测试集,支持压缩、超分等多任务评测
- 规模达现有最大数据集5倍,适配实时通信研究
口型视频是实时通信中的主流内容,但公开可用的数据集稀缺且信号质量有限。本文开源一个相机原生的口型视频数据集,包含838段约210分钟、每段15秒的视频,来自799名参与者在443个摄像头类别下的自然环境录制。所有视频采用FFV1无损编码,保留原始信号——未压缩占比24.7%,或使用MJPEG编码占比75.3%,无额外有损处理。每段视频附带平均意见分(MOS)及十项感知质量标签,联合解释64.4%的MOS方差。从中筛选出120段分层测试集,涵盖原始、背景虚化、背景替换三种内容条件。在四个数据集和四种编码器(H.264、H.265、H.266、AV1)上评估压缩效率,结果显示相对于H.264,H.266最高可节省71.3%的比特率,且编码器×数据集(η²=0.112)、编码器×内容条件(η²=0.149)存在显著交互作用,表明内容类型与背景处理影响压缩效率。初步超分辨率实验验证,该数据集显著影响绝对性能但保持模型排名一致,证明其适用于超分等任务。本数据集规模为现有最大同类数据集(160段)的5倍,且完整保留相机原生信号,为视频压缩、超分辨率、质量评估与增强模型提供基准资源。
原文摘要 · Abstract (English)
Talking-head videos constitute a predominant content type in real-time communication, yet publicly available datasets for video processing research in this domain remain scarce and limited in signal fidelity. In this paper, we open-source a camera-native dataset of 838 talking-head recordings (approximately 210 minutes), each 15s in duration, captured from 799 participants across 443 camera-code categories in their natural environments. All recordings are stored using the FFV1 lossless codec, preserving the camera-native signal---uncompressed (24.7%) or MJPEG-encoded (75.3%)---without additional lossy processing. Each recording is annotated with a Mean Opinion Score (MOS) and ten perceptual quality tokens that jointly explain 64.4% of the MOS variance. From this corpus, we curate a stratified benchmarking subset of 120 clips in three content conditions: original, background blur, and background replacement. Codec efficiency evaluation across four datasets and four codecs, namely H.264, H.265, H.266, and AV1, yields VMAF BD-rate savings up to $-71.3%$ (H.266) relative to H.264, with significant encoder$\times$dataset ($η_p^2 = .112$) and encoder$\times$content condition ($η_p^2 = .149$) interactions, demonstrating that both content type and background processing affect compression efficiency. A preliminary super-resolution evaluation with four SR models confirms that the dataset significantly affects absolute performance while preserving model rankings, demonstrating applicability beyond codec benchmarking. The dataset offers 5$\times$ the scale of the largest prior talking-head webcam dataset (838 vs. 160 clips) and preserves the camera-native signal without additional lossy compression, establishing a resource for benchmarking video compression, super-resolution, quality assessment, and enhancement models in real-time communication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。