arXiv:2604.24947cs.CV2026-04

构建首个大规模主观视频人物区域裁剪数据集,提升移动端视频适配质量。

Subjective Portrait Region Cropping in Landscape Videos with Temporal Annotation Smoothing

论文配图:Subjective Portrait Region Cropping in Landscape Videos with Temporal Annotation Smoothing
图 1 · 摘自论文原文
  • 基于90人标注1800段视频,建立主观人物区域裁剪数据库
  • 引入时序平滑滤波器优化标注,减少帧间波动
  • 可直接用于视频重构与视觉显著性研究,适合移动端视频处理方向

随着移动设备显示分辨率和朝向多样化的普及,视频适配不同宽高比面临挑战。静态裁剪或加边常损害画质,拉伸变形则可能扭曲视频原意。本文提出一种时序化裁剪策略,在保留关键内容的同时最小化失真。为解决数据匮乏问题,我们构建了LIVE-YT VC数据集,包含1800段来自YouTube-UGC和LSVQ的数据,由90名参与者主观标注。进一步推出LIVE-YT VC++版本,采用新型帧内时序滤波器平滑标注结果。通过SmartVidCrop算法和先进视频定位模型验证其有效性,证明该数据集可作为未来研究基准。标签与视频显著性具有相似性,因此额外分析了二者关联。同时,将前沿视频定位模型迁移至本任务并微调,以支持适配。项目计划开源,助力社区发展。

原文摘要 · Abstract (English)

With the rise of mobile video consumption on diverse handheld display resolutions and orientation modes, altering videos to aspect ratios poses challenges. Static cropping and border padding often compromises visual quality, while warping may distort a video's intended meaning. Here we advocate for a more effective approach: cropping significant regions within video frames in a temporal manner, while minimizing distortion and preserving essential content. One barrier to solving this problem is the lack of sufficiently large-scale database devoted to informing these tasks. Towards filling this gap, we introduce the LIVE-YouTube Video Cropping (LIVE-YT VC) database, featuring 1800 videos, annotated by 90 human subjects. Using videos sourced from the YouTube-UGC and LSVQ Databases, this new resource is the largest publicly-available subjective video portrait region cropping database. We also introduce a post-processed version of the database, called LIVE-YT VC++, whereby a novel intra-frame temporal filter was deployed to smooth subjective annotations within each video. We demonstrate the usefulness of this new data resource using the SmartVidCrop algorithm and state-of-the-art video grounding models, in hopes of establishing our subjective dataset as a benchmark for future research. Our contributions offer a resource for advancing video aspect ratio transformation models towards ensuring that reshaped mobile-friendly video content retains its quality and meaning. Since our labels bear resemblances to video saliency annotations, we also conducted an additional analysis to explore the similarity between our labels and video saliency predictions. Finally, we repurposed state-of-the-art video grounding models for aspect ratio change tasks, and fine-tuned them on our dataset. As a service to the research community, we plan to open source the project.

视频裁剪主观标注移动端适配数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。