用大模型模拟用户观看历史,生成个性化视频亮点数据集
HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting
- 用大模型生成真实用户观看习惯,构建个性化视频数据
- 涵盖20,400个视频,170类语义,2,040条带评分的历史记录
- 适合做个性化推荐、视频摘要和用户行为建模的研究者
视频内容的爆炸式增长使得个性化视频亮点提取成为关键任务,因用户偏好高度多样且复杂。现有视频数据集常缺乏个性化,仅依赖孤立视频或简单文本查询,无法捕捉用户行为的深层特征。本文提出HIPPO-Video,一个基于大语言模型的用户模拟器生成的真实观看历史数据集,包含2,040个(观看历史,显著性得分)对,覆盖20,400个视频,涉及170个语义类别。为验证该数据集,我们提出HiPHer方法,利用个性化观看历史预测条件化的片段显著性得分。大量实验表明,该方法优于现有通用及查询驱动方法,展现出在真实场景中实现高度用户中心化视频亮点提取的巨大潜力。
原文摘要 · Abstract (English)
The exponential growth of video content has made personalized video highlighting an essential task, as user preferences are highly variable and complex. Existing video datasets, however, often lack personalization, relying on isolated videos or simple text queries that fail to capture the intricacies of user behavior. In this work, we introduce HIPPO-Video, a novel dataset for personalized video highlighting, created using an LLM-based user simulator to generate realistic watch histories reflecting diverse user preferences. The dataset includes 2,040 (watch history, saliency score) pairs, covering 20,400 videos across 170 semantic categories. To validate our dataset, we propose HiPHer, a method that leverages these personalized watch histories to predict preference-conditioned segment-wise saliency scores. Through extensive experiments, we demonstrate that our method outperforms existing generic and query-based approaches, showcasing its potential for highly user-centric video highlighting in real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。