arXiv:2510.27179cs.CVcs.CR2025-10中稿 · 26th Privacy Enhan…

通过模糊字幕轮廓识别视频,可远程追踪用户观看内容。

SilhouetteTell: Practical Video Identification Leveraging Blurred Recordings of Video Subtitles

  • 利用字幕轮廓的空间与时间特征构建指纹
  • 可在40米外准确识别在线和离线视频
  • 适合隐私安全研究者关注视频泄露风险

视频识别攻击对隐私构成重大威胁,可能暴露用户观看的视频内容,进而揭示其兴趣爱好、宗教信仰、政治倾向、性取向及健康状况。观看历史还可用于用户画像或广告推送,导致网络欺凌、歧视或勒索。现有视频推断技术多依赖流媒体产生的网络流量分析。本文观察到字幕内容决定其在屏幕上呈现的轮廓,且连续字幕间存在时间差。提出SilhouetteTell,将字幕轮廓的时空特征融合为时序指纹,挖掘录制字幕轮廓与原始字幕文件间的时空关联。该方法可推断在线与离线视频。在市售智能手机上进行的全面实验验证了其高效性,可在高达40米的距离下准确识别视频标题与片段。

原文摘要 · Abstract (English)

Video identification attacks pose a significant privacy threat that can reveal videos that victims watch, which may disclose their hobbies, religious beliefs, political leanings, sexual orientation, and health status. Also, video watching history can be used for user profiling or advertising and may result in cyberbullying, discrimination, or blackmail. Existing extensive video inference techniques usually depend on analyzing network traffic generated by streaming online videos. In this work, we observe that the content of a subtitle determines its silhouette displayed on the screen, and identifying each subtitle silhouette also derives the temporal difference between two consecutive subtitles. We then propose SilhouetteTell, a novel video identification attack that combines the spatial and time domain information into a spatiotemporal feature of subtitle silhouettes. SilhouetteTell explores the spatiotemporal correlation between recorded subtitle silhouettes of a video and its subtitle file. It can infer both online and offline videos. Comprehensive experiments on off-the-shelf smartphones confirm the high efficacy of SilhouetteTell for inferring video titles and clips under various settings, including from a distance of up to 40 meters.

视频识别隐私安全字幕分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。