arXiv:2607.20903cs.CVcs.HC2026-07

richer视觉表征不一定更符合人类判断,静态图像在某些场景下表现更优

Sidewalk Moments: Are Richer Representations Always More Human-Aligned? Evidence from City-Walk Videos

论文配图:Sidewalk Moments: Are Richer Representations Always More Human-Aligned? Evidence from City-Walk Videos
图 1 · 摘自论文原文
  • 用四种模态对比城市步行视频的参与度感知,发现动态内容用视频更好
  • 在判断高/低参与度时,平均图像(TAI)表现不输甚至优于完整视频
  • 适合做感知评分模型的数据压缩,尤其对构图主导的场景有优势

我们通过61段第一人称城市步行视频(共超过50,000个10秒片段),在四种模态下评估视觉表示的人类对齐程度:时空视频特征、时序平均图像(TAIs)、音频嵌入和基于文本的语义描述。斯皮尔曼相关分析显示,按时间丰富度排序,视频特征与人类判断最一致。但在二分类(高/低参与度)任务中,该规律失效:多数分类器和分位数阈值下,TAI表现与视频相当或更优。独立的两选一强制选择实验(亚马逊机械土耳其人平台)验证:参与者从TAI和完整视频中识别参与度的准确率相近,而文本表现差,音频接近随机。差距分析揭示功能分离:视频特征在动态活动场景占优,而TAI在以稳定空间结构为主的构图主导场景更贴近人类判断。结果挑战了‘更丰富表征更人类对齐’的假设,表明基于感知的时序压缩可成为全视频编码的合理替代方案。

原文摘要 · Abstract (English)

We examine whether richer visual representations yield more human-aligned measures of urban engagement, using 61 first-person city-walk videos from YouTube segmented into over 50,000 ten-second clips and represented across four modalities: spatiotemporal video features, temporally averaged images (TAIs), audio embeddings, and text-based semantic descriptions. Spearman correlation analysis reveals the expected ordering along the temporal-richness continuum, with video features showing the strongest continuous alignment. However, this ordering breaks down under binary classification of high- versus low-engagement moments (the paradigm most commonly used to train perceptual scoring models), where TAIs consistently match or outperform video across most classifiers and quantile thresholds. An independent two-alternative forced-choice study on Amazon Mechanical Turk confirms that this parity reflects human judgment: participants identified engaging moments with comparable accuracy from TAIs and full video clips, while text performed substantially worse and audio remained near chance. Gap analysis reveals a functional dissociation: video features are advantaged in activity-driven scenes with dynamic content, whereas TAIs better align with human judgments in composition-driven scenes dominated by stable spatial structure. These findings challenge the assumption that richer representations are inherently more human-aligned, and suggest that perceptually grounded temporal compression can be a principled alternative to full video encoding.

视频理解人类对齐表征学习数据压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。