针对直播中人与物交互检测的物体偏见问题,提出原型嵌入优化方法。
Prototype Embedding Optimization for Human-Object Interaction Detection in Livestreaming
- 通过原型嵌入优化缓解物体主导带来的交互识别偏差。
- 在VidHOI和BJUT-HOI数据集上,稀有类别准确率分别提升至26.20%和30.37%。
- 适用于直播内容理解、行为监管等需精准识别交互场景的应用。
直播中主播与物品的互动对内容理解和监管至关重要。尽管通用视频任务中人-物交互(HOI)检测已有进展,但在直播场景下,现有方法往往过度关注物体而忽视其与主播的交互,导致物体偏见。为此,本文提出原型嵌入优化的人-物交互检测方法(PeO-HOI)。首先利用目标检测与跟踪技术预处理直播流,提取人-物对特征;随后采用原型嵌入优化缓解物体偏见影响;最后建模人-物对之间的时空上下文,通过预测头输出交互结果。实验表明,该方法在公开数据集VidHOI上达到37.19%@full、51.42%@non-rare、26.20%@rare,在自建数据集BJUT-HOI上达到45.13%@full、62.78%@non-rare、30.37%@rare,显著提升了直播场景下的HOI检测性能。
原文摘要 · Abstract (English)
Livestreaming often involves interactions between streamers and objects, which is critical for understanding and regulating web content. While human-object interaction (HOI) detection has made some progress in general-purpose video downstream tasks, when applied to recognize the interaction behaviors between a streamer and different objects in livestreaming, it tends to focuses too much on the objects and neglects their interactions with the streamer, which leads to object bias. To solve this issue, we propose a prototype embedding optimization for human-object interaction detection (PeO-HOI). First, the livestreaming is preprocessed using object detection and tracking techniques to extract features of the human-object (HO) pairs. Then, prototype embedding optimization is adopted to mitigate the effect of object bias on HOI. Finally, after modelling the spatio-temporal context between HO pairs, the HOI detection results are obtained by the prediction head. The experimental results show that the detection accuracy of the proposed PeO-HOI method has detection accuracies of 37.19%@full, 51.42%@non-rare, 26.20%@rare on the publicly available dataset VidHOI, 45.13%@full, 62.78%@non-rare and 30.37%@rare on the self-built dataset BJUT-HOI, which effectively improves the HOI detection performance in livestreaming.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。