arXiv:2608.25344cs.CVcs.AI2026-08

用粗粒度标注训练模型,自动找出行车风险的细节证据。

CoRE: Weakly Supervised Coarse-to-Fine Risk Evidence Learning in Driving Videos

论文配图:CoRE: Weakly Supervised Coarse-to-Fine Risk Evidence Learning in Driving Videos
图 1 · 摘自论文原文
  • 从视频级标签出发,通过干预测试定位风险发生的时间与实体
  • 在无细粒度标注情况下实现精准时间定位与异常检测
  • 适合做自动驾驶风险感知和弱监督视频分析的研究者

驾驶中的感知风险随时间演变,可能由特定场景元素支撑,但通常仅有粗粒度视频级标注。传统方法需耗时的人工标注来获取时间与实体级支持证据。本文提出CoRE框架,利用弱监督的粗到细学习机制,仅凭视频级标签即可挖掘细粒度支持信息。该方法先训练一个视频级预测器并冻结,再对候选时间区域或实体轨迹施加结构化干预,测量其对整体预测的影响,生成分级的预测效应目标。这些目标被蒸馏至学生模型,使其直接从原始视频中预测时间与实体支持,无需推理时干预。在三个互补设置中验证:RISEE基于主观片段级判断评估风险支持;DoTA提供独立的时间事件标注用于交通异常定位;UCF-Crime检验其在非驾驶异常检测任务上的泛化能力。结果表明,CoRE在无细粒度标注下仍能学习到有意义的细粒度支持信息,在DoTA上实现强时间定位,在UCF-Crime上表现竞争力。证明粗粒度预测可为恢复其支持证据提供有效监督,无需对应细粒度标签。

原文摘要 · Abstract (English)

Perceived risk in driving evolves over time and may be supported by specific scene entities, yet supervision is typically limited to coarse video-level judgments. Learning \emph{when} supporting evidence emerges and \emph{which entities} support a risk predictor would ordinarily require costly temporal- and entity-level annotations. We introduce \textbf{CoRE}, a weakly supervised coarse-to-fine framework that learns fine-grained prediction support from coarse video supervision. CoRE first trains a video-level predictor and then freezes it. Structured interventions over candidate temporal regions or entity tracks measure how each candidate changes the coarse prediction, producing graded prediction-effect targets. These targets are distilled into a student that directly predicts temporal and entity support from the original video, without requiring interventions at inference. We evaluate this learning principle across three complementary settings: RISEE tests perceived-risk support from subjective clip-level judgments without temporal or entity-level risk annotations; DoTA provides independent temporal event annotations for evaluating weakly supervised traffic-anomaly localization; and UCF-Crime tests whether the same coarse-to-fine mechanism extends to a standard non-driving anomaly-detection benchmark. Across these settings, CoRE learns informative fine-grained support from coarse supervision, with strong temporal localization on DoTA and competitive performance on UCF-Crime. These results show that coarse video predictions can provide useful supervision for recovering the fine-grained evidence supporting them, without requiring corresponding fine-grained labels.

弱监督风险感知视频分析异常检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。