arXiv:2602.09154cs.CVcs.AI2026-02中稿 · publication at the…

提出可追溯的命名实体提取框架,解决新闻视频字幕识别难题

A Hybrid Deterministic Framework for Named Entity Extraction in Broadcast News Video

  • 构建模块化确定性流程,逐阶段追踪信息提取过程
  • 检测器[email protected]达95.8%,提取精度79.9%、召回74.4%
  • 适合需透明溯源的新闻分析场景,避免生成模型幻觉

随着视频新闻内容激增,亟需透明可靠的屏幕信息提取方法。现有图形布局、字体风格和平台设计差异大,人工标注已不可行。本文提出一套自动检测并提取广播与社交媒体新闻视频中人名的完整框架。构建了包含多样化新闻图文的标注帧语料库,并设计可解释、模块化的确定性处理流程,确保每步操作可审计。在与生成式多模态方法对比中,该框架虽准确率(F1: 77.08%)略低于生成模型(F1: 84.18%),但具备全程可追溯性。底层检测器在[email protected]上达到95.8%,实现稳定图形元素定位。该流程避免幻觉,支持全流程溯源。用户调研显示59%受访者难以阅读快节奏播报中的字幕,凸显任务实际价值。研究为现代新闻媒体中的混合多模态信息提取提供了方法严谨且可解释的基准。

原文摘要 · Abstract (English)

The growing volume of video-based news content has heightened the need for transparent and reliable methods to extract on-screen information. Yet the variability of graphical layouts, typographic conventions, and platform-specific design patterns renders manual indexing impractical. This work presents a comprehensive framework for automatically detecting and extracting personal names from broadcast and social-media-native news videos. It introduces a curated and balanced corpus of annotated frames capturing the diversity of contemporary news graphics and proposes an interpretable, modular extraction pipeline designed to operate under deterministic and auditable conditions. The pipeline is evaluated against a contrasting class of generative multimodal methods, revealing a clear trade-off between deterministic auditability and stochastic inference. The underlying detector achieves 95.8% [email protected], demonstrating operationally robust performance for graphical element localisation. While generative systems achieve marginally higher raw accuracy (F1: 84.18% vs 77.08%), they lack the transparent data lineage required for journalistic and analytical contexts. The proposed pipeline delivers balanced precision (79.9%) and recall (74.4%), avoids hallucination, and provides full traceability across each processing stage. Complementary user findings indicate that 59% of respondents report difficulty reading on-screen names in fast-paced broadcasts, underscoring the practical relevance of the task. The results establish a methodologically rigorous and interpretable baseline for hybrid multimodal information extraction in modern news media.

命名实体识别视频理解可解释性新闻媒体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。