arXiv:2601.00299cs.SDcs.MM2026-01中稿 · ISMIR 2025 Late-Br…

自动提取台湾歌仔戏电视片的字幕并修正,提升研究数据效率。

Timed text extraction from Taiwanese Kua-á-hì TV series

  • 结合OCR与语音音乐活动检测,分两步定位唱段。
  • 从老旧视频中精准提取唱段与歌词,支持后续音乐分析。
  • 适合对地方戏曲数字化、音乐信息检索感兴趣的学者。

台湾歌仔戏作为重要的地方戏剧传统,曾由如尹丽华等先驱广泛改编为电视节目。这些珍贵影像虽具学术价值,但画质差且需大量人工处理。为此,我们开发了一个交互式实时OCR校正系统,并提出一种两阶段方法:先通过OCR驱动分割,再结合语音与音乐活动检测(SMAD),高精度识别档案剧集中的演唱片段。最终生成包含演唱段落及其对应歌词的数据集,可支持歌词识别、曲调检索等音乐信息检索任务。代码已公开于 https://github.com/z-huang/ocr-subtitle-editor。

原文摘要 · Abstract (English)

Taiwanese opera (Kua-á-hì), a major form of local theatrical tradition, underwent extensive television adaptation notably by pioneers like Iûnn Lē-hua. These videos, while potentially valuable for in-depth studies of Taiwanese opera, often have low quality and require substantial manual effort during data preparation. To streamline this process, we developed an interactive system for real-time OCR correction and a two-step approach integrating OCR-driven segmentation with Speech and Music Activity Detection (SMAD) to efficiently identify vocal segments from archival episodes with high precision. The resulting dataset, consisting of vocal segments and corresponding lyrics, can potentially supports various MIR tasks such as lyrics identification and tune retrieval. Code is available at https://github.com/z-huang/ocr-subtitle-editor .

音频分析字幕提取戏曲数字化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。