arXiv:2507.04182cs.IRcs.CL2025-07

用AI生成插图帮用户快速浏览海量语音档案。

Navigating Speech Recording Collections with AI-Generated Illustrations

  • 用语言和多模态生成模型构建交互式思维导图
  • 基于TED-LIUM3数据集,组织2000+个演讲音频
  • 适合需要高效探索语音资料的研究者或内容创作者

尽管可获取的语音内容持续增长,从语音录音中提取信息仍具挑战性。除了改进传统信息检索方法(如语音搜索、关键词识别),还需探索新型导航与搜索方式。本文提出一种利用近期语言及多模态生成模型的语音档案导航新方法。我们开发了一个网页应用,通过交互式思维导图和图像生成工具,将数据结构化呈现。系统基于包含超过2000个演讲录音与音频文件的TED-LIUM3数据集实现。初步用户测试采用系统可用性量表(SUS)评估,表明该应用具有简化大规模语音集合探索的潜力。

原文摘要 · Abstract (English)

Although the amount of available spoken content is steadily increasing, extracting information and knowledge from speech recordings remains challenging. Beyond enhancing traditional information retrieval methods such as speech search and keyword spotting, novel approaches for navigating and searching spoken content need to be explored and developed. In this paper, we propose a novel navigational method for speech archives that leverages recent advances in language and multimodal generative models. We demonstrate our approach with a Web application that organizes data into a structured format using interactive mind maps and image generation tools. The system is implemented using the TED-LIUM~3 dataset, which comprises over 2,000 speech transcripts and audio files of TED Talks. Initial user tests using a System Usability Scale (SUS) questionnaire indicate the application's potential to simplify the exploration of large speech collections.

语音导航AI绘图信息检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。