arXiv:2603.28103cs.DLcs.AI2026-03

用视觉语言模型自动识别意大利议会演讲内容与发言人

Transcription and Recognition of Italian Parliamentary Speeches Using Vision-Language Models

  • 结合视觉布局与文本内容,用大模型联合分析文档
  • 转录准确率和发言人标注效果显著优于传统方法
  • 适合处理历史扫描档案的语义分析任务

议会记录是计算分析的宝贵资源,但仅以扫描件保存时极具挑战性。现有意大利议会演讲转录多依赖传统OCR流程,导致转录错误频发且语义标注有限。本文提出一种基于视觉语言模型的端到端流水线,实现意大利议会演讲的自动转录、语义分段与实体链接。该流程首先使用专用OCR模型提取文本并保留阅读顺序,再通过大规模视觉语言模型联合推理视觉版式与文本内容,完成转录修正、元素分类与发言者识别。提取的发言者通过SPARQL查询与多策略模糊匹配,链接至众议院知识库。在公开基准上的评估显示,该方法在转录质量和发言人标注方面均有显著提升。

原文摘要 · Abstract (English)

Parliamentary proceedings represent a rich yet challenging resource for computational analysis, particularly when preserved only as scanned historical documents. Existing efforts to transcribe Italian parliamentary speeches have relied on traditional Optical Character Recognition pipelines, resulting in transcription errors and limited semantic annotation. In this paper, we propose a pipeline based on Vision-Language Models for the automatic transcription, semantic segmentation, and entity linking of Italian parliamentary speeches. The pipeline employs a specialised OCR model to extract text while preserving reading order, followed by a large-scale Vision-Language Model that performs transcription refinement, element classification, and speaker identification by jointly reasoning over visual layout and textual content. Extracted speakers are then linked to the Chamber of Deputies knowledge base through SPARQL queries and a multi-strategy fuzzy matching procedure. Evaluation against an established benchmark demonstrates substantial improvements both in transcription quality and speaker tagging.

视觉语言模型语音转录知识图谱议会数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。