arXiv:2505.08439cs.CL2025-05被引 2

构建意大利最高法院判例主题建模数据集,解决法律文本分析缺乏公开数据难题。

A document processing pipeline for the construction of a dataset for topic modeling based on the judgments of the Italian Supreme Court

  • 融合布局分析、文字识别与匿名化,自动化处理法院判决文书。
  • 生成数据集在主题多样性与连贯性上分别达0.6198和0.6638,优于纯OCR方法。
  • 适用于法律信息挖掘、大模型辅助司法分析的研究者与开发者。

意大利法律研究中的主题建模因缺乏公开数据集而受限,难以深入分析最高法院判例中的法律议题。为此,我们开发了一套文档处理流水线,生成适合主题建模的匿名化数据集。该流程集成文档布局分析(YOLOv8x)、光学字符识别(OCR)及文本匿名化模块。其中,布局分析模块在mAP@50达到0.964,mAP@50-95为0.800;OCR检测器的mAP@50-95为0.9022;文字识别器(TrOCR)字符错误率为0.0047,词错误率为0.0248。相较纯OCR方法,本数据集在主题建模中取得0.6198的多样性分数和0.6638的连贯性分数。采用BERTopic提取主题,并用大语言模型生成标签与摘要,经领域专家评估,Claude Sonnet 3.7在标签生成上的BERTScore F1为0.8119,在摘要生成上为0.9130。

原文摘要 · Abstract (English)

Topic modeling in Italian legal research is hindered by the lack of public datasets, limiting the analysis of legal themes in Supreme Court judgments. To address this, we developed a document processing pipeline that produces an anonymized dataset optimized for topic modeling. The pipeline integrates document layout analysis (YOLOv8x), optical character recognition, and text anonymization. The DLA module achieved a mAP@50 of 0.964 and a mAP@50-95 of 0.800. The OCR detector reached a mAP@50-95 of 0.9022, and the text recognizer (TrOCR) obtained a character error rate of 0.0047 and a word error rate of 0.0248. Compared to OCR-only methods, our dataset improved topic modeling with a diversity score of 0.6198 and a coherence score of 0.6638. We applied BERTopic to extract topics and used large language models to generate labels and summaries. Outputs were evaluated against domain expert interpretations. Claude Sonnet 3.7 achieved a BERTScore F1 of 0.8119 for labeling and 0.9130 for summarization.

法律文本主题建模数据集构建大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。