用图表解析法律文书,让外行也能看懂判决核心内容。
LegalViz: Legal Text Visualization by Text To Diagram Generation
- 用DOT语言生成法律文本对应的可视化图谱,提取实体与关系
- 构建跨23语言的7010对法律文书与图表数据集,支持多语言理解
- 提出结合图结构、语义相似性与法律内容的新评估指标
法律文书如判决书和法院命令需高度专业法律知识才能理解。为向非专业人士揭示专家知识,我们探索通过易懂图示可视化法律文本,并提出新数据集LegalViz,涵盖23种语言和7,010个法律文书与可视化配对,采用Graphviz的DOT图描述语言。LegalViz可从复杂法律文本中快速识别法律主体、交易行为、法律依据及陈述内容,呈现关键信息。此外,我们设计了融合图结构、文本相似性和法律内容的新评估指标。在小样本和微调大模型生成法律图谱上进行实证研究,基于这些指标在23种语言中评估模型表现,结果表明使用LegalViz训练的模型优于现有GPT类模型,验证了数据集的有效性。
原文摘要 · Abstract (English)
Legal documents including judgments and court orders require highly sophisticated legal knowledge for understanding. To disclose expert knowledge for non-experts, we explore the problem of visualizing legal texts with easy-to-understand diagrams and propose a novel dataset of LegalViz with 23 languages and 7,010 cases of legal document and visualization pairs, using the DOT graph description language of Graphviz. LegalViz provides a simple diagram from a complicated legal corpus identifying legal entities, transactions, legal sources, and statements at a glance, that are essential in each judgment. In addition, we provide new evaluation metrics for the legal diagram visualization by considering graph structures, textual similarities, and legal contents. We conducted empirical studies on few-shot and finetuning large language models for generating legal diagrams and evaluated them with these metrics, including legal content-based evaluation within 23 languages. Models trained with LegalViz outperform existing models including GPTs, confirming the effectiveness of our dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。