用多人协作框架解决中日法律术语跨语言映射难题
Building from Scratch: A Multi-Agent Framework with Human-in-the-Loop for Multilingual Legal Terminology Mapping
- 构建多智能体系统,人类与AI分工协作处理法律文本
- 在35部中文法规上实现高精度、一致的多语言术语映射
- 适合法律翻译、司法AI研发人员参考
跨语言法律术语精准映射仍是重大挑战,尤其在中文与日语这类存在大量同形异义词的语言对之间。现有资源和标准化工具匮乏。为此,我们提出一种人机协同的多语言法律术语数据库构建方法,基于多智能体框架,融合大语言模型与法律领域专家,贯穿原始文档预处理、篇章级对齐、术语抽取、映射及质量保障全过程。不同于单一自动化流程,本方法强调人类专家在多智能体系统中的参与方式:AI负责重复性任务如OCR、文本切分、语义对齐与初步术语提取;人类专家则利用上下文知识与法律判断进行关键审核与监督。基于包含35部核心中国法律及其英日文译本的三语平行语料库测试显示,该人机协同多智能体工作流不仅提升了术语映射的准确率与一致性,相比传统人工方法更具可扩展性。
原文摘要 · Abstract (English)
Accurately mapping legal terminology across languages remains a significant challenge, especially for language pairs like Chinese and Japanese, which share a large number of homographs with different meanings. Existing resources and standardized tools for these languages are limited. To address this, we propose a human-AI collaborative approach for building a multilingual legal terminology database, based on a multi-agent framework. This approach integrates advanced large language models and legal domain experts throughout the entire process-from raw document preprocessing, article-level alignment, to terminology extraction, mapping, and quality assurance. Unlike a single automated pipeline, our approach places greater emphasis on how human experts participate in this multi-agent system. Humans and AI agents take on different roles: AI agents handle specific, repetitive tasks, such as OCR, text segmentation, semantic alignment, and initial terminology extraction, while human experts provide crucial oversight, review, and supervise the outputs with contextual knowledge and legal judgment. We tested the effectiveness of this framework using a trilingual parallel corpus comprising 35 key Chinese statutes, along with their English and Japanese translations. The experimental results show that this human-in-the-loop, multi-agent workflow not only improves the precision and consistency of multilingual legal terminology mapping but also offers greater scalability compared to traditional manual methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。