首个香港判決語篇數據集,助力理解法院認定事實、推理與判決邏輯。
HKJudge: A Legal Discourse-Annotated Corpus for Interpreting What Courts Find, How They Reason, and What They Rule

- 以26種修辭角色細粒度標註判決句,涵蓋事實認定、推理與判決三要素。
- 包含近29萬句、650萬詞,由10位專家標註,一致性達κ=0.8。
- 提供首個基於BERT和LLM的法律判決結構分析基準,適用於法學與NLP研究者。
判決是法律實務與法理的核心,但由於缺乏專家標註語料,香港判決的語篇分析長期受到限制。本文提出香港判決語篇數據集(HKJudge),這是首個在句子級別由法律語言學專家標註的法律語篇語料庫。HKJudge涵蓋香港司法體系五級法院的刑事判決,總計約29萬句、650萬詞,全部經專家標註。設計雙層語篇框架:句子級標註26種修辭角色,段落級標註三類量刑要素(指控、監禁期、罰款)。10名法律語言學專家完成標註,組內一致係數κ=0.8。本文提出兩個任務:修辭角色分類與法律要素提取,並對四種BERT模型、兩種開源大模型(零樣本與微調)、四種商業大模型進行首個基準評估。結果顯示,句子級語篇標註對建模香港判決結構具有重要價值,為未來法律判決預測研究奠定數據基礎。數據與代碼已公開於https://github.com/xuanxixi/HKJudge。
原文摘要 · Abstract (English)
Court judgments are central to legal practice and jurisprudence, yet discourse analysis of Hong Kong judgments has received limited attention, owing largely to the absence of expert-annotated corpora. We introduce the Hong Kong Judgment Discourse Dataset (HKJudge), the first sentence-level expert-annotated legal discourse corpus. HKJudge includes criminal judgments across all five levels of HK's court hierarchy, comprising $\sim$290k sentences and $\sim$6.5 million tokens, fully annotated by legal linguistics experts. We design a two-tier discourse schema that captures what facts a court finds, how it reasons, and what it rules. At the sentence level, each sentence is assigned one of 26 rhetorical roles. At the span level, sentences are further annotated with three sentencing elements (charge, imprisonment term, fine). Ten legal linguistics annotators produced the annotations with an inter-annotator agreement of $κ= 0.8$. We formulate two tasks on HKJudge, termed rhetorical role classification and legal element extraction, and provide the first benchmark evaluation of four BERT-based models, two open-source LLMs under zero-shot and fine-tuning settings, and four commercial LLMs on both tasks. Our work demonstrates the value of sentence-level discourse annotation for modeling the structure of HK judgments and provides a rich data foundation for future work on legal judgment prediction. The HKJudge dataset and code are available at https://github.com/xuanxixi/HKJudge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。