首个印地语法律隐喻语料库,助力低资源语言司法文本理解
Figurative Justice: Detecting metaphors in Hindi judgements with qualitative assessment and transformers

- 构建印地语法律隐喻语料库HiLeMe,专家标注符合MIPVU标准
- 基于mBERT的Transformer模型在印地语法律隐喻检测中表现优异
- 为22种印度官方语言的司法语义分析提供可扩展框架
隐喻是词语的修辞用法,用于概念映射。在法律语境中,隐喻检测至关重要,因其能塑造法律意义与概念,产生显著后果。法官、律师和立法者在法律话语中使用隐喻性表述,直接影响个人命运、司法决策、法律解释及公众认知,尤其在人权侵犯案件中,语言选择决定惩罚严重性、社会评价与判决结果。尽管英语、西班牙语等主流语言已有自动隐喻检测研究,但低资源语言如印地语仍缺乏相关工作。印地语法律语料库匮乏导致自然语言处理模型难以训练。现有研究仅使用卷积神经网络对保释判决进行分类,尚未有专门针对隐喻检测的模型。本文从印地语法律数据语料库(HLDC)中提取判决文书,构建了印地语法律隐喻语料库(HiLeMe),由法律专家采用MIPVU框架标注隐喻结构。我们基于mBERT微调开发了变压器架构的隐喻检测模型,在法律分类任务中表现优于传统模型,为解读司法心理提供了新视角。本研究推动了低资源语言法律语篇自动化分析的发展,并有望推广至22种印度宪法第八表语言。
原文摘要 · Abstract (English)
Metaphors are figurative use of words for conceptual mapping. Metaphor detection in the legal context has been crucial as metaphors are persuasive juridical means of creating legal meaning and concepts resulting in significant consequences. Metaphorical framing in legal discourse by judges, lawyers, and legislators brings about real-time implications upon individuals and influences judicial decision-making, argumentation and interpretation of laws. This is crucial in Human Rights infringement cases where language determines severity of punishment, public perception and judicial outcomes. While automatic metaphor detection in major languages like English, Spanish, Polish, Lithuanian have aided in understanding inherent intentions of metaphorical use of language, there is no such attempt in low-resource languages like Hindi. The dearth of annotated legal corpora in Hindi makes it difficult to develop NLP models and detect metaphors in judicial proceedings. In the Indian context, Convolutional Neural Networks (CNNs) have been used for classification of bail judgements, however there are no existing models designed for metaphor detection. We present a Hindi Legal Metaphor Corpus (HiLeMe) by isolating judgements from Hindi Legal Data Corpus (HLDC). Legal experts annotated HiLeMe to classify metaphorical constructions using the MIPVU schema. We downstreamed an mBERT on Hindi legal metaphor detection task. We built a transformer-based architecture for metaphor detection that are known to outperform traditional models in legal classification tasks. This model provides insights into the judicial psyche for decoding judicial decisions. Our research contributes to advancing automated models in legal discourse in low-resource languages like Hindi and envisages adoption into 22 Indian schedule languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。