首个面向印度保释判決的公开数据集,助力法律NLP研究
IndianBailJudgments-1200: A Multi-Attribute Dataset for Legal NLP on Indian Bail Orders
- 构建1200份印度保释判決的多属性标注数据集
- 涵盖20余项属性,支持判決结果预测等任务
- 适用于法律AI、司法公平性分析的研究者
由于缺乏结构化数据,印度地区的法律自然语言处理研究进展缓慢。我们提出了IndianBailJudgments-1200,一个包含1200份印度法院保释判决的基准数据集,涵盖保释结果、刑法条款(IPC)、罪名类型和法律推理等20多个属性。标注通过提示工程优化的GPT-4o管道生成,并进行了一致性验证。该资源支持判決结果预测、摘要生成和公平性分析等多种法律NLP任务,是首个专注于印度保释法理学的公开数据集。
原文摘要 · Abstract (English)
Legal NLP remains underdeveloped in regions like India due to the scarcity of structured datasets. We introduce IndianBailJudgments-1200, a new benchmark dataset comprising 1200 Indian court judgments on bail decisions, annotated across 20+ attributes including bail outcome, IPC sections, crime type, and legal reasoning. Annotations were generated using a prompt-engineered GPT-4o pipeline and verified for consistency. This resource supports a wide range of legal NLP tasks such as outcome prediction, summarization, and fairness analysis, and is the first publicly available dataset focused specifically on Indian bail jurisprudence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。