构建首个公开的印度信息公开决策数据集,助力自动化分析与上诉预测。
RTI-Bench: A Structured Dataset for Indian Right-to-Information Decision Analysis
- 构建包含结果标签、免责引用和推理结构的结构化数据集
- 覆盖1516份判决,标签准确率达95.3%,零样本模型准确率超57%
- 适合政策分析、法律AI研究者使用
印度《信息获取法》2005年赋予公民向公共机构申请信息的权利,但实践中多数人难以理解中央信息委员会(CIC)裁决中使用的复杂行政语言,更无法判断申诉是否值得提起。本文提出RTI-Bench,首个公开的结构化印度信息公开行政裁决数据集,包含裁决结果标签、免责条款引用、IRAC式推理组件及程序时间线。数据来自两个来源:1,218份公开指令-响应语料库(通过规则提取添加结构字段),以及从委员会官网收集的298份裁决PDF,涵盖五位委员、三类文档格式,时间跨度为2023至2026年。指令-响应语料库标签覆盖率达89%;239份主要裁决的子集标签覆盖率为51%(本版首次发布)。随机抽取50份标注案例人工审核,标签精确度达95.3%。在100个样本上,零样本Mistral 7B基线在结果预测任务上取得57.3%准确率和37.0%宏平均F1,显著高于多数类基线的14.3%宏平均F1。RTI-Bench已开源:https://huggingface.co/datasets/joyboseroy/rti-bench
原文摘要 · Abstract (English)
India's Right to Information Act, 2005 gives every citizen the right to demand information from public authorities, yet in practice most people cannot make sense of the dense administrative language used in Central Information Commission (CIC) decisions, let alone predict whether an appeal is worth filing. This paper introduces RTI-Bench, a structured dataset of CIC decisions with outcome labels, exemption citations, IRAC-style reasoning components, and procedural timelines. To the best of our knowledge it is the first publicly released structured dataset for Indian RTI administrative decisions. The dataset draws from two sources: 1,218 cases from a publicly available instruction-response corpus (with structured fields added through rule-based extraction), and 298 CIC decision PDFs collected directly from the Commission portal, spanning five commissioners and three document format generations from 2023 to 2026. Label coverage reaches 89% on the instruction-response corpus. For the PDF subset of 239 primary decisions, coverage is 51% in this first release. A random sample of 50 labelled cases was manually reviewed, yielding a label precision of 95.3%. A zero-shot Mistral 7B baseline on 100 cases gives 57.3% accuracy and 37.0% macro-F1 on outcome prediction, well above the majority-class baseline of 14.3% macro-F1. RTI-Bench is available at https://huggingface.co/datasets/joyboseroy/rti-bench
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。