arXiv:2410.20104cs.CL2024-10被引 4

用深度学习预测印尼法院判刑时长,提升司法透明度

Hybrid Deep Learning for Legal Text Analysis: Predicting Punishment Durations in Indonesian Court Rulings

  • 融合CNN、BiLSTM与注意力机制,捕捉法律文本的局部与长期特征
  • 使用前30%高频词使预测准确率提升,R²达0.5893
  • 优化文本清洗流程,改善拼写错误对模型的影响,适合法律AI研究者

印尼法院系统因公众理解不足和判决不一致引发广泛不满,给法官带来压力。本研究构建基于深度学习的量刑时长预测系统,采用结合CNN、BiLSTM与注意力机制的混合模型,在法律文本中有效捕捉局部模式与长期依赖关系,取得R²为0.5893的预测性能。文档摘要效果不佳,而仅保留前30%最频繁词汇显著提升表现,表明聚焦核心法律术语可在信息保留与计算效率间取得平衡。同时,改进的文本标准化流程解决了常见拼写错误与词语误合并问题,大幅提高模型表现。该研究为法律文书自动化处理提供支持,助力专业人士与公众理解判决结果,推动印尼司法系统的透明化与可理解性,为实现更一致、可访问的司法决策铺路。

原文摘要 · Abstract (English)

Limited public understanding of legal processes and inconsistent verdicts in the Indonesian court system led to widespread dissatisfaction and increased stress on judges. This study addresses these issues by developing a deep learning-based predictive system for court sentence lengths. Our hybrid model, combining CNN and BiLSTM with attention mechanism, achieved an R-squared score of 0.5893, effectively capturing both local patterns and long-term dependencies in legal texts. While document summarization proved ineffective, using only the top 30% most frequent tokens increased prediction performance, suggesting that focusing on core legal terminology balances information retention and computational efficiency. We also implemented a modified text normalization process, addressing common errors like misspellings and incorrectly merged words, which significantly improved the model's performance. These findings have important implications for automating legal document processing, aiding both professionals and the public in understanding court judgments. By leveraging advanced NLP techniques, this research contributes to enhancing transparency and accessibility in the Indonesian legal system, paving the way for more consistent and comprehensible legal decisions.

法律AI深度学习文本分析NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。