用预测编码分析即时消息,提升法律文档分类效率。
A Feasibility Experiment on the Application of Predictive Coding to Instant Messaging Corpora
- 按天分组消息,结合特征选择与逻辑回归建模
- 通过降维提升模型性能,尤其优化量化特征
- 在彭博即时消息数据集上验证可行性与成本优势
预测编码常用于法律领域的文档分类,但即时消息因语言非正式、内容短小,带来额外挑战。本文提出一种数据管理流程,将消息按天聚合为聊天单元,结合特征选择与逻辑回归分类器,构建经济可行的预测编码方案。通过降维技术优化模型表现,重点关注量化特征的处理。实验基于富含定量信息的 Instant Bloomberg 数据集,验证了方法的有效性,并展示了该方案带来的成本节约潜力。
原文摘要 · Abstract (English)
Predictive coding, the term used in the legal industry for document classification using machine learning, presents additional challenges when the dataset comprises instant messages, due to their informal nature and smaller sizes. In this paper, we exploit a data management workflow to group messages into day chats, followed by feature selection and a logistic regression classifier to provide an economically feasible predictive coding solution. We also improve the solution's baseline model performance by dimensionality reduction, with focus on quantitative features. We test our methodology on an Instant Bloomberg dataset, rich in quantitative information. In parallel, we provide an example of the cost savings of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。