对比四种嵌入模型在工程延期纠纷文档中的分类效果,提升法律文本分析效率。
Empirical Evaluation of Embedding Models in the Context of Text Classification in Document Review in Construction Delay Disputes
- 用KNN和逻辑回归测试四种嵌入模型的分类性能
- 在标注数据集上实现延迟相关语句的二分类准确识别
- 适合法律科技与建筑纠纷分析领域的研究者参考
文本嵌入是将词语、短语或整篇文档转换为实数向量的技术,能捕捉文本间的语义关系。本文通过对比四种嵌入模型,在标注数据集上使用KNN和逻辑回归进行二分类任务,判断文本片段是否与'延迟'相关。研究聚焦于工程延期争议文档审查中的文本分类应用,验证了嵌入模型在提升法律文本分析效率与准确性方面的潜力,为复杂调查场景下的决策提供支持。
原文摘要 · Abstract (English)
Text embeddings are numerical representations of text data, where words, phrases, or entire documents are converted into vectors of real numbers. These embeddings capture semantic meanings and relationships between text elements in a continuous vector space. The primary goal of text embeddings is to enable the processing of text data by machine learning models, which require numerical input. Numerous embedding models have been developed for various applications. This paper presents our work in evaluating different embeddings through a comprehensive comparative analysis of four distinct models, focusing on their text classification efficacy. We employ both K-Nearest Neighbors (KNN) and Logistic Regression (LR) to perform binary classification tasks, specifically determining whether a text snippet is associated with 'delay' or 'not delay' within a labeled dataset. Our research explores the use of text snippet embeddings for training supervised text classification models to identify delay-related statements during the document review process of construction delay disputes. The results of this study highlight the potential of embedding models to enhance the efficiency and accuracy of document analysis in legal contexts, paving the way for more informed decision-making in complex investigative scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。