对比NLP模型在法律谎言检测中的表现,发现领域适配很重要。
Semantics of Subterfuge: Benchmarking Legal Deception Detection Against General-domain State-of-the-Art
- 用统一实验比较多种模型在法律与通用领域的谎言识别能力
- 通用领域用微调模型效果好,法律领域少样本LLM更稳定
- 思维链提示常不如直接分类,适合需要可解释性的法律场景
谎言检测对司法程序、执法和网络安全部门具有重要意义。尽管人类判断存在准确性和可扩展性局限,自然语言处理(NLP)提供了数据驱动的替代方案。本文系统回顾并对比了面向法律领域的自动谎言检测(ADD)研究进展,涵盖从特征工程机器学习到大语言模型(LLM)的演进。我们在七个数据集(两个法律类,五个通用领域)上进行统一实证评估,比较六种微调的Transformer模型和七种LLM,在四种提示策略下的表现。结果表明,模型表现显著受领域影响:在数据丰富的通用领域,微调模型表现优异;在资源稀缺的法律领域,少样本LLM仍具竞争力。思维链(Chain-of-Thought)提示策略通常表现较差,远不如直接分类。这些发现凸显了高风险法律场景中领域自适应与可解释系统的重要性。
原文摘要 · Abstract (English)
Deception detection has critical implications for legal proceedings, law enforcement, and online security. Although human judgment is limited in accuracy and scalability, Natural Language Processing (NLP) offers a data-driven alternative. We present a survey and comparative analysis of NLP-based Automatic Deception Detection (ADD) focusing on the legal domain, reviewing the evolution from feature-based machine learning to Large Language Model (LLM) approaches. We conduct a unified empirical evaluation across seven datasets (two legal, five general-domain), comparing six fine-tuned transformer models and seven LLMs under four prompting strategies. The results show strong domain sensitivity, with fine-tuned models excelling in data-rich general domains and few-shot LLMs remaining competitive in low-resource legal settings. Chain-of-Thought prompting often underperforms direct classification. These findings highlight the need for domain adaptation and interpretable systems in high-stakes legal contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。