arXiv:2510.18108cs.CL2025-10

对比四款AI处理法律任务,专用模型表现更优

Na Prática, qual IA Entende o Direito? Um Estudo Experimental com IAs Generalistas e uma IA Jurídica

  • 用法律理论+专业人员评估,测试AI在真实工作场景的表现
  • 专用AI JusIA 在准确性、逻辑一致性上全面领先其他三款
  • 适合法律从业者和研究者参考,验证专用模型必要性

本研究开展了一项关于通用型AI在法律领域应用的实验评估,结合法律理论(如实质正确性、体系一致性、论证完整性)与48名法律专业人士的实证评价。测试了四种系统(JusIA、ChatGPT免费版、ChatGPT专业版、Gemini),任务模拟律师日常实务。结果显示,专用模型JusIA始终优于通用模型,证明领域专业化与理论驱动的评估对生成可靠法律AI输出至关重要。

原文摘要 · Abstract (English)

This study presents the Jusbrasil Study on the Use of General-Purpose AIs in Law, proposing an experimental evaluation protocol combining legal theory, such as material correctness, systematic coherence, and argumentative integrity, with empirical assessment by 48 legal professionals. Four systems (JusIA, ChatGPT Free, ChatGPT Pro, and Gemini) were tested in tasks simulating lawyers' daily work. JusIA, a domain-specialized model, consistently outperformed the general-purpose systems, showing that both domain specialization and a theoretically grounded evaluation are essential for reliable legal AI outputs.

法律AI通用模型评测方法领域专用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。