arXiv:2501.00106cs.SEcs.AI2025-01被引 4

用AI自动分析数据集许可协议,准确率超六成,提速94%。

LicenseGPT: A Fine-tuned Foundation Model for Publicly Available Dataset License Compliance

  • 基于500份专家标注的许可协议微调大模型,专攻数据集授权分析
  • 准确率提升至64.3%,比现有最佳模型高20.55个百分点
  • 律师实测效率提高94%,适合法律与AI开发团队协作使用

数据集许可合规是商业AI产品开发中的关键挑战,尤其在公开数据集广泛应用的背景下。许可证条款模糊带来重大法律风险,即使对软件知识产权律师也难准确解读。本文提出LicenseGPT,一个针对数据集许可合规分析优化的微调基础模型。我们评估了现有法律类基础模型(FM),发现表现最好的模型预测一致率(PA)仅43.75%。LicenseGPT在500份由法律专家标注的许可证数据集上微调后,将PA提升至64.30%,显著优于法律及通用基础模型。通过与软件知识产权律师的A/B测试和用户研究,证明其可将每份许可证分析时间从108秒降至6秒,效率提升94.44%,且不牺牲准确性。律师认为LicenseGPT是高效辅助工具,但复杂案例仍需人工判断。本工作展示了专用AI工具在法律实践中的潜力,并提供公开资源供从业者与研究者使用。

原文摘要 · Abstract (English)

Dataset license compliance is a critical yet complex aspect of developing commercial AI products, particularly with the increasing use of publicly available datasets. Ambiguities in dataset licenses pose significant legal risks, making it challenging even for software IP lawyers to accurately interpret rights and obligations. In this paper, we introduce LicenseGPT, a fine-tuned foundation model (FM) specifically designed for dataset license compliance analysis. We first evaluate existing legal FMs (i.e., FMs specialized in understanding and processing legal texts) and find that the best-performing model achieves a Prediction Agreement (PA) of only 43.75%. LicenseGPT, fine-tuned on a curated dataset of 500 licenses annotated by legal experts, significantly improves PA to 64.30%, outperforming both legal and general-purpose FMs. Through an A/B test and user study with software IP lawyers, we demonstrate that LicenseGPT reduces analysis time by 94.44%, from 108 seconds to 6 seconds per license, without compromising accuracy. Software IP lawyers perceive LicenseGPT as a valuable supplementary tool that enhances efficiency while acknowledging the need for human oversight in complex cases. Our work underscores the potential of specialized AI tools in legal practice and offers a publicly available resource for practitioners and researchers.

AI法律许可分析大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。