测试大模型识别服务条款中不公平条款的能力,发现表现仅略好于随机。
Are LLM-based methods good enough for detecting unfair terms of service?
- 构建12个问题的隐私政策数据集,评估模型对条款的理解能力。
- 最佳模型(ChatGPT4)准确率仍远低于实用水平,整体表现仅略高于随机。
- 开源模型部分表现优于某些商业模型,但整体仍不足够可靠。
每天有无数用户在使用各类应用和网站时签署大量服务条款(ToS),这些条款通常长达十几页,用户往往未加细读便点击同意,可能无意中放弃数据隐私等权利。大型语言模型(LLMs)擅长解析长文本,或可帮助识别条款中的可疑内容。为此,我们构建了一个包含12个问题的数据集,针对从热门网站爬取的隐私政策进行评估。随后,对多个开源及商用聊天机器人(如ChatGPT)提问,并将答案与人工标注的基准答案对比。结果显示,部分开源模型表现优于某些商用模型,但最优结果来自商用模型ChatGPT4。然而,所有模型的表现均仅略高于随机水平。因此,当前模型尚不足以大规模用于此任务,需显著提升性能。
原文摘要 · Abstract (English)
Countless terms of service (ToS) are being signed everyday by users all over the world while interacting with all kinds of apps and websites. More often than not, these online contracts spanning double-digit pages are signed blindly by users who simply want immediate access to the desired service. What would normally require a consultation with a legal team, has now become a mundane activity consisting of a few clicks where users potentially sign away their rights, for instance in terms of their data privacy, to countless online entities/companies. Large language models (LLMs) are good at parsing long text-based documents, and could potentially be adopted to help users when dealing with dubious clauses in ToS and their underlying privacy policies. To investigate the utility of existing models for this task, we first build a dataset consisting of 12 questions applied individually to a set of privacy policies crawled from popular websites. Thereafter, a series of open-source as well as commercial chatbots such as ChatGPT, are queried over each question, with the answers being compared to a given ground truth. Our results show that some open-source models are able to provide a higher accuracy compared to some commercial models. However, the best performance is recorded from a commercial chatbot (ChatGPT4). Overall, all models perform only slightly better than random at this task. Consequently, their performance needs to be significantly improved before they can be adopted at large for this purpose.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。