用大模型从自然语言生成数字版权政策,准确率达91.95%
From Instructions to ODRL Usage Policies: An Ontology Guided Approach
- 基于ODRL本体和文档构建提示词,引导大模型生成策略
- 在文化领域数据空间中测试,知识图谱构建准确率最高达91.95%
- 适合需要自动化制定数据使用规则的机构或系统开发者
本研究提出一种利用大语言模型(如GPT-4)从自然语言指令自动生成W3C开放数字权利语言(ODRL)使用策略的方法。该方法以ODRL本体及其文档为核心提示内容,假设经筛选优化的本体文档能更有效引导策略生成。研究设计多种启发式方法,将ODRL本体及文档适配至端到端知识图谱构建流程。在数据空间场景下评估,即多个参与组织间用于文化领域的可信数据交换分布式基础设施。构建包含12个不同复杂度用例的基准测试集。评估结果表明,知识图谱构建准确率最高达91.95%。
原文摘要 · Abstract (English)
This study presents an approach that uses large language models such as GPT-4 to generate usage policies in the W3C Open Digital Rights Language ODRL automatically from natural language instructions. Our approach uses the ODRL ontology and its documentation as a central part of the prompt. Our research hypothesis is that a curated version of existing ontology documentation will better guide policy generation. We present various heuristics for adapting the ODRL ontology and its documentation to guide an end-to-end KG construction process. We evaluate our approach in the context of dataspaces, i.e., distributed infrastructures for trustworthy data exchange between multiple participating organizations for the cultural domain. We created a benchmark consisting of 12 use cases of varying complexity. Our evaluation shows excellent results with up to 91.95% accuracy in the resulting knowledge graph.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。