为生成式内容创建者建立可追踪的访问协议,让创作者在模型使用中获得应有归属。
Sovereign Context Protocol: An Open Attribution Layer for Human-Generated Content in the Age of Large Language Models
- 设计开放协议,实现内容访问时的实时溯源与授权记录。
- 支持六项核心功能,涵盖搜索、验证、评分及审计,通过REST和MCP接口调用。
- 适用于内容创作者、平台方及关注数据伦理的研究者。
大型语言模型(LLMs)大量使用人类生成内容进行训练与实时推理,但内容创作者在价值链条中仍难以被识别。现有方法或基于模型内部梯度追踪影响,或依赖法律政策透明度要求与版权诉讼,均缺乏运行时的创作归属机制。本文提出主权上下文协议(Sovereign Context Protocol, SCP),一个开源协议规范与参考架构,作为LLM与创作者拥有数据之间的归因感知数据访问层。受Anthropic的模型上下文协议(MCP)启发,SCP标准化了LLM与创作者数据的连接方式,每次访问均被记录、授权并可追溯。协议定义了六项核心方法:创作者档案、语义搜索、内容检索、可信度/价值评分、真实性验证与访问审计,通过REST与MCP兼容接口暴露。我们形式化了消息封装结构,提出包含五类攻击者的威胁模型,设计了一种按日志比例分配收益的归因模型,并报告了基于FastAPI、ChromaDB与NetworkX的参考实现的初步延迟性能。将SCP置于欧盟《人工智能法案》第53条训练数据透明性要求及美国正在进行的版权诉讼背景下,主张需通过协议层级干预,使归因成为数据访问的默认属性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) consume vast quantities of human-generated content for both training and real-time inference, yet the creators of that content remain largely invisible in the value chain. Existing approaches to data attribution operate either at the model-internals level, tracing influence through gradient signals, or at the legal-policy level through transparency mandates and copyright litigation. Neither provides a runtime mechanism for content creators to know when, by whom, and how their work is being consumed. We introduce the Sovereign Context Protocol (SCP), an open-source protocol specification and reference architecture that functions as an attribution-aware data access layer between LLMs and human-generated content. Inspired by Anthropic's Model Context Protocol (MCP), which standardizes how LLMs connect to tools, SCP standardizes how LLMs connect to creator-owned data, with every access event logged, licensed, and attributable. SCP defines six core methods (creator profiles, semantic search, content retrieval, trust/value scoring, authenticity verification, and access auditing) exposed over both REST and MCP-compatible interfaces. We formalize the protocol's message envelope, present a threat model with five adversary classes, propose a log-proportional revenue attribution model, and report preliminary latency benchmarks from a reference implementation built on FastAPI, ChromaDB, and NetworkX. We situate SCP within the emerging regulatory landscape, including the EU AI Act's Article 53 training data transparency requirements and ongoing U.S. copyright litigation, and argue that the attribution gap requires a protocol-level intervention that makes attribution a default property of data access.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。