为AI生成内容添加可验证元数据,确保其可靠复用。
A Prompt-Aware Structuring Framework for Reliable Reuse of AI-Generated Content in the Agentic Web
- 生成时自动附加模块化提示、上下文和置信度等结构化元数据
- 通过可验证凭证实现内容可信度评估与合规性检查
- 适合需要安全复用AI内容的模型微调与知识蒸馏场景
大型语言模型(LLMs)及其构建的软件代理(AI代理)正推动互联网从以人类为中心向由AI代理驱动的“智能体网络”转型。然而,对于预计将主导网络的AI生成内容(AIGC),目前缺乏机制来验证其可靠性、可重现性或版权合规性。这种透明度缺失可能导致链式幻觉和合规问题。为此,本文提出一个框架,在生成阶段自动为AIGC附加结构化元数据,包括模块化提示、上下文、思维过程、模型信息、超参数及置信度,并封装于可验证凭证中,支持对AIGC的可信评估与安全复用。该框架可高效管理结构化AIGC,促进其在模型微调与知识蒸馏等应用中的安全使用。
原文摘要 · Abstract (English)
The evolution of Large Language Models (LLMs) and the software agents built on them (AI agents) marks a turning point in the transition from a human-centric Web to an ``Agentic Web'' driven by AI agents. However, for AI-Generated Content (AIGC), which is expected to dominate the Web, there is currently no mechanism for agents to verify its reliability, reproducibility, or license compliance during generation. This lack of transparency risks causing chained hallucinations and compliance violations through the reuse of AIGC. Consequently, a framework to manage the provenance and generation conditions of AIGC is essential. In this paper, we present a framework that automatically attaches structured metadata to AIGC at generation time, including modularized prompts, contexts, thoughts, model information, hyperparameters, and confidence. The metadata is enveloped together with verifiable credentials to support the reliable assessment and reuse of AIGC. This framework enables efficient curation of structured AIGC and facilitates its safe use for applications such as fine-tuning and knowledge distillation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。