arXiv:2505.07834cs.NIcs.AI2025-05被引 1

ai.txt让AI更精准地遵守网页访问规则,提升合规性。

ai.txt: A Domain-Specific Language for Guiding AI Interactions with the Internet

  • 用自然语言定义网页元素级访问权限,比传统规则更灵活
  • 支持代码补全和自动XML生成,便于实际部署
  • 适合关注AI伦理与网络合规的研究者与开发者

我们提出ai.txt,一种新型领域特定语言(DSL),旨在明确规范人工智能模型、智能体与网络内容之间的交互行为,解决广泛采用的robots.txt标准在精度与语义表达上的不足。随着AI越来越多地参与训练、摘要和内容修改等任务,现有监管手段缺乏足够的粒度和语义能力以确保伦理与法律合规。ai.txt通过引入基于元素级别的精细控制,并融合可被AI理解的自然语言指令,超越传统的URL级访问控制。为促进实际应用,我们提供集成开发环境,支持代码自动补全与自动XML生成。此外,提出两种合规机制:基于XML的程序化执行和自然语言提示集成,并通过初步实验与案例研究验证其有效性。本方法旨在推动AI-互联网交互的治理,助力数字生态中负责任的AI使用。

原文摘要 · Abstract (English)

We introduce ai.txt, a novel domain-specific language (DSL) designed to explicitly regulate interactions between AI models, agents, and web content, addressing critical limitations of the widely adopted robots.txt standard. As AI increasingly engages with online materials for tasks such as training, summarization, and content modification, existing regulatory methods lack the necessary granularity and semantic expressiveness to ensure ethical and legal compliance. ai.txt extends traditional URL-based access controls by enabling precise element-level regulations and incorporating natural language instructions interpretable by AI systems. To facilitate practical deployment, we provide an integrated development environment with code autocompletion and automatic XML generation. Furthermore, we propose two compliance mechanisms: XML-based programmatic enforcement and natural language prompt integration, and demonstrate their effectiveness through preliminary experiments and case studies. Our approach aims to aid the governance of AI-Internet interactions, promoting responsible AI use in digital ecosystems.

AI治理网页爬虫自然语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。