为大模型作为知识主体设计可信框架,确保其可靠且符合人类认知规范。
Architecting Trust in Artificial Epistemic Agents
- 提出可信知识型AI的三要素:认知能力、可证伪性与伦理行为
- 强调需构建保护人类认知韧性的知识安全区与溯源系统
- 适合关注AI治理、认知安全与人机协同知识体系的研究者
大型语言模型日益扮演知识型代理角色——既能自主追求知识目标,又能主动塑造我们共享的知识环境。它们负责筛选我们接收到的信息,常取代传统搜索方式,并广泛用于生成个人及高度专业化的建议。这些模型如何执行功能,是否可靠且与个体和集体认知规范相匹配,对我们的决策具有深远影响。我们主张,知识型AI在复杂多智能体交互中对知识创造、整理与整合的影响,催生了新的信息依赖关系,亟需在评估与治理上实现根本转变。一个校准良好的生态系统可增强人类判断力与集体决策,而对齐不当的代理则可能引发认知能力退化与认知偏移,因此将模型与人类认知规范对齐至关重要。为此,我们提出以构建与培育知识型AI可信度为核心框架:使其与人类知识目标一致;强化周边社会认知基础设施。可信的AI代理应展现认知能力、强可证伪性与认知美德行为,依托技术溯源系统与‘知识圣殿’以保护人类认知韧性。这一规范性路线图提供了未来AI作为稳健包容知识生态可靠伙伴的实现路径。
原文摘要 · Abstract (English)
Large language models increasingly function as epistemic agents -- entities that can 1) autonomously pursue epistemic goals and 2) actively shape our shared knowledge environment. They curate the information we receive, often supplanting traditional search-based methods, and are frequently used to generate both personal and deeply specialized advice. How they perform these functions, including whether they are reliable and properly calibrated to both individual and collective epistemic norms, is therefore highly consequential for the choices we make. We argue that the potential impact of epistemic AI agents on practices of knowledge creation, curation and synthesis, particularly in the context of complex multi-agent interactions, creates new informational interdependencies that necessitate a fundamental shift in evaluation and governance of AI. While a well-calibrated ecosystem could augment human judgment and collective decision-making, poorly aligned agents risk causing cognitive deskilling and epistemic drift, making the calibration of these models to human norms a high-stakes necessity. To ensure a beneficial human-AI knowledge ecosystem, we propose a framework centered on building and cultivating the trustworthiness of epistemic AI agents; aligning AI these agents with human epistemic goals; and reinforcing the surrounding socio-epistemic infrastructure. In this context, trustworthy AI agents must demonstrate epistemic competence, robust falsifiability, and epistemically virtuous behaviors, supported by technical provenance systems and "knowledge sanctuaries" designed to protect human resilience. This normative roadmap provides a path toward ensuring that future AI systems act as reliable partners in a robust and inclusive knowledge ecosystem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。