arXiv:2607.17149cs.AIcs.CY2026-07

为理解AI行为根源,提出分层诊断框架

A Diagnostic Framework for AI Agent Behavior

论文配图:A Diagnostic Framework for AI Agent Behavior
图 1 · 摘自论文原文
  • 将AI行为分为基础能力层与行为调节层
  • 揭示行为差异可作为诊断依据
  • 适合研究AI治理与人机协作的学者

AI代理越来越多地介入临床、政治、科学和社会系统,这些系统正是行为科学家研究的对象。评估此类系统需进行源级诊断:相同的行为模式可能源于代理的表征基础,也可能由角色、目标、互动结构和治理规则所塑造。本文提出一种诊断框架——层级归因。基础计算层通过架构、记忆、感知、注意力和表征决定哪些行为是可能的;行为调节层则通过身份、资源、目标、社会互动、制度约束和治理机制来调节这些能力的表达。该框架阐明三个关键结论:代理有效性是模型-任务-层级的关系;人与AI的行为差异可提供诊断证据;治理必须先进行源归属再实施干预。因此,将AI代理视为行为主体,要求评估方法能确定行为源头,再决定如何解释、验证或治理。

原文摘要 · Abstract (English)

AI agents increasingly act within the same clinical, political, scientific, and social systems that behavioral scientists study. Evaluating these systems requires source-level diagnosis: the same behavioral pattern may arise from an agent representational substrate or from the roles, objectives, interaction structures, and governance rules that shape its expression. This Perspective proposes a diagnostic framework for AI agent behavior: layer attribution. The foundational computational layer defines what behaviors are possible through architecture, memory, perception, attention, and representation. The behavioral modulation layer shapes how those capacities are expressed through identity, resources, objectives, social interaction, institutional constraints, and governance. The framework clarifies three consequences: surrogate validity is a model-task-layer relation, human-AI divergence provides diagnostic evidence, and governance requires source attribution before intervention. Treating AI agents as behavioral actors therefore requires evaluation methods that determine where behavior originates before deciding how to explain, validate, or govern it.

AI治理行为诊断人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。