arXiv:2602.09286cs.AIcs.CY2026-02被引 2

不同AI角色社区对人类监管的理解截然不同,需按角色设计监督机制。

Human Control Is the Anchor, Not the Answer: Early Divergence of Oversight in Agentic AI Communities

  • 通过对比两个Reddit社区,发现人类监管的含义因角色而异。
  • 执行类社区关注行动风险,社交类社区关注身份与责任问题。
  • 为不同角色定制监督方案,比统一管控更有效。

代理型AI的监管常被视作单一目标(人类控制),但早期采用可能催生特定角色期待。本文对比了2026年1月至2月活跃的两个Reddit社区:r/OpenClaw(部署与运维)和r/Moltbook(以代理为中心的社会互动)。我们将此阶段视为监督期望形成的初期结晶期,此时规范尚未稳定。通过共享比较空间中的主题建模、粗粒度监督主题抽象、加权显著性分析及分歧检验,结果显示两社区具有强可区分性(JSD=0.418,余弦相似度=0.372,置换检验p=0.0005)。尽管“人类控制”是共同锚点术语,其具体含义存在分化:r/OpenClaw侧重执行过程中的安全护栏与恢复机制(行动风险),而r/Moltbook则聚焦公共互动中的身份、合法性与问责(意义风险)。这一差异为设计与评估适配代理角色的监督机制提供了可迁移视角,避免“一刀切”管控策略。

原文摘要 · Abstract (English)

Oversight for agentic AI is often discussed as a single goal ("human control"), yet early adoption may produce role-specific expectations. We present a comparative analysis of two newly active Reddit communities in Jan--Feb 2026 that reflect different socio-technical roles: r/OpenClaw (deployment and operations) and r/Moltbook (agent-centered social interaction). We conceptualize this period as an early-stage crystallization phase, where oversight expectations form before norms reach equilibrium. Using topic modeling in a shared comparison space, a coarse-grained oversight-theme abstraction, engagement-weighted salience, and divergence tests, we show the communities are strongly separable (JSD =0.418, cosine =0.372, permutation $p=0.0005$). Across both communities, "human control" is an anchor term, but its operational meaning diverges: r/OpenClaw} emphasizes execution guardrails and recovery (action-risk), while r/Moltbook} emphasizes identity, legitimacy, and accountability in public interaction (meaning-risk). The resulting distinction offers a portable lens for designing and evaluating oversight mechanisms that match agent role, rather than applying one-size-fits-all control policies.

AI监管代理系统社会技术

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。