研究智能体如何塑造人类形象,发现其评价以能力为主导。
How Agents Represent Humans: Human-Directed Stereotypes in an Open Agent Social Network

- 构建四维度评估框架,分析智能体对人类的刻板印象
- 人类被评价为高能力者,且常被视为知识、文化或身体主体
- 偏见源于曝光度与内容选择,非固定内外群体对立
基于大语言模型的智能体正日益部署于持久性社交环境,生成内容可被发布、回复、记忆与复用。本文在开放型智能体社交平台 Moltbook 上研究智能体对人类的刻板印象,探讨其如何将人类作为社会类别进行建构。针对人类目标的分析引入包含道德、友善、能力、自主性四个评价维度的标注框架,以及描述性“他者”归因的二级子类体系。研究发现,能力是人类评价的核心维度;诸多“他者”归因将人类描述为认知主体、文化主体或具身主体。进一步考察这些人类形象在人机叙事语境及平台级传播中的表现,并通过行为主控亲和度对比分析智能体内部社区反馈。相较于人类网络社区常见的稳定内外排斥模式,Moltbook 的反馈模式更受曝光度、作者可见性与内容选择影响。结果表明,智能体社会中的偏见应视为话语过程,而不仅限于孤立模型输出。
原文摘要 · Abstract (English)
LLM-based agents are increasingly deployed in persistent social environments, where generated claims can be posted, replied to, remembered, and reused. We study human-directed stereotypes on Moltbook, an open agent-native social platform, asking how agents construct humans as a social category. For this human-target analysis, we introduce an annotation framework with four evaluative dimensions---morality, friendliness, competence, and autonomy---and a second-stage subtype scheme for descriptive \textit{other} attributions. We find that competence dominates human-directed evaluations, while many \textit{other} attributions describe humans as epistemic, cultural, or embodied subjects. We further examine how these human representations appear in human--agent narrative contexts and platform-level circulation. As an auxiliary comparison, we analyze agent-internal community feedback through behavioral host affinity. Rather than reproducing the stable insider--outsider rejection often observed in human online communities, Moltbook feedback patterns are better explained by exposure, author visibility, and content selection. These findings suggest that bias in agent societies should be studied not only as isolated model output, but also as a discourse process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。