arXiv:2601.22547cs.IR2026-01

用个性化智能体模拟短视频用户,真实审计推荐系统中的信息茧房

PersonaAct: Simulating Short-Video Users with Personalized Agents for Counterfactual Filter Bubble Auditing

  • 通过行为分析与结构化提问生成可解释的用户画像
  • 在多模态数据上训练智能体,还原真实用户互动模式
  • 发现B站用户跳出信息茧房能力最强,适用于推荐系统评估

短视频平台依赖个性化推荐,引发信息茧房担忧。大规模审计困难,因真实用户研究成本高且涉及隐私,现有模拟器因依赖文本信号和弱个性化,难以复现真实行为。我们提出PersonaAct框架,利用真实行为轨迹训练多模态个性化智能体,实现对信息茧房广度与深度的审计。通过自动化访谈合成可解释的用户人格,再基于多模态观测使用监督微调与强化学习训练智能体。部署后评估显示,相比通用大模型基线,其行为真实性显著提升,能真实再现用户互动过程。结果表明,用户内容曝光随交互显著收窄,但哔哩哔哩展现最强逃逸潜力。我们发布首个开源多模态短视频数据集及代码,支持可复现的推荐系统审计。

原文摘要 · Abstract (English)

Short-video platforms rely on personalized recommendation, raising concerns about filter bubbles that narrow content exposure. Auditing such phenomena at scale is challenging because real user studies are costly and privacy-sensitive, and existing simulators fail to reproduce realistic behaviors due to their reliance on textual signals and weak personalization. We propose PersonaAct, a framework for simulating short-video users with persona-conditioned multimodal agents trained on real behavioral traces for auditing filter bubbles in breadth and depth. PersonaAct synthesizes interpretable personas through automated interviews combining behavioral analysis with structured questioning, then trains agents on multimodal observations using supervised fine-tuning and reinforcement learning. We deploy trained agents for filter bubble auditing and evaluate bubble breadth via content diversity and bubble depth via escape potential. The evaluation demonstrates substantial improvements in fidelity over generic LLM baselines, enabling realistic behavior reproduction. Results reveal significant content narrowing over interaction. However, we find that Bilibili demonstrates the strongest escape potential. We release the first open multimodal short-video dataset and code to support reproducible auditing of recommender systems.

信息茧房智能体模拟推荐系统多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。