arXiv:2602.15273cs.CYcs.CL2026-02

构建可模拟信息偏见演化的数据集与框架,研究算法推荐如何影响用户认知健康。

FrameRef: A Framing Dataset and Simulation Testbed for Modeling Bounded Rational Information Health

  • 用五类框架重构107万条陈述,生成带偏见倾向的模拟信息流。
  • 微小的接受度变化随时间累积,导致信息健康轨迹显著分化。
  • 适合研究算法偏见、数字健康与人机交互的学者及政策制定者。

信息生态系统正深刻影响人们对负面数字体验的内化,引发对信息健康长期后果的担忧。在现代搜索与推荐系统中,排序与个性化策略在塑造用户暴露环境及其长期影响方面起核心作用。为在受控环境下研究此类影响,我们提出 FrameRef——一个包含1,073,740条跨五类框架维度(权威性、共识性、情绪性、声望性、耸动性)的系统重构陈述的大规模数据集,并构建基于仿真的框架,用于建模顺序性信息暴露与强化动态。在此框架中,通过微调语言模型并引入框架条件损失衰减,构建具有针对性偏见的代理人格,同时保持任务能力。采用蒙特卡洛轨迹采样方法表明,接受度和信心的微小系统性变化会随时间累积,造成信息健康轨迹的显著偏离。人类评估进一步验证了生成框架对判断的可测量影响。本研究提供的数据集与框架为系统性信息健康研究提供仿真基础,补充并支持以人为本的负责任研究。代码、文档、人类评估数据及人格适配模型已开源:https://github.com/infosenselab/frameref。

原文摘要 · Abstract (English)

Information ecosystems increasingly shape how people internalize exposure to adverse digital experiences, raising concerns about the long-term consequences for information health. In modern search and recommendation systems, ranking and personalization policies play a central role in shaping such exposure and its long-term effects on users. To study these effects in a controlled setting, we present FrameRef, a large-scale dataset of 1,073,740 systematically reframed claims across five framing dimensions: authoritative, consensus, emotional, prestige, and sensationalist, and propose a simulation-based framework for modeling sequential information exposure and reinforcement dynamics characteristic of ranking and recommendation systems. Within this framework, we construct framing-sensitive agent personas by fine-tuning language models with framing-conditioned loss attenuation, inducing targeted biases while preserving overall task competence. Using Monte Carlo trajectory sampling, we show that small, systematic shifts in acceptance and confidence can compound over time, producing substantial divergence in cumulative information health trajectories. Human evaluation further confirms that FrameRef's generated framings measurably affect human judgment. Together, our dataset and framework provide a foundation for systematic information health research through simulation, complementing and informing responsible human-centered research. We release FrameRef, code, documentation, human evaluation data, and persona adapter models at https://github.com/infosenselab/frameref.

信息健康算法偏见仿真框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。