arXiv:2409.04043cs.CL2024-09被引 1

用AI模拟饮食障碍讨论,测试干预策略效果

Towards Safer Online Spaces: Simulating and Assessing Intervention Strategies for Eating Disorder Discussions

  • 构建LLM驱动的模拟平台,生成跨平台真实对话
  • 重视礼貌的干预提升正向情绪,启发式干预加剧负面情绪
  • 揭示大模型间认知差异,强调选型对模拟真实性影响

饮食障碍是全球数百万人群面临的复杂心理健康问题。社交媒体上的有效干预至关重要,但在实际环境中测试策略存在风险。本文提出一种基于大语言模型的实验平台,用于模拟和评估饮食障碍相关讨论中的干预策略。该框架能生成跨多个平台、模型及话题的合成对话,支持对不同干预方式的受控实验。我们从干预类型、生成模型、社交平台及社区主题四个维度分析干预对对话动态的影响,采用情感、情绪等认知领域指标评估效果。结果表明,以礼貌为导向的干预在所有维度上均显著提升积极情绪与情感基调,而以洞察重置为目标的策略则普遍引发更强烈的负面情绪。此外,我们发现大模型生成对话存在显著偏差,认知指标在不同模型间差异明显(Claude-3 Haiku > Mistral > GPT-3.5-turbo > LLaMA3),甚至同一模型的不同版本也存在差异。这些发现凸显了模型选择在模拟饮食障碍讨论中的关键作用,为理解其复杂互动机制及优化干预设计提供了重要依据。

原文摘要 · Abstract (English)

Eating disorders are complex mental health conditions that affect millions of people around the world. Effective interventions on social media platforms are crucial, yet testing strategies in situ can be risky. We present a novel LLM-driven experimental testbed for simulating and assessing intervention strategies in ED-related discussions. Our framework generates synthetic conversations across multiple platforms, models, and ED-related topics, allowing for controlled experimentation with diverse intervention approaches. We analyze the impact of various intervention strategies on conversation dynamics across four dimensions: intervention type, generative model, social media platform, and ED-related community/topic. We employ cognitive domain analysis metrics, including sentiment, emotions, etc., to evaluate the effectiveness of interventions. Our findings reveal that civility-focused interventions consistently improve positive sentiment and emotional tone across all dimensions, while insight-resetting approaches tend to increase negative emotions. We also uncover significant biases in LLM-generated conversations, with cognitive metrics varying notably between models (Claude-3 Haiku $>$ Mistral $>$ GPT-3.5-turbo $>$ LLaMA3) and even between versions of the same model. These variations highlight the importance of model selection in simulating realistic discussions related to ED. Our work provides valuable information on the complex dynamics of ED-related discussions and the effectiveness of various intervention strategies.

心理安全干预策略大模型模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。