关注人机交互长期影响,推动AI评估从短期生成转向行为变化监测
Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions

- 借鉴社会科学研究方法,构建长期追踪人类行为变化的测量体系
- 发现长期互动可能引发认知、情感等深层改变,短期测试难以察觉
- 适合关注AI安全、用户体验与伦理的学者及产品开发者
语言模型因其类人特性及快速融入用户日常生活,正成为一种新型技术。这种特性组合可能带来长期风险——如认知、发展与社会情感层面的人类变化,这些风险在短期交互中不易显现,却可能对用户产生持久影响。这催生了自然语言处理领域的新使命:从静态、短期的文本生成评估,转向对行为变化的长期测量,以实现对人-模型交互的历时性理解。本文借鉴社会科学研究中用于分析纵向数据的关键测量方法,探讨如何将计算方法与之结合,不仅识别长期安全风险,还可引导模型开发向有益于用户的方向演进。通过建模人类行为随交互时间的变化,可实现问题行为的在线检测,而非事后补救,应纳入对齐框架以缓解长期风险。
原文摘要 · Abstract (English)
Language models have taken on the role of a very new type of technology, by virtue of their "human-ness" and rapid integration into users' daily lives. This combination of features can introduce longitudinal risks---cognitive, developmental and socio-affective changes in humans---that might not surface during a short-term interaction, but can have lasting long-term effects on users. This forms the basis of a critical new mission for NLP: to pivot from static, short-term evaluations of text generations to long-term measurements of behavioral changes, towards a diachronic understanding of human-model interactions. In this work, we draw from measurements used in social science fields that are crucial to understand emergent phenomena in longitudinal data. We discuss how computational methods in the field of NLP need to be combined with such measurements, not only to understand long-term safety risks of human-model interactions, but to help steer model development towards positive rather than negative outcomes for users. This ability to model human behavioral shifts as a function of model interactions can facilitate online rather than post-hoc detection of problematic behaviors, and should be leveraged in alignment frameworks to mitigate long-term risks in users.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。