用多个提示优化的AI代理模拟不同人群行为,更真实还原人类多样性。
Prompt Optimization Across Multiple Agents for Representing Diverse Human Populations
- 通过少量人类示例微调多个LLM代理,实现行为多样性建模。
- 在众包与教育场景中,代理群表现优于基线方法。
- 适合需要真实人类行为模拟的研究与应用,如教育评测。
获取大规模人类响应成本高昂,大语言模型(LLMs)成为人类行为的替代方案和有前景的代理。然而,已有研究表明,LLMs常生成同质化输出,难以捕捉人类观点与行为的丰富多样性。因此,我们提出一种新框架,通过构建一组代理来共同表征特定人群的多样性,而非依赖单一模型。每个代理是经少量人类示范(任务-回应对)通过上下文学习引导的LLM。核心挑战是从指数级庞大的可能代理空间中选出代表性集合,我们从子模优化视角解决该问题,开发了在时间复杂度与性能保障间权衡不同的方法。在众包和教育领域的大量实验表明,我们的方法构建的代理群比基线更有效地代表人类群体。此外,新任务上的行为分析显示,这些代理重现了其对应学生与标注者的行为模式与视角。
原文摘要 · Abstract (English)
The difficulty and expense of obtaining large-scale human responses make Large Language Models (LLMs) an attractive alternative and a promising proxy for human behavior. However, prior work shows that LLMs often produce homogeneous outputs that fail to capture the rich diversity of human perspectives and behaviors. Thus, rather than trying to capture this diversity with a single LLM agent, we propose a novel framework to construct a set of agents that collectively capture the diversity of a given human population. Each agent is an LLM whose behavior is steered by conditioning on a small set of human demonstrations (task-response pairs) through in-context learning. The central challenge is therefore to select a representative set of LLM agents from the exponentially large space of possible agents. We tackle this selection problem from the lens of submodular optimization. In particular, we develop methods that offer different trade-offs regarding time complexity and performance guarantees. Extensive experiments in crowdsourcing and educational domains demonstrate that our approach constructs agents that more effectively represent human populations compared to baselines. Moreover, behavioral analyses on new tasks show that these agents reproduce the behavior patterns and perspectives of the students and annotators they are designed to represent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。