让用户模拟器可验证,提升推荐与搜索系统评估的透明度和公平性。
Verifiable User Simulation for Search and Recommendation Systems
- 构建七组件可审计框架,使模拟用户行为可追溯、可验证。
- 通过双实验任务检验模拟器在推荐与搜索中的真实表现。
- 适合关注可信AI与算法公平性的检索与推荐研究者。
基于大语言模型的用户模拟被广泛用于评估搜索引擎、推荐系统及检索增强生成流程,但现有模拟器大多缺乏透明性:难以判断模拟用户为何做出特定选择,或该选择是否符合预设用户画像。此外,研究表明,大语言模型可能因用户背景特征(如语言、教育水平、文化背景)产生偏见或歧视性响应,引发对少数群体与弱势群体公平对待的担忧。本半天线下教程提出一种设计-审计框架,将用户模拟器视为由七个可审计组件构成的工程实体:结构化人格档案、任务感知契约、人机行为匹配、可审计轨迹、人格对齐验证、结构化反馈以及更新人格与契约的迭代优化环。通过两个动手实验——推荐列表评估与搜索查询生成——参与者将端到端分析模拟器行为,区分诊断性差异分析与统计验证,并应用真实性、可信度与人口统计偏差检测方法。本教程面向信息检索与推荐系统的研究人员与从业者,关注用户行为模拟与负责任AI。
原文摘要 · Abstract (English)
Large-language-model (LLM) based user simulation is increasingly adopted for evaluating search engines, recommender systems, and retrieval-augmented generation pipelines, yet most simulators remain opaque: it is difficult to determine why a simulated user made a particular choice or whether that choice is consistent with the intended user profile. Compounding this, recent research shows that LLMs can produce biased or discriminatory responses depending on user background characteristics such as language, education level, and cultural context, raising concerns about the equitable treatment of minority and disadvantaged groups. This half-day, in-person tutorial introduces a proposed design-and-audit framework that treats a user simulator as a verifiable engineering artefact composed of seven auditable components - structured Persona, task-aware Contract, matched human-vs-agent Execution, auditable Trace, persona-aligned Verification, structured Feedback, and a Refinement loop that updates personas and contracts. Through two hands-on mini-labs on recommendation-list evaluation and search-query formulation, participants will inspect simulator behaviour end-to-end, distinguish diagnostic discrepancy analysis from statistical validation, and apply checks for fidelity, credibility, and demographic bias. The tutorial targets information retrieval and recommender systems researchers and practitioners interested in user behaviour simulation and responsible AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。