arXiv:2508.17644cs.IR2025-08被引 6

用大模型生成用户画像化查询,让评测更贴近真实搜索差异。

Demographically-Inspired Query Variants Using an LLM

  • 用大模型生成符合不同用户特征的同义查询变体。
  • 不同用户画像下系统排名差异显著,表现评估结果变化明显。
  • 适合关注评测公平性与用户差异的检索系统研究者。

本研究提出一种方法,通过大型语言模型(LLM)生成现有测试集中的查询变体,以反映搜索引擎用户的多样性,呼应早期关于‘理想’测试集的构想。这些变体是语义相同但体现不同用户特征(如语言水平、领域专业度)的替代查询,这些特征在信息检索(IR)文献中已被证明会影响查询构建。实验验证了LLM生成与用户画像匹配的查询变体的能力,并进一步探索其在系统评估中的效用。结果表明,这些变体影响系统排名,且不同用户画像下的系统有效性存在显著差异。该方法为系统评估提供了新视角,可同时观察用户画像对排名的影响及系统在不同用户间的性能波动。

原文摘要 · Abstract (English)

This study proposes a method to diversify queries in existing test collections to reflect some of the diversity of search engine users, aligning with an earlier vision of an 'ideal' test collection. A Large Language Model (LLM) is used to create query variants: alternative queries that have the same meaning as the original. These variants represent user profiles characterised by different properties, such as language and domain proficiency, which are known in the IR literature to influence query formulation. The LLM's ability to generate query variants that align with user profiles is empirically validated, and the variants' utility is further explored for IR system evaluation. Results demonstrate that the variants impact how systems are ranked and show that user profiles experience significantly different levels of system effectiveness. This method enables an alternative perspective on system evaluation where we can observe both the impact of user profiles on system rankings and how system performance varies across users.

信息检索大模型应用用户建模评测方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。