评测大模型学者推荐中用户干预的影响,发现不同操作有不同优劣。
Whose Name Comes Up? II: Benchmarking and Intervention-Based Auditing of LLM-Based Scholar Recommendation
- 构建基准测试框架,同时评估模型和用户干预效果。
- 温度调高降低推荐准确性,约束提示提升多样性但削弱事实性。
- 检索增强生成改善技术质量,但降低多样性和公平性。
大型语言模型(LLMs)现被用于学术专家推荐。现有审计多孤立评估推荐结果,忽略用户在使用时的干预行为。因此,推荐失败(如拒绝、幻觉、覆盖不均)是源于模型本身还是部署决策尚不明确。本文提出 LLMScholarBench,一个用于审计基于 LLM 的学者推荐的基准,联合评估模型架构与用户干预在多个任务下的表现。该基准采用九项指标衡量技术质量与社会代表性。我们在物理学领域实例化该基准,对 22 个 LLM 在温度调节、表示受限提示、以及基于网络搜索的检索增强生成(RAG)条件下进行审计。结果表明,每种干预均带来不同权衡:提高温度会降低有效性、一致性和事实性;表示受限提示可提升多样性但牺牲事实性;而 RAG 主要提升技术质量,但降低多样性和公平性。总体而言,用户干预重塑了性能权衡,而非带来统一改进。LLMScholarBench 可实现跨模型与干预方式的动态可审计性。
原文摘要 · Abstract (English)
Large language models (LLMs) are now used for academic expert recommendation. Existing audits typically evaluate such recommendations in isolation, ignoring end-user inference-time interventions. Thus, it remains unclear whether failures (e.g., refusals, hallucinations, uneven coverage) stem from model choice or deployment decisions. We introduce LLMScholarBench, a benchmark for auditing LLM-based scholar recommendation that jointly evaluates model infrastructure and end-user interventions across multiple tasks. LLMScholarBench measures technical quality and social representation using nine metrics. We instantiate the benchmark in physics expert recommendation and audit 22 LLMs under temperature variation, representation-constrained prompting, and retrieval-augmented generation (RAG) via web search. Our results show that each intervention entails distinct tradeoffs. Higher temperature degrades validity, consistency, and factuality. Representation-constrained prompting improves diversity at the expense of factuality, while RAG primarily improves technical quality while reducing diversity and parity. Overall, end-user interventions reshape trade-offs rather than providing uniform gains. LLMScholarBench makes all these dynamics auditable across models and interventions in LLM-based scholar recommendations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。