评测大模型在不同社区风格下的可控生成能力
STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models
- 基于30组对比Reddit社区构建评测数据集
- 13个模型平均准确率仅65%,距人类81%有明显差距
- 适合研究模型对多元文化与意识形态的适应性
可调节性(Steerability)指大语言模型根据不同群体特有的规范、视角和表达风格调整输出的能力,对实际应用至关重要但长期缺乏评估。本文提出Steer-Bench,一个基于对比性Reddit社区的评测基准。该基准覆盖19个领域中的30组对立子论坛,包含超10,000条指令-响应对,以及经验证的5,500道多选题及其银标签,用于测试模型对多样社区规范的对齐程度。对13个主流LLM的评估显示,人类专家使用银标签可达81%准确率,而表现最佳的模型仅达约65%,且在部分领域与人类差距超过15个百分点,揭示了当前模型在社区敏感性调节上的显著不足。Steer-Bench可用于系统评估模型理解群体特定指令的能力、抵御对抗性调节攻击的鲁棒性,以及准确呈现多元文化与意识形态观点的能力。
原文摘要 · Abstract (English)
Steerability, or the ability of large language models (LLMs) to adapt outputs to align with diverse community-specific norms, perspectives, and communication styles, is critical for real-world applications but remains under-evaluated. We introduce Steer-Bench, a benchmark for assessing population-specific steering using contrasting Reddit communities. Covering 30 contrasting subreddit pairs across 19 domains, Steer-Bench includes over 10,000 instruction-response pairs and validated 5,500 multiple-choice question with corresponding silver labels to test alignment with diverse community norms. Our evaluation of 13 popular LLMs using Steer-Bench reveals that while human experts achieve an accuracy of 81% with silver labels, the best-performing models reach only around 65% accuracy depending on the domain and configuration. Some models lag behind human-level alignment by over 15 percentage points, highlighting significant gaps in community-sensitive steerability. Steer-Bench is a benchmark to systematically assess how effectively LLMs understand community-specific instructions, their resilience to adversarial steering attempts, and their ability to accurately represent diverse cultural and ideological perspectives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。