arXiv:2602.08716cs.CL2026-02被引 3

构建首个可扩展的多视角评估基准,测试大模型理解多元观点的能力。

PERSPECTRA: A Scalable and Configurable Pluralist Benchmark of Perspectives from Arguments

  • 融合Kialo结构与Reddit语言多样性,生成3810个扩展观点
  • 在100个争议话题上验证模型对观点数量、立场匹配和极性判断的缺陷
  • 适合研究模型偏见、价值观对齐与复杂推理的学者使用

多元主义——在不将不同观点简化为单一立场的前提下,保持对多样视角的互动能力——对于开发真实反映人类差异的大语言模型至关重要。然而该特性在大模型研究中尚未受到充分关注,多数对齐研究也未涉及。辩论类数据源是研究多元主义的天然入口。以往工作依赖在线辩论资源,但受限于高昂的人工标注成本。其他丰富辩论内容的平台如Reddit和Kialo亦具潜力:Reddit提供语言多样性与规模,但缺乏清晰的论点结构;Kialo具备明确的正反论证图谱,却过于简略且脱离自然语境。我们提出PERSPECTRA,一个整合了Kialo论证结构清晰性与Reddit真实讨论语言多样性的多元主义基准。通过受控的检索-扩展流程,构建了覆盖100个争议话题、762个正反立场的3,810个增强型论点。每个观点均拓展为多个自然变体,支持对多元主义的稳健评估。我们定义三项任务:观点计数(识别不同立场)、观点匹配(将支持性论述与原始立场对齐)、极性检测(推断混合论述中的总体倾向)。对先进开源及专有大模型的实验揭示系统性失败,例如高估观点数量、误判让步结构,凸显模型在多元认知理解与推理上的根本挑战。通过结合多样性与结构,PERSPECTRA成为首个可扩展、可配置的基准,用于评估模型在多视角表征、区分与推理方面的表现。

原文摘要 · Abstract (English)

Pluralism, the capacity to engage with diverse perspectives without collapsing them into a single viewpoint, is critical for developing large language models that faithfully reflect human heterogeneity. Yet this characteristic has not been carefully examined in the LLM research community and remains absent from most alignment studies. Debate-oriented sources provide a natural entry point for pluralism research. Previous work builds on online debate sources but remains constrained by costly human validation. Other debate-rich platforms such as Reddit and Kialo also offer promising material: Reddit provides linguistic diversity and scale but lacks clear argumentative structure, while Kialo supplies explicit pro/con graphs but remains overly concise and detached from natural discourse. We introduce PERSPECTRA, a pluralist benchmark that integrates the structural clarity of Kialo debate graphs with the linguistic diversity of real Reddit discussions. Using a controlled retrieval-and-expansion pipeline, we construct 3,810 enriched arguments spanning 762 pro/con stances on 100 controversial topics. Each opinion is expanded to multiple naturalistic variants, enabling robust evaluation of pluralism. We initialise three tasks with PERSPECTRA: opinion counting (identifying distinct viewpoints), opinion matching (aligning supporting stances and discourse to source opinions), and polarity check (inferring aggregate stance in mixed discourse). Experiments with state-of-the-art open-source and proprietary LLMs, highlight systematic failures, such as overestimating the number of viewpoints and misclassifying concessive structures, underscoring the difficulty of pluralism-aware understanding and reasoning. By combining diversity with structure, PERSPECTRA establishes the first scalable, configurable benchmark for evaluating how well models represent, distinguish, and reason over multiple perspectives.

多视角模型评估大模型对齐观点识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。