arXiv:2603.17838cs.CL2026-03

构建新闻中以事件为中心的人类价值观理解基准,精准定位行为者与价值方向。

Event-Centric Human Value Understanding in News-Domain Texts: An Actor-Conditioned, Multi-Granularity Benchmark

  • 基于事件结构设计多粒度标注体系,支持细粒度价值归属分析。
  • 涵盖45,793个单元-行为者对和168,061个有向价值实例,覆盖54种细粒度价值类别。
  • 适用于评估大模型在真实新闻中的价值观识别能力,尤其适合研究价值观推理的学者。

现有价值观数据集难以直接支撑事实性新闻中的价值观理解:多数忽略行为者、依赖孤立语句或合成场景,缺乏显式事件结构与价值方向。我们提出NEVU(新闻事件中心价值观理解),一个面向行为者条件化、事件中心化、方向感知的人类价值观识别基准。NEVU评估模型能否从实证证据中识别价值线索、正确归因至行为者并判断价值方向。基于2,865篇英文新闻文章构建,标注涵盖四个语义层级(子事件、基于行为的复合事件、基于故事的复合事件、整篇文章),并对(单元,行为者)对进行细粒度标注,支持局部与复合上下文评估。通过大模型辅助的分阶段验证与针对性人工审核生成标注。采用包含54种细粒度价值和20种粗粒度类别的层次化价值空间,覆盖45,793个单元-行为者对与168,061个有向价值实例。提供专有与开源大模型统一基线,发现轻量级微调(LoRA)可稳定提升开源模型表现,表明尽管NEVU主要作为基准,也支持监督式适配,而不仅是提示评测。数据可用性见附录。

原文摘要 · Abstract (English)

Existing human value datasets do not directly support value understanding in factual news: many are actor-agnostic, rely on isolated utterances or synthetic scenarios, and lack explicit event structure or value direction. We present \textbf{NEVU} (\textbf{N}ews \textbf{E}vent-centric \textbf{V}alue \textbf{U}nderstanding), a benchmark for \emph{actor-conditioned}, \emph{event-centric}, and \emph{direction-aware} human value recognition in factual news. NEVU evaluates whether models can identify value cues, attribute them to the correct actor, and determine value direction from grounded evidence. Built from 2{,}865 English news articles, NEVU organizes annotations at four semantic unit levels (\textbf{Subevent}, \textbf{behavior-based composite event}, \textbf{story-based composite event}, and \textbf{Article}) and labels \mbox{(unit, actor)} pairs for fine-grained evaluation across local and composite contexts. The annotations are produced through an LLM-assisted pipeline with staged verification and targeted human auditing. Using a hierarchical value space with \textbf{54} fine-grained values and \textbf{20} coarse-grained categories, NEVU covers 45{,}793 unit--actor pairs and 168{,}061 directed value instances. We provide unified baselines for proprietary and open-source LLMs, and find that lightweight adaptation (LoRA) consistently improves open-source models, showing that although NEVU is designed primarily as a benchmark, it also supports supervised adaptation beyond prompting-only evaluation. Data availability is described in Appendix~\ref{app:data_code_availability}.

价值观理解新闻分析多粒度标注大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。