arXiv:2602.00913cs.CLcs.AI2026-02被引 3

用价值观层级结构提升单句价值判断,实测校准与集成效果优于硬性分类规则

Do Schwartz Higher-Order Values Help Sentence-Level Human Value Detection? A Study of Hierarchical Gating and Calibration

  • 利用价值观层级结构作为先验知识,通过校准和集成提升检测效果
  • 阈值调优使社会导向与个人导向准确率从0.41提升至0.57,增幅0.16
  • 适合关注小样本、不平衡多标签任务的实用型研究者参考

从单个句子中检测人类价值观是一个稀疏且不平衡的多标签任务。本文在ValueEval'24 / ValuesML(7.4万条英文句子)上研究了舒瓦茨高阶(HO)类别在此设定下的有效性,预算受限。不提出新架构,而是对比监督式Transformer、硬性HO→值管道、存在性→HO→值级联、紧凑指令微调大模型(LLMs)、QLoRA以及低成本改进如阈值调优和小集成。结果显示HO类别可学习:最易的一对(成长 vs 自我保护)达到宏平均F1=0.58。最可靠增益来自校准与集成:阈值调优使社会聚焦与个人聚焦从0.41升至0.57(+0.16),Transformer软投票将成长类从0.286提至0.303,而Transformer+LLM混合模型在自我保护类达到0.353。相反,硬性层级路由未持续提升最终任务性能。紧凑型LLM单独使用时表现不如监督编码器,但在混合集成中可提供有用多样性。在此基准下,HO结构更适合作为归纳偏置,而非刚性路由规则。

原文摘要 · Abstract (English)

Human value detection from single sentences is a sparse, imbalanced multi-label task. We study whether Schwartz higher-order (HO) categories help this setting on ValueEval'24 / ValuesML (74K English sentences) under a compute-frugal budget. Rather than proposing a new architecture, we compare direct supervised transformers, hard HO$\rightarrow$values pipelines, Presence$\rightarrow$HO$\rightarrow$values cascades, compact instruction-tuned large language models (LLMs), QLoRA, and low-cost upgrades such as threshold tuning and small ensembles. HO categories are learnable: the easiest bipolar pair, Growth vs. Self-Protection, reaches Macro-$F_1=0.58$. The most reliable gains come from calibration and ensembling: threshold tuning improves Social Focus vs. Personal Focus from $0.41$ to $0.57$ ($+0.16$), transformer soft voting lifts Growth from $0.286$ to $0.303$, and a Transformer+LLM hybrid reaches $0.353$ on Self-Protection. In contrast, hard hierarchical gating does not consistently improve the end task. Compact LLMs also underperform supervised encoders as stand-alone systems, although they sometimes add useful diversity in hybrid ensembles. Under this benchmark, the HO structure is more useful as an inductive bias than as a rigid routing rule.

价值观检测多标签学习模型集成校准优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。