arXiv:2601.09142cs.LGcs.CL2026-01被引 1

构建首个大规模财报问答逃避行为检测基准,识别企业高管模糊回应。

EvasionBench: A Large-Scale Benchmark for Detecting Managerial Evasion in Earnings Call Q&A

  • 基于2270万组对话构建三等级避责语义分类体系
  • 引入多模型共识标注,达成0.835的标注一致性,优于单模型
  • 发布84K训练集与1K专家标注测试集,适合金融合规与AI审计研究者

本文提出EvasionBench,首个针对财报电话会议问答中管理层逃避行为的大规模检测基准。基于S&P Capital IQ中2270万组问答对构建严谨过滤的数据集,提出包含直接、中间和完全逃避的三级分类体系。采用多模型共识(MMC)标注流程,结合双前沿大模型与三评委多数投票机制,实现0.835的科恩肯德尔一致系数。释放:(1)84,000条平衡训练集,(2)1,000条专家标注黄金标准评估集,(3)基于Qwen3-4B微调的Eva-4B分类器,在宏观F1上达84.9%,超越Claude 4.5、GPT-5.2与Gemini 3 Flash。消融实验验证多模型共识优于单模型标注。EvasionBench填补了金融自然语言处理在管理层沟通逃避检测方面的空白。

原文摘要 · Abstract (English)

We present EvasionBench, a comprehensive benchmark for detecting evasive responses in corporate earnings call question-and-answer sessions. Drawing from 22.7 million Q&A pairs extracted from S&P Capital IQ transcripts, we construct a rigorously filtered dataset and introduce a three-level evasion taxonomy: direct, intermediate, and fully evasive. Our annotation pipeline employs a Multi-Model Consensus (MMC) framework, combining dual frontier LLM annotation with a three-judge majority voting mechanism for ambiguous cases, achieving a Cohen's Kappa of 0.835 on human inter-annotator agreement. We release: (1) a balanced 84K training set, (2) a 1K gold-standard evaluation set with expert human labels, and (3) [Eva-4B], a 4-billion parameter classifier fine-tuned from Qwen3-4B that achieves 84.9% Macro-F1, outperforming Claude 4.5, GPT-5.2, and Gemini 3 Flash. Our ablation studies demonstrate the effectiveness of multi-model consensus labeling over single-model annotation. EvasionBench fills a critical gap in financial NLP by providing the first large-scale benchmark specifically targeting managerial communication evasion.

金融NLP避责检测大模型标注财报分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。