构建首个多维主观性标注数据集,评估财报问答中回答的语气与质量。
SubjECTive-QA: Measuring Subjectivity in Earnings Call Transcripts' QA Through Six-Dimensional Feature Analysis
- 从财报问答中提取六维主观特征,人工标注49,446条数据
- 大模型在清晰、相关等客观特征上表现接近,主观特征差距达10.01%
- 数据集可迁移至白宫简报等场景,跨领域适用性强
事实核查广泛研究虚假信息中的客观错误,但更隐蔽的误导来自回答虽真实却缺乏清晰性与相关性等主观特质。此类问题在金融、政治等正式问答场景中普遍存在。然而,缺乏多维度主观特征的人工标注数据集。为此,本文提出SubjECTive-QA,基于公司财报问答(ECTs)构建首个六维主观特征标注数据集,涵盖断言性、谨慎性、乐观性、具体性、清晰性和相关性,共49,446条长文本问答对。实验表明,最佳预训练模型RoBERTa-base在低主观性特征(如相关性、清晰性)上与Llama-3-70b-Chat性能相近,平均加权F1差异仅2.17%;而在高主观性特征(如具体性、断言性)上差异达10.01%。此外,在白宫简报和记者会问答上测试,最佳模型平均加权F1为65.97%,证明其跨领域泛化能力。数据集已开源,许可协议为CC BY 4.0。
原文摘要 · Abstract (English)
Fact-checking is extensively studied in the context of misinformation and disinformation, addressing objective inaccuracies. However, a softer form of misinformation involves responses that are factually correct but lack certain features such as clarity and relevance. This challenge is prevalent in formal Question-Answer (QA) settings such as press conferences in finance, politics, sports, and other domains, where subjective answers can obscure transparency. Despite this, there is a lack of manually annotated datasets for subjective features across multiple dimensions. To address this gap, we introduce SubjECTive-QA, a human annotated dataset on Earnings Call Transcripts' (ECTs) QA sessions as the answers given by company representatives are often open to subjective interpretations and scrutiny. The dataset includes 49,446 annotations for long-form QA pairs across six features: Assertive, Cautious, Optimistic, Specific, Clear, and Relevant. These features are carefully selected to encompass the key attributes that reflect the tone of the answers provided during QA sessions across different domain. Our findings are that the best-performing Pre-trained Language Model (PLM), RoBERTa-base, has similar weighted F1 scores to Llama-3-70b-Chat on features with lower subjectivity, such as Relevant and Clear, with a mean difference of 2.17% in their weighted F1 scores. The models perform significantly better on features with higher subjectivity, such as Specific and Assertive, with a mean difference of 10.01% in their weighted F1 scores. Furthermore, testing SubjECTive-QA's generalizability using QAs from White House Press Briefings and Gaggles yields an average weighted F1 score of 65.97% using our best models for each feature, demonstrating broader applicability beyond the financial domain. SubjECTive-QA is publicly available under the CC BY 4.0 license
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。