用大模型评估金融言论质量,发现其标注一致性高于人工但存性别偏见。
Argument Quality Annotation and Gender Bias Detection in Financial Communication through Large Language Models
- 用GPT-4o、Llama 3.1、Gemma 2三款大模型标注金融文本质量
- 模型标注一致率超人工,但在不同温度下存在性别偏见
- 提出对抗攻击检测偏见,适合做金融内容安全与公平性研究
金融言论在影响投资决策和公众对金融机构信任方面至关重要,但其质量评估在学术界仍研究不足。本文评估了三款先进大模型(GPT-4o、Llama 3.1、Gemma 2)在金融沟通文本中进行论点质量标注的能力,使用FinArgQuality数据集。贡献有二:一是评估模型标注结果在多次运行中的稳定性,并与人工标注对比;二是设计对抗性攻击以注入性别偏见,分析模型响应,确保公平性与鲁棒性。所有实验在三种温度设置下进行,以考察其对标注稳定性和与人工标签一致性的影响。结果表明,大模型标注的标注者间一致性高于人工,但模型仍表现出不同程度的性别偏见。本文提供多维度分析并提出实用建议,推动未来更可靠、低成本且具备偏见意识的标注方法发展。
原文摘要 · Abstract (English)
Financial arguments play a critical role in shaping investment decisions and public trust in financial institutions. Nevertheless, assessing their quality remains poorly studied in the literature. In this paper, we examine the capabilities of three state-of-the-art LLMs GPT-4o, Llama 3.1, and Gemma 2 in annotating argument quality within financial communications, using the FinArgQuality dataset. Our contributions are twofold. First, we evaluate the consistency of LLM-generated annotations across multiple runs and benchmark them against human annotations. Second, we introduce an adversarial attack designed to inject gender bias to analyse models responds and ensure model's fairness and robustness. Both experiments are conducted across three temperature settings to assess their influence on annotation stability and alignment with human labels. Our findings reveal that LLM-based annotations achieve higher inter-annotator agreement than human counterparts, though the models still exhibit varying degrees of gender bias. We provide a multifaceted analysis of these outcomes and offer practical recommendations to guide future research toward more reliable, cost-effective, and bias-aware annotation methodologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。