arXiv:2511.09854cs.CL2025-11AAAI被引 1

让大模型更懂法律金融术语,提升细粒度语义区分能力

TermGPT: Multi-Level Contrastive Fine-Tuning for Terminology Adaptation in Legal and Financial Domain

  • 构建句子图谱,结合上下文与结构生成正负样本
  • 多层级对比学习提升句子与词级语义区分度
  • 首个基于监管文件的金融术语数据集,支持可靠评估

大型语言模型在文本生成任务中表现优异,但其嵌入空间常存在各向同性问题,导致法律与金融领域特定术语的语义区分能力弱。这种术语层面表示的不足会严重影响法律判决预测、金融风险分析等下游任务,而这些任务对细微语义差异极为敏感。为此,我们提出 TermGPT,一种用于术语适配的多层级对比微调框架。首先构建句子图谱以捕捉语义与结构关系,并基于上下文和拓扑线索生成语义一致但可区分的正负样本;随后在句子与词元层面设计多层级对比学习,增强全局上下文理解与细粒度术语判别能力。为支持稳健评估,我们构建了首个源自官方监管文件的金融术语数据集。实验表明,TermGPT 在金融与法律领域的术语判别任务中优于现有基线方法。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated impressive performance in text generation tasks; however, their embedding spaces often suffer from the isotropy problem, resulting in poor discrimination of domain-specific terminology, particularly in legal and financial contexts. This weakness in terminology-level representation can severely hinder downstream tasks such as legal judgment prediction or financial risk analysis, where subtle semantic distinctions are critical. To address this problem, we propose TermGPT, a multi-level contrastive fine-tuning framework designed for terminology adaptation. We first construct a sentence graph to capture semantic and structural relations, and generate semantically consistent yet discriminative positive and negative samples based on contextual and topological cues. We then devise a multi-level contrastive learning approach at both the sentence and token levels, enhancing global contextual understanding and fine-grained terminology discrimination. To support robust evaluation, we construct the first financial terminology dataset derived from official regulatory documents. Experiments show that TermGPT outperforms existing baselines in term discrimination tasks within the finance and legal domains.

术语识别对比学习法律AI金融NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。