欢迎来到《四川大学学报(医学版)》

癫痫临床文本信息的结构化标注与抽取研究

Structured Annotation and Information Extraction of Epilepsy Clinical Texts

  • 摘要:
    目的 面向癫痫专科临床及科研需求,构建一套细粒度、高覆盖的中文临床文本命名实体标注体系,并验证其在命名实体识别(named entity recognition,NER)任务中的有效性。
    方法 从疾病、病程时间、临床表现、医疗检查、非药物治疗、用药及影响因素等7大维度设计3级标注体系,共制定25类标签,明确标签边界与特殊表达处理规则。以四川大学华西医院2009~2023年癫痫住院患者脱敏病历为数据来源,由3名具有癫痫临床背景的标注员完成804例高质量语料标注,采用双人复标、专家仲裁与实体级标注者一致性(inter-annotator agreement,IAA)评估保障标注质量。以五种中文医学预训练语言模型(Base-BERT、chinese-bert、chinese-Roberta、MC-BERT、MedBERT)分别结合BiLSTM-CRF与GlobalPointer两种序列标注框架,共10种组合方案进行NER性能验证,并在外部文献病例集上开展跨域泛化评估。
    结果 最终语料库涵盖25类癫痫相关实体,标注实例804条,实体总数28 400个;三位标注员IAA为0.86~0.88,标注一致性达到计算语言学通行标准。NER验证中,GlobalPointer框架表现优于BiLSTM-CRF,整体Micro-F1为0.906、Macro-F1为0.760;高频核心实体(如发作症状、药物名称、时间信息)F1值均超过0.90;跨域文献病例验证中高频标签F1值维持在0.80以上,低频标签(如fac-inc-amb、trt、dru-adv)F1值为0.31~0.57,主要受样本量有限与表达多样性影响。
    结论 本研究构建的癫痫专科中文临床标注体系覆盖范围广、粒度细、一致性高,可有效支持NER模型的训练与评估,为癫痫病历结构化分析及下游智能诊疗工具的开发提供了可复用的语料基础。

     

    Abstract:
    Objective To develop a fine-grained, highly comprehensive Chinese clinical texts named entity annotation schema tailored to the needs of epilepsy specialty clinical practice and research, and to validate its effectiveness in named entity recognition (NER) tasks.
    Methods A three-level annotation schema covering 25 entity types was designed across seven major dimensions, including disease, disease course timeline, clinical manifestations, medical examinations, non-pharmacological treatments, medication, and influencing factors, with explicit label boundary definitions and rules for handling special expressions. De-identified inpatient records of epilepsy patients admitted to West China Hospital, Sichuan University from 2009 to 2023 served as the data source. Three annotators with epilepsy clinical backgrounds completed high-quality annotation of 804 cases, with annotation quality ensured through double annotation, expert arbitration, and entity-level inter-annotator agreement (IAA) evaluation. NER performance was validated using 10 model combinations comprising five Chinese medical pre-trained language models (Base-BERT, chinese-bert, chinese-Roberta, MC-BERT, MedBERT) paired with two sequence labeling frameworks (BiLSTM-CRF and GlobalPointer), with additional cross-domain generalization evaluation on an external case dataset established on the basis of published literature.
    Results The final corpus contains 25 categories of epilepsy-related entities, 804 annotated cases, and a total of 28 400 entities. The IAA among the three annotators ranged from 0.86 to 0.88, indicating annotation consistency that met the accepted standards in computational linguistics. In NER validation, the GlobalPointer framework outperformed BiLSTM-CRF, achieving an overall Micro-F1 score of 0.906 and Macro-F1 score of 0.760. High-frequency core entities (e.g., seizure symptoms, drug names, and temporal information) all yielded F1 scores exceeding 0.90. In cross-domain validation on literature-based cases, high-frequency entity F1 scores remained above 0.80, while low-frequency entities (e.g., factors with incomplete/ambiguous induction fac-inc-amb, treatment information trt, and adverse drug reactions dru-adv) achieved F1 scores of 0.31-0.57, primarily attributable to limited sample size and high linguistic variability.
    Conclusion The epilepsy-specific Chinese clinical annotation schema developed in this study demonstrates broad coverage, fine granularity, and high inter-annotator consistency. It effectively supports the training and evaluation of NER models and provides a reusable corpus foundation for the structured analysis of epilepsy medical records and the development of downstream intelligent diagnostic and therapeutic tools.

     

/

返回文章
返回