Skip to main navigation Skip to search Skip to main content

A Comparison between Human and NLP-based Annotation of Clinical Trial Eligibility Criteria Text Using The OMOP Common Data Model

  • Xinhang Li
  • , Hao Liu
  • , Fabrício Kury
  • , Chi Yuan
  • , Alex Butler
  • , Yingcheng Sun
  • , Anna Ostropolets
  • , Hua Xu
  • , Chunhua Weng

Research output: Contribution to journalArticlepeer-review

Abstract

Human annotations are the established gold standard for evaluating natural language processing (NLP) methods. The goals of this study are to quantify and qualify the disagreement between human and NLP. We developed an NLP system for annotating clinical trial eligibility criteria text and constructed a manually annotated corpus, both following the OMOP Common Data Model (CDM). We analyzed the discrepancies between the human and NLP annotations and their causes (e.g., ambiguities in concept categorization and tacit decisions on inclusion of qualifiers and temporal attributes during concept annotation). This study initially reported complexities in clinical trial eligibility criteria text that complicate NLP and the limitations of the OMOP CDM. The disagreement between and human and NLP annotations may be generalizable. We discuss implications for NLP evaluation.

Original languageEnglish
Pages (from-to)394-403
Number of pages10
JournalAMIA ... Annual Symposium proceedings. AMIA Symposium
Volume2021
StatePublished - 2021

Fingerprint

Dive into the research topics of 'A Comparison between Human and NLP-based Annotation of Clinical Trial Eligibility Criteria Text Using The OMOP Common Data Model'. Together they form a unique fingerprint.

Cite this