Biomedical Information Extraction
113 items from caufieldjh/awesome-bioie ★468
-
BioGPT github.com
paper - A GPT-2 model pre-trained on 15 million PubMed abstracts, along with fine-tuned versions for several biomedical tasks.
-
mimic-code github.com
Code associated with the MIMIC-III dataset (see below). Includes some helpful tutorials.
-
ScispaCy github.com
paper - A version of the spaCy framework for scientific and biomedical documents.
-
-
-
-
BioBERT github.com
paper - code - A PubMed and PubMed Central-trained version of the BERT language model.
-
-
-
medaCy github.com
A system for building predictive medical natural language processing models. Built on the spaCy framework.
-
-
-
CORD-19 github.com
A corpus of scholarly manuscripts concerning COVID-19. Articles are primarily from PubMed Central and preprint servers, though the set also includes metadata on papers without full-text availability.
-
SemEHR github.com
paper - an IE infrastructure for electronic health records (EHR). Built on the CogStack project.
-
CRAFT github.com
paper - 67 full-text biomedical articles annotated in a variety of ways, including for concepts and coreferences. Now on version 5, including annotations linking concepts to the MONDO disease ontology.
-
II-Commons github.com
Daily-updated skill and CLI for deterministic retrieval across arXiv, PubMed/PMC, and supported US policy corpora.
-
Pubrunner github.com
A framework for running text mining tools on the newest set(s) of documents from PubMed.
-
-
DeepPhe github.com
A system for processing documents describing cancer presentations. Based on cTAKES (see above).
-
-
Comparative Toxicogenomics Database ctdbase.org
paper - A database of manually curated associations between chemicals, gene products, phenotypes, diseases, and environmental exposures. Useful for assembling ontologies of the related concepts, such as types of chemicals.
-
Biopython biopython.org
paper - code - Python tools primarily intended for bioinformatics and computational molecular biology purposes, but also a convenient way to obtain data, including documents/abstracts from PubMed (see Chapter 9 of the documentation).
-
Apache cTAKES ctakes.apache.org
paper - code - A system for processing the text in electronic medical records. Widely used and open source.
-
unmiri-ngs-fhir-schema github.com
Apache-2.0 JSON Schema (Draft 2020-12) API contract for cross-vendor somatic NGS interpretation output (Foundation Medicine, Tempus, Caris, Guardant), aligned with the HL7 FHIR Genomics IG. A standards-aligned target representation for biomedical information-extraction…
-
Zaklab zaklab.org
Group led by Dr. Isaac Kohane at Harvard Medical School's Department of Biomedical Informatics (Dr. Kohane is also a steward of the n2c2 (formerly i2b2) datasets - see Datasets below).
-
Word Sense Disambiguation (WSD) wsd.nlm.nih.gov
paper - 203 ambiguous words and 37,888 automatically extracted instances of their use in biomedical research publications. Requires UTS account.
-
BioUML wiki.biouml.org
paper - An architecture for biomedical data analysis, integration, and visualization. Conceptually based on the visual modeling language UML.
-
TurkuNLP turkunlp.org
Based at the University of Turku and concerned with NLP in general with a focus on BioNLP and clinical applications.
-
EHR DREAM Challenge synapse.org
Held along with several other more bioinformatics-focused challenges, this challenge opened in October 2019 and focuses on using electronic health record data to predict patient mortality. Uses a synthetic data set rather than real EHR contents.
-
BioCreAtIvE 1 sourceforge.net
paper - 15,000 sentences (10,000 training and 5,000 test) annotated for protein and gene names. 1,000 full text biomedical research articles annotated with protein names and Gene Ontology terms.
- next page of items loading…