Abstract

Translate

The explosive increase of biomedical literature has made information extraction an increasingly important tool for biomedical research. A fundamental task is the recognition of biomedical named entities in text (BNER) such as genes/proteins, diseases, and species. Recently, a domain-independent method based on deep learning and statistical word embeddings, called long short-term memory network-conditional random field (LSTM-CRF), has been shown to outperform state-of-the-art entity-specific BNER tools. However, this method is dependent on gold-standard corpora (GSCs) consisting of hand-labeled entities, which tend to be small but highly reliable. An alternative to GSCs are silver-standard corpora (SSCs), which are generated by harmonizing the annotations made by several automatic annotation systems. SSCs typically contain more noise than GSCs but have the advantage of containing many more training examples. Ideally, these corpora could be combined to achieve the benefits of both, which is an opportunity for transfer learning. In this work, we analyze to what extent transfer learning improves upon state-of-the-art results for BNER. We demonstrate that transferring a deep neural network (DNN) trained on a large, noisy SSC to a smaller, but more reliable GSC significantly improves upon state-of-the-art results for BNER. Compared to a state-of-the-art baseline evaluated on 23 GSCs covering four different entity classes, transfer learning results in an average reduction in error of approximately 11%. We found transfer learning to be especially beneficial for target data sets with a small number of labels (approximately 6000 or less).

Footnotes

* New manuscript version contains all changes made after peer-review and matches the to-be published copied manuscript. The largest changes include additional datasets added to the experiments and an appendix table detailing the performance results of the neural network on all datasets used in the study.

Details

Title

Transfer learning for biomedical named entity recognition with neural networks.

Author

Giorgi, John M; Bader, Gary

University/institution

Cold Spring Harbor Laboratory Press

Section

New Results

Publication year

2018

Publication date

May 30, 2018

Publisher

Cold Spring Harbor Laboratory Press

ISSN

2692-8205

Source type

Working Paper

Language of publication

English

DOI

https://doi.org/10.1101/262790

ProQuest document ID

2071228089

�� 2018. This article is published under https://creativecommons.org/publicdomain/zero/1.0/ (��the License��). Notwithstanding the ProQuest Terms and Conditions, you may use this content in accordance with the terms of the License.

Transfer learning for biomedical named entity recognition with neural networks.

Jump to:

Abstract

Details

Suggested sources