Improved Chemical Text Mining of Patents with Infinite Dictionaries and Automatic Spelling Correction

Domenii publicaţii > Chimie + Tipuri publicaţii > Articol în revistã ştiinţificã

Autori: Sayle, R.; Xie, P.H.; Muresan, S

Editorial: J. Chem. Inf. Model., 52 (1), p.51-62, 2012.

Rezumat:

The text mining of patents of pharmaceutical interest poses a number of unique challenges not encountered in other fields of text mining. Unlike fields, such as bioinformatics, where the number of terms of interest is enumerable and essentially static, systematic chemical nomenclature can describe an infinite number of molecules. Hence, the dictionary- and ontology-based techniques that are commonly used for gene names, diseases, species, etc., have limited utility when searching for novel therapeutic compounds in patents. Additionally, the length and the composition of IUPAC-like names make them more susceptible to typographic problems: OCR failures, human spelling errors, and hyphenation and line breaking issues. This work describes a novel technique, called CaffeineFix, designed to efficiently identify chemical names in free text, even in the presence of typographical errors. Corrected chemical names are generated as input for name-to-structure software. This forms a preprocessing pass, independent of the name-to-structure software used, and is shown to greatly improve the results of chemical text mining in our study.

Cuvinte cheie: chemical text mining; infinite dictionaries; CaffeineFix

URL: http://dx.doi.org/10.1021/ci200463r

Sorel Muresan
ianuarie 25, 2012
Niciun comentariu

Staff Login

Login Id

Password

Solicitam publicarea versiunii romanesti oficiale a codului ALLEA

Dobandirea calitatii de autor prin abuz de autoritate—solicitare de clarificare a exonerarii de la initiatorii legilor

Comunicat Ad-Astra: evoluția publicațiilor științifice ale României (2020–2025) și schimbările structurale ale canalelor editoriale

Inscriere cercetatori

Premii Ad Astra

Improved Chemical Text Mining of Patents with Infinite Dictionaries and Automatic Spelling Correction

Întrebări frecvente

Contacteaza-ne

Ajută-ne!

Staff Login

Login Id

Password

Search

Solicitam publicarea versiunii romanesti oficiale a codului ALLEA

Dobandirea calitatii de autor prin abuz de autoritate—solicitare de clarificare a exonerarii de la initiatorii legilor

Comunicat Ad-Astra: evoluția publicațiilor științifice ale României (2020–2025) și schimbările structurale ale canalelor editoriale

Inscriere cercetatori

Premii Ad Astra

Improved Chemical Text Mining of Patents with Infinite Dictionaries and Automatic Spelling Correction

Share

Întrebări frecvente

Contacteaza-ne

Ajută-ne!