Finding molecules and proteins in scientific literature
I have participated in various projects where clients required the analysis of scientific literature to track mentions of molecules or proteins. For instance, the molecule depicted on the right is Aspirin, which remains a trademark of Bayer in certain countries. In academic papers, it may be referenced as acetylsalicylic acid, 2-acetoxybenzenecarboxylic acid, C9H8O4, or by different identifiers like DB00945. Additionally, there may be identifiers related to other molecules or specific variations of a molecule. A frequently observed example in clinical literature is the ERBB2 gene, which plays a crucial role in specific breast cancer types. ERBB2 is also known as Erb-B2 Receptor Tyrosine Kinase, HER2, HER-2, among other designations, many of which also denote the protein produced by the gene. Many of these names resemble common English words and may not always be capitalized in text. Due to these complexities, the task of recognizing names of proteins, genes, and molecules in scientific literature presents significant challenges. I have developed several effective techniques to resolve these ambiguities. Typically, I require a set of annotated examples as a starting point, from which I train a machine learning model to learn and annotate incoming publications. This model can be implemented on the client’s servers, providing daily updates via a dashboard. This enables clients to track literature in real time concerning specific molecules, proteins, or genes and to identify trends proactively.