Skip to main navigation Skip to search Skip to main content

Semantic text mining in early drug discovery for type 2 diabetes

  • Lena K. Hansson
  • , Rasmus Borup Hansen
  • , Sune Pletscher-Frankild
  • , Rudolfs Berzins
  • , Daniel Hvidberg Hansen
  • , Dennis Madseni
  • , Sten B. Christensen
  • , Malene Revsbech Christiansen
  • , Malene Revsbech Christiansen
  • , Ulrika Boulund
  • , Ulrika Boulund
  • , Xenia Asbæk Wolf
  • , Sonny Kim Kjærulff
  • , Sonny Kim Kjærulff
  • , Martijn Van De Bunt
  • , S. ren Tulin
  • , Thomas Skøt Jensen
  • , Rasmus Wernersson*
  • , Rasmus Wernersson
  • , Jan Nygaard Jensen
  • Jan Nygaard Jensen
*Corresponding author for this work
  • Novo Nordisk Foundation
  • Intomics A/S
  • University of Copenhagen
  • University of Amsterdam
  • Technical University of Denmark
  • Boehringer Ingelheim GmbH

Research output: Contribution to journalArticleAcademicpeer-review

29 Downloads (Pure)

Abstract

Background Surveying the scientific literature is an important part of early drug discovery; and with the ever-increasing amount of biomedical publications it is imperative to focus on the most interesting articles. Here we present a project that highlights new understanding (e.g. recently discovered modes of action) and identifies potential drug targets, via a novel, data-driven text mining approach to score type 2 diabetes (T2D) relevance. We focused on monitoring trends and jumps in T2D relevance to help us be timely informed of important breakthroughs. Methods We extracted over 7 million n-grams from PubMed abstracts and then clustered around 240,000 linked to T2D into almost 50,000 T2D relevant 'semantic concepts'. To score papers, we weighted the concepts based on co-mentioning with core T2D proteins. A protein's T2D relevance was determined by combining the scores of the papers mentioning it in the five preceding years. Each week all proteins were ranked according to their T2D relevance. Furthermore, the historical distribution of changes in rank from one week to the next was used to calculate the significance of a change in rank by T2D relevance for each protein. Results We show that T2D relevant papers, even those not mentioning T2D explicitly, were prioritised by relevant semantic concepts. Well known T2D proteins were therefore enriched among the top scoring proteins. Our 'high jumpers' identified important past developmentsin the apprehension of how certain key proteins relate to T2D, indicating that our method will make us aware of future breakthroughs. In summary, this project facilitated keeping up with current T2D research by repeatedly providing short lists of potential novel targets into our early drug discovery pipeline.
Original languageEnglish
Article numbere0233956
JournalPLoS ONE
Volume15
Issue number6
DOIs
Publication statusPublished - 1 Jun 2020
Externally publishedYes

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 3 - Good Health and Well-being
    SDG 3 Good Health and Well-being

Fingerprint

Dive into the research topics of 'Semantic text mining in early drug discovery for type 2 diabetes'. Together they form a unique fingerprint.

Cite this