Abstract

In the fast growing of digital technologies, crawlers and search engines face unpredictable challenges. Focused web-crawlers are essential for mining the boundless data available on the internet. Web-Crawlers face indeterminate latency problem due to differences in their response time. The proposed work attempts to optimize the designing and implementation of Focused Web-Crawlers using Master-Slave architecture for Bioinformatics web sources. Focused Crawlers ideally should crawl only relevant pages, but the relevance of the page can only be estimated after crawling the genomics pages. A solution for predicting the page relevance, which is based on Natural Language Processing, is proposed in the paper. The frequency of the keywords on the top ranked sentences of the page determines the relevance of the pages within genomics sources. The proposed solution uses a TextRank algorithm to rank the sentences, as well as ensuring the correct classification of Bioinformatics web page. Finally, the model is validated by being compared with a breadth first search web-crawler. The comparison shows significant reduction in run time for the same harvest rate.

Details

Title
Optimized Focused Web Crawler with Natural Language Processing Based Relevance Measure in Bioinformatics Web Sources
Author
Mani Sekhar, S R 1 ; Siddesh, G M 2 ; Manvi, Sunilkumar S 1 ; Srinivasa, K G 3 

 School of C & IT, Reva University, Bangalore, India 
 Deptarment of Information Science & Engineering, Ramaiah Institute of Technology, Bangalore, India 
 National Institute of Technical Teachers Training and Research, Chandigarh, India 
Pages
146-158
Publication year
2019
Publication date
2019
Publisher
De Gruyter Poland
ISSN
13119702
e-ISSN
13144081
Source type
Scholarly Journal
Language of publication
Bulgarian; English
ProQuest document ID
3155290437
Copyright
© 2019. This work is published under http://creativecommons.org/licenses/by-nc-nd/3.0 (the “License”). Notwithstanding the ProQuest Terms and Conditions, you may use this content in accordance with the terms of the License.