A focused crawler for dark web forums

Tianjun Fu, Ahmed Abbasi, Hsinchun Chen

Research output: Contribution to journalArticle

53 Citations (Scopus)

Abstract

The unprecedented growth of the Internet has given rise to the Dark Web, the problematic facet of the Web associated with cybercrime, hate, and extremism. Despite the need for tools to collect and analyze Dark Web forums, the covert nature of this part of the Internet makes traditional Web crawling techniques insufficient for capturing such content. In this study, we propose a novel crawling system designed to collect Dark Web forum content. The system uses a human-assisted accessibility approach to gain access to Dark Web forums. Several URL ordering features and techniques enable efficient extraction of forum postings.The system also includes an incremental crawler coupled with a recall-improvement mechanism intended to facilitate enhanced retrieval and updating of collected content. Experiments conducted to evaluate the effectiveness of the human-assisted accessibility approach and the recall-improvement-based, incremental-update procedure yielded favorable results. The human-assisted approach significantly improved access to Dark Web forums while the incremental crawler with recall improvement also outperformed standard periodic-and incremental-update approaches. Using the system, we were able to collect over 100 DarkWeb forums from three regions. A case study encompassing link and content analysis of collected forums was used to illustrate the value and importance of gathering and analyzing content from such online communities.

Original languageEnglish (US)
Pages (from-to)1213-1231
Number of pages19
JournalJournal of the American Society for Information Science and Technology
Volume61
Issue number6
DOIs
StatePublished - Jun 2010

Fingerprint

Internet
World Wide Web
Websites
hate
radicalism
internet community
content analysis
Experiments
experiment
Values
Incremental

ASJC Scopus subject areas

  • Software
  • Artificial Intelligence
  • Information Systems
  • Human-Computer Interaction
  • Computer Networks and Communications

Cite this

A focused crawler for dark web forums. / Fu, Tianjun; Abbasi, Ahmed; Chen, Hsinchun.

In: Journal of the American Society for Information Science and Technology, Vol. 61, No. 6, 06.2010, p. 1213-1231.

Research output: Contribution to journalArticle

@article{d9637b6ea5374aa5b6c06fdf01da07c0,
title = "A focused crawler for dark web forums",
abstract = "The unprecedented growth of the Internet has given rise to the Dark Web, the problematic facet of the Web associated with cybercrime, hate, and extremism. Despite the need for tools to collect and analyze Dark Web forums, the covert nature of this part of the Internet makes traditional Web crawling techniques insufficient for capturing such content. In this study, we propose a novel crawling system designed to collect Dark Web forum content. The system uses a human-assisted accessibility approach to gain access to Dark Web forums. Several URL ordering features and techniques enable efficient extraction of forum postings.The system also includes an incremental crawler coupled with a recall-improvement mechanism intended to facilitate enhanced retrieval and updating of collected content. Experiments conducted to evaluate the effectiveness of the human-assisted accessibility approach and the recall-improvement-based, incremental-update procedure yielded favorable results. The human-assisted approach significantly improved access to Dark Web forums while the incremental crawler with recall improvement also outperformed standard periodic-and incremental-update approaches. Using the system, we were able to collect over 100 DarkWeb forums from three regions. A case study encompassing link and content analysis of collected forums was used to illustrate the value and importance of gathering and analyzing content from such online communities.",
author = "Tianjun Fu and Ahmed Abbasi and Hsinchun Chen",
year = "2010",
month = "6",
doi = "10.1002/asi.21323",
language = "English (US)",
volume = "61",
pages = "1213--1231",
journal = "Journal of the Association for Information Science and Technology",
issn = "2330-1635",
publisher = "John Wiley and Sons Ltd",
number = "6",

}

TY - JOUR

T1 - A focused crawler for dark web forums

AU - Fu, Tianjun

AU - Abbasi, Ahmed

AU - Chen, Hsinchun

PY - 2010/6

Y1 - 2010/6

N2 - The unprecedented growth of the Internet has given rise to the Dark Web, the problematic facet of the Web associated with cybercrime, hate, and extremism. Despite the need for tools to collect and analyze Dark Web forums, the covert nature of this part of the Internet makes traditional Web crawling techniques insufficient for capturing such content. In this study, we propose a novel crawling system designed to collect Dark Web forum content. The system uses a human-assisted accessibility approach to gain access to Dark Web forums. Several URL ordering features and techniques enable efficient extraction of forum postings.The system also includes an incremental crawler coupled with a recall-improvement mechanism intended to facilitate enhanced retrieval and updating of collected content. Experiments conducted to evaluate the effectiveness of the human-assisted accessibility approach and the recall-improvement-based, incremental-update procedure yielded favorable results. The human-assisted approach significantly improved access to Dark Web forums while the incremental crawler with recall improvement also outperformed standard periodic-and incremental-update approaches. Using the system, we were able to collect over 100 DarkWeb forums from three regions. A case study encompassing link and content analysis of collected forums was used to illustrate the value and importance of gathering and analyzing content from such online communities.

AB - The unprecedented growth of the Internet has given rise to the Dark Web, the problematic facet of the Web associated with cybercrime, hate, and extremism. Despite the need for tools to collect and analyze Dark Web forums, the covert nature of this part of the Internet makes traditional Web crawling techniques insufficient for capturing such content. In this study, we propose a novel crawling system designed to collect Dark Web forum content. The system uses a human-assisted accessibility approach to gain access to Dark Web forums. Several URL ordering features and techniques enable efficient extraction of forum postings.The system also includes an incremental crawler coupled with a recall-improvement mechanism intended to facilitate enhanced retrieval and updating of collected content. Experiments conducted to evaluate the effectiveness of the human-assisted accessibility approach and the recall-improvement-based, incremental-update procedure yielded favorable results. The human-assisted approach significantly improved access to Dark Web forums while the incremental crawler with recall improvement also outperformed standard periodic-and incremental-update approaches. Using the system, we were able to collect over 100 DarkWeb forums from three regions. A case study encompassing link and content analysis of collected forums was used to illustrate the value and importance of gathering and analyzing content from such online communities.

UR - http://www.scopus.com/inward/record.url?scp=77952994598&partnerID=8YFLogxK

UR - http://www.scopus.com/inward/citedby.url?scp=77952994598&partnerID=8YFLogxK

U2 - 10.1002/asi.21323

DO - 10.1002/asi.21323

M3 - Article

AN - SCOPUS:77952994598

VL - 61

SP - 1213

EP - 1231

JO - Journal of the Association for Information Science and Technology

JF - Journal of the Association for Information Science and Technology

SN - 2330-1635

IS - 6

ER -