Text-based video content classification for online video-sharing sites

Chunneng Huang, Tianjun Fu, Hsinchun Chen

Research output: Contribution to journalArticle

38 Citations (Scopus)

Abstract

With the emergence of Web 2.0, sharing personal content, communicating ideas, and interacting with other online users in Web 2.0 communities have become daily routines for online users. User-generated data from Web 2.0 sites provide rich personal information (e.g., personal preferences and interests) and can be utilized to obtain insight about cyber communities and their social networks. Many studies have focused on leveraging usergenerated information to analyze blogs and forums, but few studies have applied this approach to video-sharing Web sites. In this study, we propose a text-based framework for video content classification of online-video sharing Web sites. Different types of user-generated data (e.g., titles, descriptions, and comments) were used as proxies for online videos, and three types of text features (lexical, syntactic, and content-specific features) were extracted. Three feature-based classification techniques (C4.5, Naïve Bayes, and Support Vector Machine) were used to classify videos. To evaluate the proposed framework, user-generated data from candidate videos, which were identified by searching user-given keywords on You Tube, were first collected.Then, a subset of the collected data was randomly selected and manually tagged by users as our experiment data.The experimental results showed that the proposed approach was able to classify online videos based on users' interests with accuracy rates up to 87.2%, and all three types of text features contributed to discriminating videos. Support Vector Machine outperformed C4.5 and Naïve Bayes techniques in our experiments. In addition, our case study further demonstrated that accurate video-classification results are very useful for identifying implicit cyber communities on video-sharing Web sites.

Original languageEnglish (US)
Pages (from-to)891-906
Number of pages16
JournalJournal of the American Society for Information Science and Technology
Volume61
Issue number5
DOIs
StatePublished - May 2010

Fingerprint

Websites
video
Support vector machines
Blogs
Syntactics
Experiments
community
experiment
weblog
social network
candidacy
Web 2.0
Web sites

ASJC Scopus subject areas

  • Software
  • Artificial Intelligence
  • Information Systems
  • Human-Computer Interaction
  • Computer Networks and Communications

Cite this

Text-based video content classification for online video-sharing sites. / Huang, Chunneng; Fu, Tianjun; Chen, Hsinchun.

In: Journal of the American Society for Information Science and Technology, Vol. 61, No. 5, 05.2010, p. 891-906.

Research output: Contribution to journalArticle

@article{f39417905eca4ed3ae67314927a25e14,
title = "Text-based video content classification for online video-sharing sites",
abstract = "With the emergence of Web 2.0, sharing personal content, communicating ideas, and interacting with other online users in Web 2.0 communities have become daily routines for online users. User-generated data from Web 2.0 sites provide rich personal information (e.g., personal preferences and interests) and can be utilized to obtain insight about cyber communities and their social networks. Many studies have focused on leveraging usergenerated information to analyze blogs and forums, but few studies have applied this approach to video-sharing Web sites. In this study, we propose a text-based framework for video content classification of online-video sharing Web sites. Different types of user-generated data (e.g., titles, descriptions, and comments) were used as proxies for online videos, and three types of text features (lexical, syntactic, and content-specific features) were extracted. Three feature-based classification techniques (C4.5, Na{\"i}ve Bayes, and Support Vector Machine) were used to classify videos. To evaluate the proposed framework, user-generated data from candidate videos, which were identified by searching user-given keywords on You Tube, were first collected.Then, a subset of the collected data was randomly selected and manually tagged by users as our experiment data.The experimental results showed that the proposed approach was able to classify online videos based on users' interests with accuracy rates up to 87.2{\%}, and all three types of text features contributed to discriminating videos. Support Vector Machine outperformed C4.5 and Na{\"i}ve Bayes techniques in our experiments. In addition, our case study further demonstrated that accurate video-classification results are very useful for identifying implicit cyber communities on video-sharing Web sites.",
author = "Chunneng Huang and Tianjun Fu and Hsinchun Chen",
year = "2010",
month = "5",
doi = "10.1002/asi.21291",
language = "English (US)",
volume = "61",
pages = "891--906",
journal = "Journal of the Association for Information Science and Technology",
issn = "2330-1635",
publisher = "John Wiley and Sons Ltd",
number = "5",

}

TY - JOUR

T1 - Text-based video content classification for online video-sharing sites

AU - Huang, Chunneng

AU - Fu, Tianjun

AU - Chen, Hsinchun

PY - 2010/5

Y1 - 2010/5

N2 - With the emergence of Web 2.0, sharing personal content, communicating ideas, and interacting with other online users in Web 2.0 communities have become daily routines for online users. User-generated data from Web 2.0 sites provide rich personal information (e.g., personal preferences and interests) and can be utilized to obtain insight about cyber communities and their social networks. Many studies have focused on leveraging usergenerated information to analyze blogs and forums, but few studies have applied this approach to video-sharing Web sites. In this study, we propose a text-based framework for video content classification of online-video sharing Web sites. Different types of user-generated data (e.g., titles, descriptions, and comments) were used as proxies for online videos, and three types of text features (lexical, syntactic, and content-specific features) were extracted. Three feature-based classification techniques (C4.5, Naïve Bayes, and Support Vector Machine) were used to classify videos. To evaluate the proposed framework, user-generated data from candidate videos, which were identified by searching user-given keywords on You Tube, were first collected.Then, a subset of the collected data was randomly selected and manually tagged by users as our experiment data.The experimental results showed that the proposed approach was able to classify online videos based on users' interests with accuracy rates up to 87.2%, and all three types of text features contributed to discriminating videos. Support Vector Machine outperformed C4.5 and Naïve Bayes techniques in our experiments. In addition, our case study further demonstrated that accurate video-classification results are very useful for identifying implicit cyber communities on video-sharing Web sites.

AB - With the emergence of Web 2.0, sharing personal content, communicating ideas, and interacting with other online users in Web 2.0 communities have become daily routines for online users. User-generated data from Web 2.0 sites provide rich personal information (e.g., personal preferences and interests) and can be utilized to obtain insight about cyber communities and their social networks. Many studies have focused on leveraging usergenerated information to analyze blogs and forums, but few studies have applied this approach to video-sharing Web sites. In this study, we propose a text-based framework for video content classification of online-video sharing Web sites. Different types of user-generated data (e.g., titles, descriptions, and comments) were used as proxies for online videos, and three types of text features (lexical, syntactic, and content-specific features) were extracted. Three feature-based classification techniques (C4.5, Naïve Bayes, and Support Vector Machine) were used to classify videos. To evaluate the proposed framework, user-generated data from candidate videos, which were identified by searching user-given keywords on You Tube, were first collected.Then, a subset of the collected data was randomly selected and manually tagged by users as our experiment data.The experimental results showed that the proposed approach was able to classify online videos based on users' interests with accuracy rates up to 87.2%, and all three types of text features contributed to discriminating videos. Support Vector Machine outperformed C4.5 and Naïve Bayes techniques in our experiments. In addition, our case study further demonstrated that accurate video-classification results are very useful for identifying implicit cyber communities on video-sharing Web sites.

UR - http://www.scopus.com/inward/record.url?scp=77951190250&partnerID=8YFLogxK

UR - http://www.scopus.com/inward/citedby.url?scp=77951190250&partnerID=8YFLogxK

U2 - 10.1002/asi.21291

DO - 10.1002/asi.21291

M3 - Article

VL - 61

SP - 891

EP - 906

JO - Journal of the Association for Information Science and Technology

JF - Journal of the Association for Information Science and Technology

SN - 2330-1635

IS - 5

ER -