Short Text Classification Using An Enhanced Term Weighting Scheme And Filter-Wrapper Feature Selection

Social networks and their usage in everyday life have caused an explosion in the amount of short electronic documents. Social networks, such as Twitter, are common mechanisms through which people can share information. The utilization of data that are available through social media for many applicat...

全面介绍

Saved in:
书目详细资料
主要作者: Alsmadi, Issa Mohammad Ibrahim
格式: Thesis
语言:English
出版: 2018
主题:
在线阅读:http://eprints.usm.my/46679/1/short%20text%20classification%20using%20an%20enhanced%20term%20weghiting%20scheme%20and%20filter-wrapper%20feture%20selection24.pdf
标签: 添加标签
没有标签, 成为第一个标记此记录!
实物特征
总结:Social networks and their usage in everyday life have caused an explosion in the amount of short electronic documents. Social networks, such as Twitter, are common mechanisms through which people can share information. The utilization of data that are available through social media for many applications is gradually increasing. Redundancy and noise in short texts are common problems in social media and in different applications that use short text. However, the shortness and high sparsity of short text lead to poor classification performance. Employing a powerful short-text classification method significantly affects many applications in terms of efficiency enhancement. This research aims to investigate and develop solutions for feature discrimination and selection in short texts classification. For feature discrimination, we introduce a term weighting approach namely, simple supervised weight (SW), which considers the special nature of short text in terms of term strength and distribution. To address the drawbacks of using existing feature selection with short text, this thesis proposes a filter-wrapper feature selection approach. In the first stage, we propose an adaptive filter-based feature selection method that is derived from the odd ratio method, used in reducing the dimensionality of feature space. In the second stage, grey wolf optimization (GWO) algorithm, a new heuristic search algorithm, uses the SVM accuracy as a fitness function to find the optimal subset feature.