TY - GEN
T1 - Impact of membership and non-membership features on classification decision
T2 - 24th IEEE International Conference on Automation and Computing, ICAC 2018
AU - Abbasi, Bushra Zaheer
AU - Hussain, Shahid
AU - Bibi, Shaista
AU - Shah, Munam Ali
N1 - Publisher Copyright:
© 2018 Chinese Automation and Computing Society in the UK - CACSUK.
PY - 2018/9
Y1 - 2018/9
N2 - In text categorization, the discriminative power of classifiers, dataset characteristics, and construction of the more representative feature set play an important role in classification decisions. Subsequently, in text categorization, filter based feature selection methods are used rather than wrapper and embedded methods. In terms of construction of an illustrative feature set, a number of global and local filter based feature selection methods are used with their respective pros and cons. The inclusion and exclusion of membership and non-membership features in a constructed feature set depends on the discriminative power of the feature selection method. Though, there are few studies which have reported the impact of non-membership features on the classification decision. However, to best of our knowledge, there is no detail study, which calibrates the effectiveness of the feature selection method in terms of inclusion of non-membership features to improve the classification decisions. Consequently, in this paper, we conduct an empirical study to investigate the effectiveness of four well-known filter based feature selection methods, namely IG, \chi 2, RF, and DF. Subsequently, we perform a case study in the context of classification of the Gang-of-Four software design patterns. The results show that the balance consideration of membership and non-membership features has a positive impact on the performance of the classifier and classification decision can be improved. It has also been concluded that random forest is best among existing methods in considering an equal number of membership and non-membership features and the classifiers show better performance with this method as compare to others.
AB - In text categorization, the discriminative power of classifiers, dataset characteristics, and construction of the more representative feature set play an important role in classification decisions. Subsequently, in text categorization, filter based feature selection methods are used rather than wrapper and embedded methods. In terms of construction of an illustrative feature set, a number of global and local filter based feature selection methods are used with their respective pros and cons. The inclusion and exclusion of membership and non-membership features in a constructed feature set depends on the discriminative power of the feature selection method. Though, there are few studies which have reported the impact of non-membership features on the classification decision. However, to best of our knowledge, there is no detail study, which calibrates the effectiveness of the feature selection method in terms of inclusion of non-membership features to improve the classification decisions. Consequently, in this paper, we conduct an empirical study to investigate the effectiveness of four well-known filter based feature selection methods, namely IG, \chi 2, RF, and DF. Subsequently, we perform a case study in the context of classification of the Gang-of-Four software design patterns. The results show that the balance consideration of membership and non-membership features has a positive impact on the performance of the classifier and classification decision can be improved. It has also been concluded that random forest is best among existing methods in considering an equal number of membership and non-membership features and the classifiers show better performance with this method as compare to others.
UR - http://www.scopus.com/inward/record.url?scp=85069206804&partnerID=8YFLogxK
UR - http://www.scopus.com/inward/citedby.url?scp=85069206804&partnerID=8YFLogxK
U2 - 10.23919/IConAC.2018.8749009
DO - 10.23919/IConAC.2018.8749009
M3 - Conference contribution
AN - SCOPUS:85069206804
T3 - ICAC 2018 - 2018 24th IEEE International Conference on Automation and Computing: Improving Productivity through Automation and Computing
BT - ICAC 2018 - 2018 24th IEEE International Conference on Automation and Computing
A2 - Ma, Xiandong
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 6 September 2018 through 7 September 2018
ER -