Addressing imbalance in multilabel classification: Measures and random resampling algorithms

Francisco Charte; A.J. Rivera-Rivas; M. J. del Jesus; F. Herrera

Submitted by fcharte on Thu, 28/02/2019 - 13:08

Title	Addressing imbalance in multilabel classification: Measures and random resampling algorithms
Publication Type	Journal Article
Year of Publication	2015
Authors	Charte, Francisco, Rivera-Rivas A.J., del Jesus M. J., and Herrera F.
Journal	Neurocomputing
Volume	163
Start Page	3
Pagination	3–16
Abstract	The purpose of this paper is to analyze the imbalanced learning task in the multilabel scenario, aiming to accomplish two different goals. The first one is to present specialized measures directed to assess the imbalance level in multilabel datasets (MLDs). Using these measures we will be able to conclude which MLDs are imbalanced, and therefore would need an appropriate treatment. The second objective is to propose several algorithms designed to reduce the imbalance in MLDs in a classifier-independent way, by means of resampling techniques. Two different approaches to divide the instances in minority and majority groups are studied. One of them considers each label combination as class identifier, whereas the other one performs an individual evaluation of each label imbalance level. A random undersampling and a random oversampling algorithm are proposed for each approach, giving as result four different algorithms. All of them are experimentally tested and their effectiveness is statistically evaluated. From the results obtained, a set of guidelines directed to show when these methods should be applied is also provided.
Notes	TIN2012-33856,TIN2011-28488,P10-TIC-6858,P11-TIC-7765
DOI	10.1016/j.neucom.2014.08.091

Fichero:

2015-Neucom-AddressingImbalance.pdf