Clustering with density based initialization and Bhattacharyya based merging

Köse, Erdem; Hocaoğlu, Ali

doi:10.3906/elk-2105-44

Clustering with density based initialization and Bhattacharyya based merging

Köse E., Hocaoğlu A. K.

Turkish Journal of Electrical Engineering and Computer Sciences, cilt.30, sa.3, ss.502-517, 2022 (SCI-Expanded)

Yayın Türü: Makale / Tam Makale
Cilt numarası: 30 Sayı: 3
Basım Tarihi: 2022
Doi Numarası: 10.3906/elk-2105-44
Dergi Adı: Turkish Journal of Electrical Engineering and Computer Sciences
Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Academic Search Premier, Applied Science & Technology Source, Compendex, Computer & Applied Sciences, INSPEC, TR DİZİN (ULAKBİM)
Sayfa Sayıları: ss.502-517
Anahtar Kelimeler: Infinite mixture models, density estimation, Jensen inequality, bandwidth selection, optimal number of, DISTANCE, NUMBER
Kocaeli Üniversitesi Adresli: Hayır

Özet

© TÜBİTAKCentroid based clustering approaches, such as k-means, are relatively fast but inaccurate for arbitrary shape clusters. Fuzzy c-means with Mahalanobis distance can accurately identify clusters if data set can be modelled by a mixture of Gaussian distributions. However, they require number of clusters apriori and a bad initialization can cause poor results. Density based clustering methods, such as DBSCAN, overcome these disadvantages. However, they may perform poorly when the dataset is imbalanced. This paper proposes a clustering method, named clustering with density initialization and Bhattacharyya based merging based on the fuzzy clustering. The initialization is carried out by density estimation with adaptive bandwidth using k-Nearest Orthant-Neighbor algorithm to avoid the effects of imbalanced clusters. The local peaks of the point clouds constructed by the k-Nearest Orthant-Neighbor algorithm are used as initial cluster centers for the fuzzy clustering. We use Bhattacharyya measure and Jensen inequality to find overlapped Gaussians and merge them to form a single cluster. We carried out experiments on a variety of datasets and show that the proposed algorithm has remarkable advantages especially for imbalanced and arbitrarily shaped data sets.