EDMIX: an entropy-based dissimilarity measure to cluster mixed data comprising of numerical–nominal–ordinal attributes

Publications

EDMIX: an entropy-based dissimilarity measure to cluster mixed data comprising of numerical–nominal–ordinal attributes

Year : 2025

Publisher : Springer Science and Business Media Deutschland GmbH

Source Title : Knowledge and Information Systems

Document Type :

Abstract

Existing (dis)similarity measures for mixed data often ignore data distribution and ordering information along numerical and categorical features, respectively, during distance computation. Additionally, combining numerical and categorical (dis)similarities is often done naively, failing to judiciously assign relative weights. This work introduces a unified dissimilarity measure for mixed data (EDMIX) that addresses these challenges. It employs an entropy-based measure (SEND) for numerical attributes which utilizes intra-attribute inhomogeneity efficiently. A Boltzmann’s entropy-based measure (EDMD) is adopted for categorical features, which effectively handles both nominal and ordinal attributes. Decay of attribute weight is analyzed to determine the threshold weights for numerical and categorical parts. Number of attributes whose weights are above the threshold value in each part decides the mixing proportion of the respective parts. Notably, EDMIX requires no input parameters and is adaptable to diversified mixed data. Experimental results highlight its superiority in terms of cluster quality, accuracy, discrimination ability, and execution time across diverse mixed datasets for clustering applications.