Abstract
The shift from paper-based health records to Electronic Health Records (EHR) resulted in a vast volume of digital patient information. The knowledge derived from this data can be used for better decision-making and im-proving Personalized Care. It is very challenging to analyze EHR due to its high dimensional and heterogeneous nature. Deep Learning techniques can be used in this scenario. This work proposes a personalized care framework using Patient Similarity (DeepPCPS). DeepPCPS incorporates autoencoders for dimensionality reduction and uses various techniques for deducing patient similarity to find clusters o f similar patients. Experiments were conducted with the Jaccard index and the Sorensen Dice similarity index technique to calculate simi-larity scores. Based on similarity score, K-means, DBSCAN and agglomerative clustering algorithms were applied to form patient cohorts. The result suggests that the Sorensen Dice similarity with K-means has better results than the Jaccard index similarity. A quantitative analysis is conducted to support the claim. Qualitative analysis was conducted on the clusters formed from Sorensen similarity to show how the cohorts were used for customized treatment. The results obtained were promising, and further investigation is needed by incorporating additional treatment and timeline features for personalized care.