Learning latent representations of bank customers with the Variational Autoencoder

Andrade Mancisidor, Rogelio; Kampffmeyer, Michael; Aas, Kjersti; Jenssen, Robert

Andrade Mancisidor, Rogelio; Kampffmeyer, Michael; Aas, Kjersti; Jenssen, Robert

Journal article, Peer reviewed

Submitted version

Åpne

RevisedManuscript_clean.pdf (2.554Mb)

Permanent lenke

https://hdl.handle.net/11250/2732556

Utgivelsesdato

2020

Metadata

Vis full innførsel

Samlinger

Originalversjon

Expert systems with applications. 2020, 164 . 10.1016/j.eswa.2020.114020

Sammendrag

Learning data representations that reflect the customers’ creditworthiness can improve marketing campaigns, customer relationship management, data and process management or the credit risk assessment in retail banks. In this research, we show that it is possible to steer data representations in the latent space of the Variational Autoencoder (VAE) using a semi-supervised learning framework and a specific grouping of the input data called Weight of Evidence (WoE). Our proposed method learns a latent representation of the data showing a well-defied clustering structure. The clustering structure captures the customers’ creditworthiness, which is unknown a priori and cannot be identified in the input space. The main advantages of our proposed method are that it captures the natural clustering of the data, suggests the number of clusters, captures the spatial coherence of customers’ creditworthiness, generates data representations of unseen customers and assign them to one of the existing clusters. Our empirical results, based on real data sets reflecting different market and economic conditions, show that none of the well-known data representation models in the benchmark analysis are able to obtain well-defined clustering structures like our proposed method. Further, we show how banks can use our proposed methodology to improve marketing campaigns and credit risk assessment.

Tidsskrift

Expert systems with applications

Med mindre annet er angitt, så er denne innførselen lisensiert som Navngivelse-Ikkekommersiell-DelPåSammeVilkår 4.0 Internasjonal