Faculties » Faculty of Management and Economics » Institute of Management

Machine Learning and Accounting Fraud Prediction - Adding Context to Improve Prediction Accuracy

 

Abstract

This study sheds light on the question why machine learning performs poorly in account- ing fraud prediction and suggests ways to improve prediction accuracy. In theory, machine learning enables the identification of complex patterns in the data, such as unusual combina- tions of financial variables that are characteristic of specific fraud schemes, and should thereby be more accurate than simpler techniques. However, based on SEC enforcement action data and using prior research’s latest machine learning algorithms, we find that algorithms do not exploit such combinations. Instead, they mainly rely on one specific pattern when classifying firms’ compliance: similarity in the absolute levels of input factors. Because absolute levels are largely driven by firm size and industry structure rather than by fraud-related characteristics, this renders the algorithm’s positive classifications a reflection of financial proximity to known fraud firms in the training data rather than a firm’s actual fraud risk profile. Addressing this limitation, we introduce a novel approach that optimizes the algorithm in the validation period by conditioning positive classifications on contextual information about industry characteristics, governance structures, and signs of financial pressure. Applying this approach, we are able to improve prediction performance by two to three times. Collectively, our results suggest that providing machine learning algorithms with theoretically grounded contextual information that allows them to distinguish meaningful from spurious financial similarity can substantially im- prove prediction accuracy.