ORCID

Jeremy Bertomeu, https://orcid.org/0000-0001-6746-5767

Language

English (en)

Publication Date

2018

Abstract

Machine learning offers empirical methods to sift through accounting data sets with a large number of variables and limited a-priori knowledge about functional forms. In this study, we show that these methods can help detect and interpret pat- terns present in ongoing accounting misstatements. We use a wide set of variables from accounting, capital markets, governance, and auditing datasets to detect mate- rial misstatements. Relative to traditional methods, we show that machine learning algorithms substantially improve out-of-sample detection accuracy. We also show that accounting variables become the most important factors only if considered in conjunction with other variables, especially audit variables. We also analyze differ- ences between misstatements and irregularities, compare algorithms, examine one- year and two-year ahead predictions, and interpret the output of the model with cutoff rules on observable variables identifying groups at greater risk of misstatements. ∗J. Bertomeu, E. Cheynel, and E. Floyd are from the Rady School of Management, University of Califor- nia San Diego. J. Bertomeu is an Associate Professor of Accounting, E. Cheynel is an Assistant Professor of Accounting, E. Floyd is an Assistant Professor of Accounting and Finance. W. Pan is a Ph.D. student at Columbia Business School, Columbia University. We gratefully thank B. Cadman, P. Dechow, C. Lennox, S.X. Li, D. Macciocchi, M. Plumlee, X. Peng and seminar participants at LSE, University of Utah, the USC-UCLA-UCSD-UCI conference, MIT, and the CMU Accounting Mini Conference for valuable feed- back. We also thank J. Engelberg for many suggestions central in seeding the project. 1 Machine learning is a broad discipline that has designed learning algorithms which can drive cars, recognize spoken language, and discover hidden regularities in growing volumes of data. Archival financial research is no exception with data streams ranging from firm characteristics, governance attributes, audit reports, market data and accounting variables. Machine learning algorithms detect complex patterns in the data, select the best variables to explain an outcome variable, and uncover suitable combinations of variables to make accurate out-of-sample predictions. They are the keys to unlocking the large - and growing - financial data sources to make better predictions and smarter decisions. This paper offers preliminary steps to applying this technology in accounting. We motivate the method by answering a practical question: How do we detect ongoing accounting misstatements? We focus on restatement items 4.02(a) “Non-Reliance on Previously Issued Finan- cial Statements or a Related Audit Report or Completed Interim Review,” which are, in principle, restatements that materially affect the interpretation of accounting numbers by investors. These misstatements differ from irregularities because they need not be frauds or carry evidence of managerial intent. Still, the audit procedures in place should have prevented these misstatements from occurring. The reasons for such misstatements can be complex and related to a large set of factors. Relative to a linear statistical model, machine learning algorithms are best suited for such problems in which the set of vari- ables, their interactions and the mapping into outcomes is not theoretically obvious. For our research question, many accounting numbers could serve as input variables to detect misstatements. An increase in accruals, for example, may either reflect growth or a firm overstating its earnings. Many other characteristics, such as industry, changes in balance sheet accounts including cash or inventories, or special transactions such as leases, do not have a clear theoretical relationship with misstatements. Intuitively, a machine learning algorithm can be thought to replicate the process of a researcher seeking a good empirical model for the relation between outcomes and input variables. This search involves many moving parts, such as finding the right specification and interactions between variables, which variables to select, and when to stop searching. The relation between the dependent and independent variables needs not be monotonic and the interactions between variables are a-priori unknown. Machine learning algorithms capture the relationship between misstatements and financial data, enabling the account- ing profession, auditors, managers, regulators and investors to evaluate more carefully the risks of misstatements, and offering a starting point in taking preventive actions. For our baseline analysis, we use the gradient boosted regression tree (GBRT) devel- oped in Friedman (2001) and Hastie, Tibshirani and Friedman (2001). GBRT is a regres- 2

Document Type

Working Paper

Author's Department

Accounting

Author's School

Olin Business School

Included in

Business Commons

Share

COinS