Handling Missing Values in Information Systems Research: A Review of Methods and Assumptions
Jiaxu Peng, Jungpil Hahn, Ke‐Wei Huang
Information Systems Research
- 주제정보시스템 연구방법론 · 경영정보·의사결정
- 방법
- 현상
Data have never been more essential to the success of decision making. However, data are often messy. A perennial data challenge is missing values, which frequently occur in real-world data, such as unreported data items in public firms’ financial statements and skipped product ratings from consumers. What is the influence of missing values and how should they be handled? Although we are in a big data era, missing values are not ignorable if data are missing for nonrandom reasons. In the case of product ratings, if only people who favor the product provide ratings while others put aside the product and do not respond, then even a simple mean estimation of the product rating would be significantly biased. Such bias challenges the validity of data analysis, and it cannot be eliminated simply by increasing the sample size of the data. To correct the bias arising from nonrandom missing values, it is necessary to examine and model what causes the missing values. We propose and demonstrate the superior performance of a Monte Carlo likelihood approach to correct the bias. Overall, we recommend well-designed data collection processes with documentation of the possible reasons for missing values, cautious adoption of missing value handling methods, and structured missing value reporting practices.
불러오는 중…
불러오는 중…
불러오는 중…
불러오는 중…
- 저널Information Systems Research · 34(1) · 5–26
- 토픽Big Data and Business Intelligence · Management Information Systems
- DOI10.1287/isre.2022.1104
- 저자Jiaxu Peng, Jungpil Hahn, Ke‐Wei Huang