Actionable Insights Backed by Research
Over the last decade, manufacturers like Volkswagen and Nissan have been at the center of massive scandals due to their falsification of emissions data. But while these egregious violations made headlines, they were likely only the tip of the iceberg. After all, sustainability laws are only as useful as regulators’ ability to actually identify the companies breaking them — which is a major challenge when most of the data regulators use to determine compliance is self-reported. As a result, it’s probable that only a small proportion of violators are caught through random inspections or whistleblowers.
To bridge this enforcement gap, Price’s Mei Li and coauthors collaborated with China’s Ministry of Ecology and Environment (MEEC) to create a tool that uses advanced machine learning techniques to analyze reported emissions data and identify likely cases of fraud and falsification. In a real-world test, this tool helped MEEC identify a factory that was clearly falsifying its data as well as several other suspicious cases, validating that it can effectively help regulators move beyond random inspections toward a much more efficient, scalable approach to identifying falsified data and actually enforcing essential environmental regulations.
Author: Mei Li, University of Oklahoma, Associate Professor of Supply Chain Management
In 2015, one of the biggest emissions scandals in history broke: Volkswagen had intentionally cheated on government-mandated emissions tests, reporting emissions rates up to 40 times lower than its actual rates. Over the next several years, global regulators levied billions of dollars in fines, and Volkswagen’s stock plummeted. The scandal highlighted just how dramatically data falsification can hinder the enforcement of environmental regulations — but of course, in all likelihood, this was only the tip of the iceberg.
While a handful of companies have been found out for violating environmental laws, these regulatory efforts are often stymied by a major enforcement problem. After all, sustainability regulations typically rely on self-reported data, which can easily be falsified, to determine compliance. This can seriously hinder regulators’ ability to actually identify companies that knowingly break the rules, making it likely that only a small proportion of violators are caught through random inspections or whistleblowers.
Specifically, many countries rely on Continuous Emission Monitoring Systems, or CEMS, to gather data from manufacturers that emit harmful chemicals. These systems typical entail a physical piece of hardware that is installed on-site at a factory or manufacturing facility. The device tracks emissions of waste gases and other pollutants and automatically sends that data to a public database. The resulting dataset is extensive, but since the collection and reporting process is entirely controlled by the firm, it is relatively easy for companies violating legal emissions limits to falsify their self-reported data.
In contrast, other publicly available data is more reliable, but a lot less extensive. In China, for instance, a Violation and Punishment Dataset (VPD) captures data on known environmental regulation violators, but it is far from complete, since most violators are never caught. This poses a major challenge when it comes to actually enforcing vital sustainability laws: How can regulators effectively enforce regulations that rely on self-reported data?
To help bridge this gap, I worked with coauthors at several leading academic institutions as well as China’s Ministry of Ecology and Environment (MEEC) to create a tool that uses advanced machine learning techniques to analyze reported emissions data and identify likely cases of fraud and falsification.
We began by analyzing 18 months of publicly available CEMS data for more than 7,000 factories across China. Using a technique known as Feature Engineering, we identified six key features of the self-reported data that tend to correlate with falsification. Through interviews with senior MEEC regulators and specialists from CEMS device manufacturers as well as a comprehensive literature review, we uncovered the following features that may be indicative of red flags:
For each of these metrics, we calculated the ratio between the period of suspicious data and the total data collection period for a given firm. So, if a firm had a lot of normal data and just a small amount of suspicious data, they wouldn’t be flagged — but if a manufacturer had a relatively large amount of suspicious data, that suggested they may be falsifying data.
Then, once we defined these features, we used a machine learning algorithm trained on both the CEMS and VPD data to create a model that uses those six features to identify cases of likely falsification. This model essentially connects the dots between the large dataset of self-reported emissions data and the small dataset of known violators to automatically identify facilities that may be falsifying their data or engaging in other suspicious practices.
Importantly, this tool cannot offer a 100% guarantee that flagged firms are actually violating regulations or falsifying data. But what it can do is produce a list of the most suspicious firms recommended for human inspection, enabling regulators to prioritize their limited inspection resources much more efficiently.
To validate the tool, we conducted a real-world test: We used it to identify the top five most suspicious firms in China, and then we had MEEC inspect those five flagged firms. And what did we find? Of those five firms, one was engaged in blatant data falsification, one could not explain its abnormal readings, and two others were engaged in highly suspicious activity — an impressive success rate for an regulatory organization that was used to relying on random inspections alone.
The most egregious violator was a large steel manufacturer that had obviously tampered with its emissions monitoring device. As shown in the image below, inspectors found a hidden break in the intake area of the CEMS device. The break allowed extra oxygen to enter the area, diluting the pollutants and thus distorting the readings on their monitoring devices.
Two of the other companies flagged by our tool had replaced their monitoring devices immediately prior to the inspection, suggesting that they may have been tipped off and used the advance warning to cover up evidence of malfeasance. Indeed, our interviews with subject matter experts confirmed that both of these cases were highly suspicious and that both firms likely would have been caught circumventing the monitoring systems if the devices hadn’t been changed right before the inspections. The fourth company could not explain its unusual testing results, and the fifth company had just upgraded its equipment in a way that reduced its pollution rates.
In other words, the tool achieved a fairly high accuracy rate: Four of the five flagged companies were in fact engaged in suspicious or outright fraudulent activities. Taken together, this validation in the field suggest that our tool can help regulators identify many more rule-breakers than they could through random inspections. As such, our findings offer important practical takeaways for both regulators and firms.
On the one hand, our results demonstrate that regulators can immediately improve their effectiveness by using this tool or one like it to flag suspicious firms. While our study focused on China, the EPA in the U.S. uses the same CEMS devices and thus has access to the same kinds of data as that used in our study. Moreover, similar data sources are also available in Canada, Europe, and other global regions. As such, regulators around the world can use this tool to complement random inspections with more targeted inspections of flagged firms
On the other hand, as firms become aware of this tool and its effectiveness in identifying falsified data, they will be increasingly incentivized to avoid falsifying data and instead actually adhere to environmental regulations. So, to avoid becoming the next Volkswagen, firms should get ahead of these improved enforcement capabilities by proactively ensuring compliance and accurate self-reporting.
To be sure, as with any tool, this new approach to identifying suspicious firms is only as effective as the person or organization that’s wielding it. In particular, our real-world test highlighted that information leaks pose a real hurdle to effective enforcement. If violators are tipped off that inspectors are coming, they can avoid getting caught, as likely happened with two of the five companies in our study. As such, implementation of this tool will be more effective if accompanied by stringent confidentiality and security practices.
Also, as noted above, this tool does not provide a 100% guarantee that a firm is or isn’t breaking the rules. It should be used to guide regulatory activities, but should never be used to replace human inspection and verification, as it is still far from a perfect predictor of violations.
Nevertheless, despite these limitations, our research demonstrates that new, AI-powered tools can serve as important additions to the toolbox of global sustainability regulators. As environmental issues grow increasingly urgent, it’s more critical than ever that regulators are empowered to actually enforce sustainability laws. By leveraging machine learning models trained on publicly available data, regulators can make substantial strides in identifying firms trying to break the rules — ultimately helping them do their jobs more effectively, incentivizing companies to avoid breaking the rules in the first place, and keeping our planet safer and less polluted for generations to come.