Emissions Data Is a Data Engineering Problem Nobody Is Taking Seriously Yet

Ask an operator who owns their emissions numbers and you’ll usually get pointed at HSE, or a sustainability consultant, or whoever drew the short straw on the last ESG report. Ask where the numbers come from and the room goes quiet.

That’s the tell. Emissions reporting gets treated as a reporting problem, a thing a person assembles in a spreadsheet once a year and hands to someone who formats it into a PDF. But the number in that PDF is the output of a data pipeline, whether anyone built it deliberately or not. Flare volumes come from SCADA. Methane attribution needs well-level production. Scope 1 depends on an equipment inventory that lives in a maintenance system nobody has cleaned in years. The report is the easy part. Producing a defensible number is a data engineering problem, and almost nobody is treating it like one.

The regulatory ground moved, and it doesn’t matter as much as you think

Here’s the part that surprises people. The federal pressure that was supposed to force this got weaker, not stronger.

The SEC’s climate disclosure rules, adopted in March 2024, were stayed that April and never took effect.[1] In March 2025 the Commission voted to stop defending them in court, and in May 2026 it formally proposed to rescind them in their entirety.[2] On the methane side, EPA’s Waste Emissions Charge, the per-ton fee on methane from petroleum and natural gas systems, had its compliance rule overturned by Congress in 2025, and the charge itself is now deferred to emissions reported for 2034 and later.[3] EPA has even proposed scaling back most of the Greenhouse Gas Reporting Program.

So the mandate that was going to make everyone panic has, for now, softened. If your emissions program was built entirely on fear of the SEC, you got a reprieve.

We’d argue that changes almost nothing about the work. Federal enforcement calendars swing every four years. The rules that got stayed can get un-stayed, and a rule proposed for rescission in 2026 is one election away from being reproposed. Meanwhile the parties that actually move on this timescale, lenders, equity partners, and the buyer’s diligence team, never stopped asking. A private-equity sponsor wants a defensible carbon intensity number before they underwrite. A reserve-based lender’s borrowing base review asks for it. And when you put an asset on the market, the buyer’s technical team treats your emissions data exactly like they treat your production history: something to audit, not something to take on faith.

The regulatory teeth come and go. The commercial teeth don’t. Either way, you need a number you can stand behind, and the work to produce one is identical regardless of which pressure is currently pointed at you.

What actually feeds an emissions calculation

Walk backward from a Scope 1 number for an upstream operator and you land in the same systems you already fight with every month.

Flaring and venting. Flare volumes come from SCADA, sometimes from a dedicated flare meter, more often estimated from a control setpoint or a difference between produced and sold gas. Those are three different numbers with three different error bars, and which one you use changes the answer materially.

Combustion. Compressor engines, heaters, generators. Fuel gas burned is fuel gas not sold, so it shows up as a shrink in your gas balance, if your gas balance is any good. Compressor run hours drive a big share of the estimate, and run hours live in SCADA or in a maintenance log or in a driver’s memory, depending on the vintage of the site.

Pneumatic devices. This is the one that quietly dominates the methane estimate for a lot of gas-weighted operators, and it’s the worst-maintained data of the bunch. The calculation is basically device count times device type times an emission factor. That device count comes from an equipment inventory, and the equipment inventory is a spreadsheet somebody built during an acquisition and never updated after the field crews swapped a dozen high-bleed controllers for low-bleed ones.

Fugitives and everything else. Tank flashing, dehydrator vents, liquids unloading, equipment leaks found on a survey. Each has its own source, its own factor, its own way of being wrong.

The methane frameworks that matter here, EPA’s Subpart W and the UN-backed OGMP 2.0, both push you toward source-level accounting rather than a single top-line estimate.[4][5] OGMP 2.0 explicitly grades reporting by how granular and measurement-based it is, from an asset-level aggregate at Level 1 up to a source-level inventory reconciled against site measurements at Level 5.[5:1] You cannot report at the higher levels out of a spreadsheet. You need the source data joined, cleaned, and traceable, which is a pipeline.

Why these sources don’t reconcile

If this list of systems sounds familiar, it should. It’s the same reconciliation problem we’ve written about with land and production data, wearing a different hat.

The well identity problem is right back. SCADA knows the site by a device tag or a controller address. Production knows the well by API number. The equipment inventory knows it by whatever the field called it during the last acquisition. Before you can attribute a flare volume or a compressor’s fuel burn to a well or a facility, you have to resolve those identities to one master, and that is exactly the entity-resolution work most operators have been deferring for a decade.

The allocation problem is back too. A tank battery serves six wells. The flash emissions off that battery have to be split among them to get a per-well intensity, and the split follows an allocation methodology that may or may not match the one accounting uses for revenue. Now you have two allocation methods in the same company producing two versions of “what came off well 4,” and an auditor is going to ask why.

The equipment inventory is the genuinely new problem, and it’s the ugliest. Production data at least gets looked at every month during close, so its errors get some scrutiny. Nobody closes the books on a pneumatic controller count. That data goes stale silently, and a stale device inventory produces a wrong methane number that looks exactly as confident as a right one. We made this point about production data generally in what goes wrong with upstream data quality; it’s sharper for emissions, because there’s no monthly reconciliation catching the drift.

A number you can defend versus one you can’t

This is the distinction that matters, and it has nothing to do with whether the number is “accurate” in some abstract sense. Every emissions number is an estimate. The question an auditor or a buyer’s team asks is not “is this exactly right.” It’s “can you show me how you got here, and would you get the same answer if you ran it again.”

A number you can defend has a documented path from source to total. You can point at the flare volume, say which meter or estimate it came from, name the emission factor and cite its version, show the device count and the date the inventory was last verified, and reproduce the whole calculation from raw inputs. If a factor changed mid-year, the number knows which periods used which factor.

A number you can’t defend is a cell in a workbook that references twelve other workbooks, three of which the analyst who built them has since left. It gives an answer. It cannot survive the question “where did the 4,200 come from,” and in a diligence room that question gets asked about every line that moves the total.

The difference is entirely data engineering. Lineage, versioned rules, reproducibility, an audit trail. This is the same argument we made in the 42 Gallons series about production data: the industry measures every physical barrel with custody-grade rigor and then runs the data on the honor system. Emissions is the same failure, with a regulator or a buyer eventually holding the meter ticket.

What a minimum viable emissions pipeline looks like

You don’t need a platform. You need the same medallion-style discipline you’d apply to any operational data, pointed at these sources.

Land the raw inputs, unchanged. Flare and fuel volumes from SCADA, production allocations, compressor run hours, the equipment inventory. Bronze layer, one place per source, decorated with where it came from and when. Don’t clean it on the way in; you want the raw values preserved so the calculation is reproducible later.

Resolve identity and allocate. Join every source to a single well and facility master so a flare volume can actually be attributed to the right entity. This is where PPDM earns its place. The well master you built for production and land reconciliation is the same foundation emissions attribution needs, which means the operators who already did the master-data work get emissions attribution nearly for free, and the ones who didn’t hit the same wall twice.

Apply factors as versioned, owned rules. The emission factors and the allocation methods are business logic. They belong in code or in dbt models with tests, not in a formula bar, and every factor should carry the source and the effective date. When EPA updates a factor or you re-survey a field and the device counts change, that’s a versioned change with a date, not a silent edit to last year’s spreadsheet.

Test the inputs before they reach the total. The checks worth writing first are the boring ones: device counts that dropped to zero, run hours exceeding hours in the month, a flare volume that’s negative, an allocation that doesn’t sum to the metered total. These are the same accepted-range and sum-check tests that catch bad production data, and they catch bad emissions data for the same reasons.

Keep the audit trail. Every override, every manual adjustment, every factor change, captured with a person, a timestamp, and a reason. When someone asks in eighteen months why the 2026 number looks the way it does, the answer is a query.

None of this is exotic. It’s the SCADA-to-warehouse work and the ingestion discipline we already do for operators, aimed at a total that happens to end up in a sustainability report instead of a production dashboard.

Where to start

Pick your largest emission source and trace it end to end. For most gas-weighted operators that’s flaring, venting, or pneumatics. Find out which system the input actually comes from, how current it is, and whether you could reproduce last year’s number from raw data today. You almost certainly can’t, and that gap is the whole project in miniature.

The federal deadline pressure eased for the moment. The buyer’s diligence team, the lender’s borrowing base, and the equity partner’s underwriting did not, and the regulatory pressure is one policy cycle from returning. An emissions number built on a clean, attributed, versioned pipeline is defensible to all of them and cheap to reproduce every year after the first. One rebuilt by hand each spring is a number you’re hoping nobody looks at too closely. In this business, somebody always does.


Get in touch


  1. U.S. Securities and Exchange Commission, “The Enhancement and Standardization of Climate-Related Disclosures for Investors” (adopted March 6, 2024; effectiveness stayed April 4, 2024). https://www.sec.gov/rules-regulations/2026/05/s7-2026-19 ↩︎

  2. U.S. Securities and Exchange Commission, “SEC Proposes Rescission of Climate-Related Disclosure Rules,” Press Release 2026-49 (May 29, 2026); comment period open through August 3, 2026. The Commission voted to end its defense of the rules on March 27, 2025. https://www.sec.gov/newsroom/press-releases/2026-49-sec-proposes-rescission-climate-related-disclosure-rules ↩︎

  3. Congressional Research Service, “Inflation Reduction Act Methane Emissions Charge: Overview and Developments,” Report R48475 (2025). The Waste Emissions Charge compliance rule was disapproved under the Congressional Review Act (H.J.Res. 35, Pub. L. 119-2, March 2025), and the charge was subsequently deferred to emissions reported for calendar year 2034 and later. https://www.congress.gov/crs-product/R48475 ↩︎

  4. U.S. Environmental Protection Agency, “Subpart W: Petroleum and Natural Gas Systems,” Greenhouse Gas Reporting Program (40 CFR Part 98, Subpart W). Facilities emitting 25,000 metric tons or more of GHGs per year, expressed as carbon dioxide equivalents, are required to report. https://www.epa.gov/ghgreporting/subpart-w-petroleum-and-natural-gas-systems ↩︎

  5. United Nations Environment Programme, “OGMP 2.0 Reporting Framework” (2020). The framework grades methane reporting across five levels, from asset-level aggregates (Level 1) to a source-level inventory reconciled with site-level measurements (Level 5). https://www.ogmpartnership.org/sites/default/files/resources/2025-04/OGMP_20_Reporting_Framework.pdf ↩︎ ↩︎