The output of a monitoring program is evidence
A state air bureau, a municipal wastewater utility, a refinery environmental group, and a Department of Energy remediation site all run the same shaped problem. Instruments produce numbers continuously. A small fraction of those numbers become the basis of a permit determination, a public health advisory, a penalty, or a lawsuit. The distance between a number and a defensible result is the record behind it: where the probe sat, when it was last calibrated against a traceable standard, who had custody of the sample, which qualifier flag the value carries, and what decision rule was written down before anyone saw the data. Machine learning is genuinely useful inside that world. It is only useful inside that record.
Most environmental AI projects fail at the same place. A team builds a model against a clean historical extract, shows a promising validation curve, then meets production data carrying instrument fault codes, non-detects reported as blanks, three different units in one column, and a QA flag vocabulary that changed when the state migrated laboratory systems. Modeling was never the hard part. Regulatory data have provenance obligations that ordinary analytics data do not.

Where modeling pays off in a monitoring program
Editorial weighting from public regulatory guidance and practitioner reading. Illustrative, not a measured statistic.
The air side: a reference network with a sensor tier bolted on
Regulatory ambient air monitoring in the United States runs on 40 CFR Part 58. Appendix D sets network design criteria, Appendix E sets probe and inlet siting, and Appendix A sets the quality assurance requirements that make the data usable for comparison against the National Ambient Air Quality Standards. Instruments have to be designated Federal Reference Methods or Federal Equivalent Methods under 40 CFR Part 53. Data land in EPA's Air Quality System. Design values get computed by the arithmetic in 40 CFR Part 50, Appendix N, and those design values are what determine whether an area is in attainment.
The thresholds that matter right now: the primary annual PM2.5 standard was tightened to 9.0 micrograms per cubic meter in EPA's February 2024 final rule, with the 24-hour standard staying at 35. The 8-hour ozone standard remains 0.070 parts per million, evaluated on the fourth-highest daily maximum averaged over three years. Areas that sat comfortably inside the old annual PM2.5 value now sit near the line, and near the line is where data completeness and QA flags start deciding outcomes.
Appendix A prescribes the QA burden directly. Gaseous analyzers get a one-point quality control check at least once every two weeks. Continuous PM monitors get flow rate verifications and semiannual audits. Roughly 15 percent of PM2.5 monitors carry collocated quality control samplers so precision can be estimated. Independent audits arrive through the National Performance Audit Program and, for particulate, the Performance Evaluation Program. Each of those checks leaves a record, and each record is a feature a drift detection model can use. QA data are often better training signal than the concentration data.
The sensor tier is where agencies are moving fastest and where the defensibility question is sharpest. Low-cost optical PM and electrochemical gas sensors cost two to three orders of magnitude less than an FEM analyzer and deploy by the hundreds. EPA's 2021 performance testing protocols for fine particulate sensors set the targets most teams now quote: a coefficient of determination of at least 0.70 against a collocated FRM or FEM, and normalized root mean square error at or below 30 percent. Optical PM sensors also overestimate badly at high relative humidity, a well-characterized and correctable problem. Correction models fitted on collocation data plus humidity are standard practice, and the corrections belong in the published record, not hidden inside a dashboard.
Satellite air data have improved sharply. NASA's TEMPO instrument, in geostationary orbit since April 2023, delivers hourly daytime nitrogen dioxide, formaldehyde, and ozone over North America at roughly two by four and a half kilometers. Sentinel-5P TROPOMI has provided daily global nitrogen dioxide, sulfur dioxide, and methane columns since 2017 at coarser resolution. Both are strong for regional pattern and trend work. Neither substitutes for a monitor when the question is whether one facility exceeded one limit on one day.
The water side: permits, DMRs, and continuous records
On the water side the compliance object is the permit. A National Pollutant Discharge Elimination System permit issued under 40 CFR Part 122 names outfalls, parameters, limits, averaging periods, monitoring frequencies, and analytical methods. The permittee reports through Discharge Monitoring Reports under 40 CFR 122.41(l)(4), filed electronically through NetDMR under the Part 127 requirements, into ICIS-NPDES. The result is public within weeks through ECHO, so anyone, including a citizen suit plaintiff, can pull the same record the regulator sees.
Two mechanics drive most of the analytic work. First, methods are prescribed: 40 CFR Part 136 lists the approved methods for Clean Water Act compliance, and Appendix B to Part 136 defines the method detection limit procedure, revised in 2017 to incorporate method blank data. A result below the detection limit is a censored value, and how censored values enter an average is a decision rule that belongs in the permit or the plan, never in a modeler's preprocessing step. Second, exceedances are graded. The Significant Noncompliance criteria in Appendix A to 40 CFR Part 123.45 apply technical review criteria of 1.4 for Group I pollutants and 1.2 for Group II on monthly average limits, and separately flag chronic patterns across a six-month window. Those numbers define what an alert actually means.
Continuous water quality monitoring brings the richest data and the strictest correction discipline. A deployed sonde measuring turbidity, specific conductance, dissolved oxygen, pH, and temperature drifts and fouls. USGS Techniques and Methods 1-D3 defines the fouling and drift correction procedure for continuous records and assigns record grades from excellent through poor based on the size of the corrections applied. That grading is a gift to anyone building analytics on top, because every point arrives with a documented reliability tier. USGS Techniques and Methods 3-C4 extends the same discipline to surrogate regression, estimating suspended sediment from turbidity or chloride from specific conductance, with published prediction intervals attached.
Drinking water sits under a separate regime with sharper reporting limit implications. EPA's April 2024 national primary drinking water regulation for PFAS set a maximum contaminant level of 4.0 parts per trillion for PFOA and PFOS, close enough to routine laboratory quantitation limits that method choice, blank contamination control, and non-detect handling become the whole analysis. The Lead and Copper Rule Revisions pushed systems to complete service line material inventories, which for most utilities meant reconciling historical tap cards, permit records, and field verification into one dataset. That is a records reconciliation problem with a sampling component, and it is where document extraction and entity matching earn their cost.
Remote sensing: pick the sensor for the legal question
Satellite imagery gets oversold in environmental work because the demos are visually persuasive. The useful framing is narrow: match the source to the question, and know before you start whether the output is an enforcement input or a targeting input.
| Source | Resolution and revisit | What it supports |
|---|---|---|
| Landsat 8 / 9 (OLI, TIRS) | 30 m multispectral, 100 m thermal; 16-day each, roughly 8-day combined | Long baselines back to the 1980s, thermal signatures of discharge plumes, Collection 2 Level-2 surface reflectance and surface temperature products |
| Sentinel-2 (MSI) | 10 m and 20 m bands; about 5-day revisit | Land disturbance acreage, riparian buffer compliance, stockpile and impoundment extent, bloom screening on larger water bodies |
| Sentinel-1 (C-band SAR) | Roughly 5 by 20 m; 6 to 12 day repeat | Cloud-free flood and inundation mapping, surface change on tailings and ash impoundments, ground displacement through interferometry |
| Sentinel-5P (TROPOMI) | About 5.5 by 3.5 km at nadir, daily global | Regional nitrogen dioxide, sulfur dioxide, and methane anomalies; screening, not facility-level attribution |
| Sentinel-3 (OLCI) | 300 m; near-daily | Cyanobacteria and chlorophyll trends on lakes and reservoirs, the basis of the interagency CyAN products |
| Harmonized Landsat Sentinel-2 | 30 m, 2 to 3 day effective revisit | Change detection where cadence matters more than pixel size, on a single consistent surface reflectance grid |
The engineering under these products is unglamorous and decisive. Cloud and shadow masking through the Landsat Collection 2 quality bands or an equivalent classifier. Consistent surface reflectance, so a shift in a spectral index reflects the ground rather than the atmosphere. Co-registration tight enough that a boundary polygon means the same thing across dates. Terrain and illumination correction in relief. Get those wrong and the change detector will faithfully find changes in the weather.
Exceedance detection is a legal event with an alert attached
The single most common design error is treating exceedance as an anomaly detection problem. It is not. An exceedance is the output of a decision rule that was fixed in advance: a parameter, an averaging period, a completeness threshold, a rounding convention, a censoring rule, and a comparison to a numeric limit. The model does not get a vote on that rule. If the permit says monthly average and the permit's method is Part 136 Method 200.8, then the monthly average computed from Method 200.8 results is the compliance number, whatever a model would have predicted.
Completeness and substitution rules are where analytics quietly go wrong. Ambient air data carry quarterly completeness requirements, generally 75 percent, before a quarter counts toward a design value. The acid rain program under 40 CFR Part 75 requires 95 percent monitor data availability and prescribes missing data substitution procedures that are deliberately conservative, so a gap costs the source money. Those are legally defined imputation rules. A model that fills the same gap with a better estimate produces a number that is more accurate and less reportable. Both values can live in one system. Only one goes on the form.
Three alert classes belong on separate paths in any system that pages a human: instrument fault, process excursion, and reportable deviation. A fault routes to the field technician. An excursion routes to operations, ideally early enough to prevent the third class. A reportable deviation routes to the environmental manager with a clock attached, since permits carry notice deadlines, commonly 24 hours orally and five days in writing for NPDES noncompliance that may endanger health or the environment. Systems that collapse the three into one stream get muted within a month.
False positive control is a design parameter. EPA's Unified Guidance for statistical analysis of groundwater monitoring data at RCRA facilities builds its whole method around holding a site-wide false positive rate near 10 percent per year across all wells and constituents, using prediction limits combined with formal retesting. That is the right instinct for any network: set the tolerable annual false alarm budget first, then pick the statistic and the retest that meet it. On the other side, the exceptional events provisions at 40 CFR 50.14 let an agency exclude data influenced by events such as wildfire smoke, but only through a documented demonstration with its own notification and submission deadlines. Flagging candidate periods automatically is useful. Excluding them automatically is not defensible.
Data quality objectives come before the model
EPA's systematic planning process, documented in Guidance on Systematic Planning Using the Data Quality Objectives Process (EPA QA/G-4, EPA/240/B-06/001), is seven steps and takes a few weeks. Skipping it is the most expensive shortcut in this field, because the sampling design determines what questions the data can ever answer, and no amount of modeling recovers a design that was underpowered or spatially misaligned with the decision.
The seven-step DQO process, as we run it
Step 6 sets the measurement quality objectives, and the traditional data quality indicators still organize that conversation: precision, accuracy and bias, representativeness, comparability, completeness, and sensitivity. Each becomes a numeric acceptance criterion. Precision might be a relative percent difference limit on field duplicates. Bias might be a percent recovery window on matrix spikes. Completeness might be 90 percent of scheduled samples. Sensitivity is the quantitation limit relative to the standard being compared against, which is exactly what makes a 4.0 part per trillion limit hard.
Building a measurement chain that survives challenge
When a result is contested, the questions are predictable and mechanical. Who calibrated the instrument, against what traceable standard, and when. What was the drift at the next check. What qualifier flag does the value carry, and what does that flag mean in the documented vocabulary. Who had custody of the sample. Was the laboratory accredited under the TNI standard or ISO/IEC 17025 for that method. Can the reported value be recomputed from the raw record and produce the same digits. Every one of those is answerable in advance by system design.
- An approved Quality Assurance Project Plan exists before collection, following EPA QA/R-5 element structure, and the version is recorded with every dataset
- Raw instrument output is immutable and stored separately from any corrected or modeled layer
- Every transformation is versioned code with a pinned environment, so a recomputation two years later is bit-identical
- Every substituted, corrected, or model-adjusted value carries a qualifier and the magnitude of the adjustment
- Calibration, custody, audit, and maintenance records are joined to the measurements, not filed separately
- Electronic submissions meet the Cross-Media Electronic Reporting Rule at 40 CFR Part 3 for signature, copy of record, and integrity
- Retention meets or exceeds the applicable requirement, including the three-year minimum at 40 CFR 122.41(j)(2) for NPDES records
The architectural principle behind that list is a two-layer record. One layer is the regulatory record, computed by the prescribed rule, immutable once submitted, with a full audit trail. The second layer is the analytic environment, where models run, gaps get filled with better estimates, and forecasts are produced. The analytic layer reads from the regulatory layer and never writes back to it. Teams that blur those two layers end up unable to explain, under questioning, why the number in their system differs from the number on the form.
Where models genuinely earn their place
Fault and fouling detection. A fouling sonde produces a slowly diverging residual against its neighbors and against its own post-cleaning reading. Learned residual models catch that days earlier than a fixed threshold, which converts a discarded record into a graded one. The payoff is measured in recovered record days.
Collocation-based correction. Fitting a correction for a low-cost sensor network against a small number of reference monitors, with humidity and temperature as covariates, is a well-posed regression problem with published evaluation targets. Report the corrected value, the raw value, and the correction, and the network becomes usable for screening and public communication without pretending to be a reference method.
Surrogate estimation. Turbidity to suspended sediment, specific conductance to chloride, and similar relationships are established USGS practice. What makes them defensible is publishing the fitted model, the calibration sample set, and the prediction interval alongside every estimate.
Imagery change detection for inspection targeting. Segmenting disturbed acreage against a permitted boundary polygon, tracking impoundment extent, or finding thermal anomalies near an outfall gives a small inspection staff a ranked worklist. The output is a work order. The field visit remains the evidence.
Report assembly, never report numbers. Language models draft the narrative sections of a semiannual monitoring report or a Title V deviation summary well from structured inputs. They should read compliance values from the regulatory layer verbatim and never compute or restate them. Deterministic extraction and templating handle the numbers; the model handles the prose around them.
Common questions on the boundary between analytics and compliance
Can a modeled value ever be a reported compliance value?
Only when the applicable rule provides for it. Part 75 prescribes missing data substitution. Some permits allow calculated loading from flow and concentration. Some state programs accept documented surrogate relationships. Absent an explicit authorization, a modeled value is an internal estimate and belongs in the analytic layer with a qualifier.
How do low-cost sensor networks coexist with regulatory monitors?
As a spatial screening and communication tier anchored to the reference network. The reference monitors provide the truth for correction and for any regulatory comparison. The sensor tier provides density, which is what answers neighborhood-scale and fenceline questions that a network of a few dozen sites cannot.
What is the biggest hidden cost in an environmental data project?
Historical harmonization. Units, method changes, detection limit changes, station relocations, and qualifier vocabulary shifts across a twenty-year record consume more effort than modeling. Budget for it explicitly, and treat the harmonization mapping as a reviewable deliverable.
Does a citizen suit change how the system should be designed?
It raises the standard on reproducibility, since the plaintiff has access to the same public ECHO and AQS records. The practical answer is to design as though every reported value will be recomputed by someone else from the public record, because it can be.
Frequently asked questions
An approved quality assurance plan written before collection, an approved analytical method, traceable calibration, documented custody, qualifier flags with a published vocabulary, and a recomputable path from raw instrument output to the reported value. The statistical method matters less than the completeness of that chain.
It is routinely used to detect change, target inspections, and establish timelines. Enforcement determinations generally rest on ground truth: field observation, sampling, or an instrument record. Treat imagery as the layer that decides where to look, and design the field workflow that follows it.
As censored data, with the substitution rule stated in advance. Simple substitution at one half the detection limit is common and biased. Kaplan-Meier and regression-on-order-statistics approaches are better supported, and EPA's ProUCL implements them. What matters most is that the rule is fixed before the analysis, not chosen after seeing the result.
A quantitative statement of how good the data must be to support a specific decision, including tolerable false positive and false negative rates. The program office or permit holder sets it during systematic planning, following EPA QA/G-4, and it drives sample size, siting, and method selection.
Three years is the common floor, including for NPDES monitoring records under 40 CFR 122.41(j)(2), extendable by the permitting authority and often longer under other programs or during litigation holds. Design retention around the longest applicable requirement and keep raw data beyond the reported summaries.
