Contamination Without Borders: Why America's Fractured Brownfield Data Infrastructure Is Costing Remediation Professionals Time, Money, and Accuracy
Ask an environmental consultant working a brownfield assessment in Gary, Indiana, what they know about the contamination plume two miles north—the one that crosses the county line and falls under a different state agency's jurisdiction—and the answer is often some version of: not much, and not easily findable. This is not a failure of professional diligence. It is a structural failure of the data systems on which the entire remediation enterprise depends.
The United States does not have a unified national database for soil and groundwater contamination. What it has is an archipelago of state databases, EPA program-specific registries, county records, and private consultant archives that were built independently, maintained inconsistently, and were never designed to communicate with one another. The consequences of this fragmentation are not merely administrative. They shape which contamination gets identified, how thoroughly it is characterized, and whether remediation efforts address the full extent of a problem or only the portion that happens to fall within a single agency's data horizon.
Mapping the Fragmentation
The architecture of contamination data in the United States reflects the federated nature of environmental regulation itself. The EPA administers several program-specific databases—the Superfund site inventory, the Resource Conservation and Recovery Act (RCRA) corrective action database, the Brownfields program tracking system—each capturing a subset of contaminated sites according to program-specific criteria and data standards.
State environmental agencies maintain parallel systems, typically organized around state-specific remediation programs that may use different chemical detection thresholds, different site classification criteria, and different data formats than their federal counterparts or their neighboring states. A site that qualifies for active remediation under New Jersey's standards may not meet the threshold for listing under Pennsylvania's system, even if the contamination plume crosses the state line.
Below the state level, county health departments and municipal planning agencies often maintain their own records of known or suspected contamination, derived from local permitting histories, complaint investigations, and legacy industrial records. These records are rarely digitized in any consistent format, and their relationship to state databases is typically informal at best.
Private environmental consulting firms add another layer. Firms that conduct Phase I and Phase II Environmental Site Assessments generate substantial contamination characterization data that, in most states, is reported to the relevant agency but not retained in any publicly accessible standardized format. The data exists in consultant archives, client files, and agency submission records—distributed across formats and access protocols that make systematic cross-referencing essentially impossible without extensive manual effort.
What Fragmentation Costs in Practice
The practical consequences of this landscape manifest in several distinct ways.
Duplicated characterization work is perhaps the most straightforward cost. When a consultant begins a brownfield assessment without access to prior investigation data for adjacent or nearby properties, they may repeat sampling and analysis work that has already been conducted—sometimes multiple times by different firms over successive years. Industry practitioners have estimated that duplicated subsurface investigation represents a significant fraction of total brownfield assessment expenditure nationally, though precise figures are difficult to establish precisely because the duplication itself is invisible when it occurs.
Missed plume connections represent a more consequential problem. Groundwater contamination does not respect parcel boundaries, county lines, or state jurisdictions. A chlorinated solvent plume originating from a dry-cleaning operation in one municipality may migrate through the aquifer for miles before emerging as a drinking water concern in a neighboring jurisdiction. When the agencies responsible for those jurisdictions maintain separate databases with no automated cross-referencing capability, the connection between source and receptor may not be identified until significant additional exposure has occurred—or may never be identified at all.
Inconsistent remediation standards compound both problems. When different jurisdictions apply different cleanup criteria to the same contaminants, the result is a patchwork of remediation outcomes that does not reflect the actual risk distribution across a region. A plume remediated to one state's standards at its point of origin may still carry residual concentrations that exceed another state's standards at the point where it crosses the border—a discrepancy that neither agency's data system is equipped to surface.
The Architecture of a National Registry
Proposals for a unified national contamination database are not new. They have circulated in environmental policy discussions for at least two decades, consistently encountering the same set of objections: jurisdictional resistance from state agencies protective of their regulatory autonomy, liability concerns from property owners and responsible parties worried about expanded data disclosure, and the sheer technical challenge of harmonizing incompatible legacy systems.
These objections are legitimate, and any viable national registry architecture must address them directly rather than dismissing them as obstacles to progress.
On the jurisdictional question, the most promising models are federated rather than centralized. Rather than requiring states to surrender control of their data to a federal repository, a federated architecture would establish shared data standards, application programming interfaces (APIs), and query protocols that allow state systems to remain independently administered while becoming interoperable. This approach has precedent in other federal data domains, including aspects of the National Environmental Policy Act's environmental review data infrastructure and certain public health surveillance systems.
On liability, the registry's design would need to incorporate clear statutory protections distinguishing data disclosure for public interest purposes from admissions of liability. Similar frameworks exist in other regulatory contexts—the Emergency Planning and Community Right-to-Know Act's Toxics Release Inventory, for instance, requires disclosure without treating reported releases as automatic admissions of wrongful conduct. Extending analogous protections to a contamination registry would require legislative action but is not without precedent.
On technical harmonization, the challenge is substantial but tractable. A phased approach—beginning with standardized minimum data elements that all state systems would be required to capture and report, before moving toward deeper integration of legacy records—would reduce the upfront technical burden while establishing the interoperability foundation that more ambitious integration could eventually build upon.
What the Environmental Science Community Can Contribute
The case for a national contamination registry is ultimately an empirical one, and the environmental and chemical science community is well-positioned to build it. Systematic documentation of duplicated investigation work, quantitative analysis of plume connectivity failures, and comparative studies of remediation outcomes across jurisdictional boundaries would collectively constitute an evidence base that policy advocates and legislative staff could act upon.
Professional organizations in the environmental consulting and remediation sectors have begun to engage with this issue, but the research dimension remains underdeveloped. Academic environmental science programs, in collaboration with practicing consultants and state agency data managers, could contribute substantially by developing rigorous methodologies for measuring the cost of fragmentation—both in direct expenditure terms and in public health outcomes.
Conference environments that bring together researchers, regulators, and practitioners from multiple states and disciplines are particularly well-suited to advancing this kind of cross-jurisdictional analysis. The contamination data fragmentation problem is, at its core, a coordination failure—and coordination failures are most effectively addressed when the parties who bear their costs are in the same room, working from shared evidence.
The Argument for Acting Now
The brownfield redevelopment pipeline in the United States is substantial. EPA estimates that there are more than 450,000 brownfield sites nationwide, and demand for brownfield remediation and reuse has increased significantly as urban land constraints and infrastructure investment priorities have converged. That pipeline will be navigated more efficiently, more accurately, and more equitably if the data systems supporting it are fit for purpose.
Fit for purpose, in this context, means interoperable, comprehensive, and accessible to the professionals who depend on them. The current system is none of these things. Building something better will require political will, technical investment, and sustained engagement from the scientific and professional communities that understand most clearly what is being lost in the gaps between databases.