Open Data Initiatives and Public Value Creation

The GovLab's Open Data's Impact framework identifies four pathways through which open data creates value: improving government, empowering citizens, creating economic opportunity, and solving large public problems. They are not equally weighted, and the mechanisms differ enough that treating them as interchangeable will send you down the wrong road entirely.
Improving government is the least glamorous pathway and often the most immediate. Government is frequently its own first beneficiary. When agencies publish data in shared, standardized formats, they reduce redundant acquisition across ministries, lower internal costs, and make cross-agency analysis possible for the first time. The efficiency gain happens before any citizen or entrepreneur touches the dataset. Nobody writes headlines about it. It happens anyway.
Empowering citizens gets the most rhetorical attention, and for good reason. Access to government information is a foundational tenet of democratic governance; scholars Zuiderwijk and Janssen have argued directly that open data policies connect to citizens' rights to public information. But empowerment requires more than access. It requires data legible enough for civic organizations and constituents to act on. Publishing data and sharing trustworthy information are different acts with different outcomes, and conflating them is how well-intentioned programs quietly fail.
The economic opportunity pathway is where the numbers become striking. According to UNDP, citing OECD research, open government data fosters job creation, resource savings, and productivity gains, with entrepreneurs transforming raw data into products and services that generate value far beyond the original dataset. The market for public sector information already runs into the tens of billions globally and is projected to reach roughly $200 billion by 2030.
The cleanest single illustration of how this works: the United Kingdom published hospital heart surgery survival rates. Hospitals began competing on outcomes. Survival rates improved by 50%. Publication enabled comparison, comparison drove competition, competition produced better medicine. Causation is visible. No interpretation required.
Solving large public problems is the fourth pathway, and the most contingent. It requires capable actors, policy environments that reward evidence-based decisions, and time horizons long enough for data-driven interventions to mature. The potential is significant. The realization is uneven. That gap is not random, and pretending otherwise is the reason so many open data initiatives underdeliver.
What the NOAA Weather Data Case Reveals About the Distance Between Data and Value
The National Oceanic and Atmospheric Administration collects over 3.5 billion weather observations per day. Around 96% of the U.S. public obtains some 301 billion forecasts per year, the vast majority touching NOAA data somewhere in the chain. Annual benefits from this ecosystem have been estimated at $31.5 billion, against roughly $5.1 billion spent annually by public and private weather bureaus combined. That ratio is the clearest available demonstration of what open data can do when conditions are right.
NOAA did not create that consumer value. Private actors did. The Weather Channel reaches roughly 97.3 million American households using NOAA data as its foundation. The Climate Corporation built a weather insurance business on the same underlying data and sold to Monsanto in 2013 for $930 million. The government built the infrastructure and released the data. The market built the products. Those are two distinct functions, and conflating them is where a lot of open data strategy goes sideways.
What made this possible is not complicated, but it is demanding. NOAA's data is consistent and high-quality across decades. Formats are standardized enough that commercial developers can build without bespoke engineering for every application. The legal status is unambiguous, with no reuse restrictions creating compliance risk. And over time, an ecosystem of intermediaries developed the capacity and incentive to invest in building on top of it. All four of those conditions had to be true simultaneously, and in most government datasets, none of them are.
The NOAA case is exceptional, which is exactly why it is instructive. It does not describe the average open data release. It describes what the average open data release would need to become before it can generate comparable returns. Most open data optimists would rather not dwell on how long that journey actually takes.
The Structural Barriers That Keep Published Data From Becoming Usable Data
Most open data failures are failures of infrastructure, quality, and intermediary capacity rather than failures of political will. The remedies are completely different depending on which problem you actually have. Political will problems call for advocacy. Infrastructure problems call for investment and technical standards. Conflating them produces the wrong interventions, and wrong interventions waste the political will you do have.
The technical barriers are the most tractable and the most commonly neglected. Non-machine-readable formats, inconsistent schemas, absent metadata, no API access: under these conditions, data is technically published but practically unusable. The U.S. OPEN Government Data Act addresses this directly, requiring federal agencies to publish in standardized machine-readable formats and include metadata in the Data.gov catalog. The mandate exists because the default, without it, is inconsistency. Full stop.
Quality and trust are subtler and more damaging over time. A researcher or entrepreneur who builds on a dataset and later discovers it is unreliable does not simply stop using that dataset. They become skeptical of the entire category. That skepticism spreads through professional networks faster than any government communications campaign can counteract it, and rebuilding trust after a quality failure costs more in time and credibility than maintaining quality would have in the first place.
The intermediary gap is the least-discussed barrier and possibly the most consequential. Raw data rarely creates value directly. It requires actors with the technical capacity and the incentive to transform it into something actionable. NOAA did not produce The Weather Channel. An ecosystem of intermediaries did, over years, with real capital at risk. Most open data releases enter a world where that ecosystem is thin, underfunded, or simply does not yet exist. Releasing data into that vacuum and then measuring low uptake as a data failure is a misdiagnosis. It is a market development problem.
Institutional siloing compounds everything else. Agencies often cannot find or effectively access their own data, which means the efficiency gains government itself would capture go unrealized before any external user even enters the picture. And even where data is technically open, ambiguous licensing or fragmented regulatory regimes create hesitation among commercial actors facing compliance risk. A developer who is uncertain whether they can legally build a product on a dataset will often simply move on. That hesitation is invisible in usage statistics and almost never gets counted in assessments of why a dataset failed to generate value. It should be.
How Major Policy Frameworks Are Trying to Close the Gap, and Where They Differ in Approach
The U.S. approach is fundamentally about mandate and infrastructure. The OPEN Government Data Act establishes enforceable publishing requirements: standardized formats, machine-readable outputs, metadata in the Data.gov catalog. The National Secure Data Service, implemented by the National Science Foundation and active as of late 2025, provides infrastructure for discovering and linking statistical data across agencies. It targets the cross-agency silo problem directly, which is notable because that problem has historically been treated as cultural rather than structural. It is both, but the structural fix has to come first.
The European Union centers its approach on designation rather than general mandate. Commission Implementing Regulation 2023/138, applied across all Member States since 2024, identifies six high-value dataset categories: geospatial, earth observation and environment, meteorological, statistics, company ownership, and mobility. Public bodies must make these available free of charge, in machine-readable formats, via APIs and bulk downloads, with cross-border reuse as an explicit design goal. The logic is prioritization. Not all data is equally valuable, so urgency should be allocated accordingly. The Digital Omnibus proposal, under development through 2025 and 2026, takes a different angle: consolidating the Open Data Directive, the Data Governance Act, and related instruments into a simplified framework. The diagnosis is that regulatory fragmentation itself has become a barrier, with overlapping rules creating compliance uncertainty that chills reuse even when the data is technically available. Too many rules can produce the same outcome as too few.
Emerging jurisdictions are broadening the frame in ways that deserve attention. Singapore's Public Sector (Governance) (Amendment) Bill extends data sharing to trusted external partners, including social service organizations and industry associations, with ministerial authorization and contractual safeguards. The framing is coordinated service delivery, not transparency as an end in itself. The African Union, through ACHPR Resolution 620 in 2024, articulates a human rights-based framework calling for maximum disclosure, national open data policies, and independent oversight. A rights framing gaining traction beyond high-income contexts matters because it changes who can make enforceable demands. Ukraine amended its open data regulations to strengthen the "open by default" principle, expand mandatory datasets, and introduce AI-assisted dataset moderation, embedding artificial intelligence into the governance mechanism itself rather than treating AI solely as a downstream use case.
The direction of travel across all of these is consistent. Every jurisdiction examined is moving from voluntary disclosure toward structured, enforceable obligations. The era of publishing what you like and calling it open data is closing.
The Economic Potential That Gives These Policy Efforts Their Urgency
The public sector information market, standing at roughly $60 billion and projected to reach $200 billion by 2030, is a conservative and measurable signal of what is at stake. It measures commercial activity around public data, not total societal benefit. Total societal benefit is larger and harder to count, which is exactly why the market figure is useful: it undersells the case, and it still commands attention.
McKinsey's widely cited estimate, that open data could unlock trillions of dollars in annual value across multiple sectors of the U.S. economy alone, predates the maturation of the smartphone era, the platform economy, and AI-driven data use. Treat it as an illustration of magnitude rather than a precise forecast. The actual potential is likely larger and differently distributed than that analysis captured. The direction of the estimate is not in dispute; the exact number matters far less than the order of magnitude it signals.
Economic value from open data materializes in three places. Commercial innovation, as the NOAA case demonstrates, produces new products and services that generate revenue and employment. Efficiency gains within government reduce redundancy and lower operating costs in ways that compound over time, mostly invisibly. Research and evidence-based policymaking are harder to quantify but central to long-run governance quality. A government that makes better decisions because it has better information generates diffuse, durable value that never appears on a balance sheet and rarely gets credited to data policy. It should be, but the accounting frameworks do not exist yet to capture it.
The critical qualifier is distribution. The economic case for open data is well-established in aggregate. Value concentrates where high-quality data meets capable intermediaries in a permissive legal environment. Everywhere that combination is absent, potential sits unrealized, sometimes indefinitely.
A Countercurrent: The Forces Now Restricting Data Access and What a "Data Winter" Would Cost
Stefan Verhulst of NYU and The GovLab has described the current moment as the beginning of a data winter: a period marked by the growing inaccessibility of data critical for science, governance, and innovation. The analogy to previous AI winters is deliberate. Those were periods of stagnation caused not by lack of interest but by conditions that made progress impossible regardless of intent. The warning is that data access is becoming one of those conditions.
Verhulst identifies eight interrelated forces driving the contraction: government open data cutbacks, institutional risk aversion, AI-induced data hoarding by companies and governments treating datasets as competitive assets, research data lockdowns, scarcity of high-quality training data, risks from synthetic data substitution, geopolitical fragmentation, and private-sector closure. These are not independent variables. A government that cuts open data programs creates a vacuum that private actors fill with proprietary systems, which drives more hoarding, which makes publicly available alternatives scarcer, which makes the empirical case for open data harder to argue, which makes further cuts easier to justify politically. The cycle is self-reinforcing, and once it gains momentum, it is very hard to reverse.
The core problem is this: demand for data is expanding rapidly, driven by AI development, research needs, and complex governance challenges, while the access models, trust frameworks, and reuse governance structures that enable the data ecosystem to function are not keeping pace. What a data winter would actually cost is not speculative. Evidence-based policymaking becomes harder when the data governments need for their own decisions is locked in proprietary or fragmented systems. AI models trained on closed or synthetic data compound existing biases and reduce reliability for public-interest applications. The economic multiplier effects documented in cases like NOAA's require an intermediary ecosystem that remains open and accessible. Closure does not pause that ecosystem. It forecloses the applications that would otherwise exist, applications we cannot name yet because they depend on access we have not yet granted.
How Generative AI Is Reshaping Both the Opportunity and the Obligation Around Open Data
The Open Data Policy Lab has described the current moment as a potential fourth wave: open data and generative AI used in tandem, with open data serving as training and fine-tuning material and generative AI serving as an interface that makes datasets accessible to non-technical users. Previous waves each required new infrastructure to realize their potential. The release era needed portals. The interoperability era needed standards. The impact measurement era needed evaluation frameworks. This wave needs AI-ready data and deliberate governance of the intersection between models and sources. The infrastructure requirement is different in kind, not just degree.
The access democratization argument has real force. Generative AI enables users without technical backgrounds to engage with complex datasets through conversational interfaces, directly addressing the intermediary gap that has constrained open data's reach for years. The U.S. Department of Commerce launched an AI and Open Government Data Working Group in late 2023 specifically to assess how to improve this intersection. The logic is straightforward: if you can reduce the barrier to data use from "hire a data scientist" to "ask a question," the pool of people who can derive value from open data expands considerably. That is not a marginal gain. It is a structural shift in who the policy actually serves.
Getting agency data AI-ready requires the same things good open data has always required: clean datasets, better descriptions, consistent machine-processable formats. What is new is that "AI-ready" reframes data quality from a bureaucratic compliance obligation into a genuine strategic asset. A dataset that a language model can use effectively can reach orders of magnitude more people than one that requires a specialist to interpret. That changes the cost-benefit calculus around investment in data quality, and it changes it in favor of investment, which has historically been the harder political argument to make. AI has arguably accidentally solved that particular problem.
The tensions AI introduces are real, though. Large language models trained on open data can surface that data more widely, but they also abstract away the source, making provenance and trust harder to maintain for downstream users who do not know where a given piece of information originated. Data hoarding by AI companies, one of the forces Verhulst identifies as driving the data winter, reduces the pool of high-quality open training data available for public-interest model development. The same technology that expands open data's reach simultaneously creates incentives to restrict the inputs that make open data valuable. That tension does not resolve cleanly. It requires active governance, not optimism.
GSA has responded by drafting an AI Strategic Plan that explicitly treats open government data as infrastructure for training, testing, and validating AI models in partnership with academia, industry, and civil society. Open data is not just a transparency mechanism or a civic good. It is the substrate on which trustworthy AI for public purposes gets built, and protecting it, improving it, and governing the relationship between open datasets and AI systems is not a separate mission from open data. It is where the open data mission is going next.


