Blog
What Open Data Stewards Can Learn from Corporate Guidance on AI-Ready Data
Posted on 28th of July 2026 by Stefaan Verhulst
For two decades, the open data community has worked to articulate what it means to make data fit for use, converging on the now-canonical FAIR principles: Findable, Accessible, Interoperable, and Reusable. As we argue in Moving Toward the FAIR-R Principles, this framework, though foundational, is no longer sufficient in an era of artificial intelligence. It must be extended to include a fifth commitment: Ready for AI.
The conversation over data readiness for the AI age is not confined to the public-interest domain. Parallel discussions have unfolded in the corporate sector, where the imperatives of AI readiness have been confronted at scale, and it is worth asking what the open data movement might learn from them.
McKinsey’s recent article, AI Data Readiness: The Key to Scaling Impact, is instructive in this regard. On its surface, the article reads as a memo to chief data officers diagnosing why corporate AI pilots so often stall before reaching production. It also offers some valuable lessons and principles for data stewards operating in the public interest, and who may be considering how to apply FAIR-R principles in practice.
In what follows, we treat this private-sector-focused article as a source of transferable knowledge while still remaining attentive to its limits. The caveat about limits is consequential, and we return to it below: in general, while enterprises optimize primarily for reliable outputs and cost avoidance, public-interest stewardship must also account for equity, openness, and public accountability. Some lessons therefore transfer cleanly, while others may require translation (or simply not apply at all).
To be clear, our premise is not that public-interest data work should become more corporate. It is, rather, that firms have been contending for some time with the mechanics of AI-ready data (at scale, under audit, and with material consequences for failure) and that the open data movement stands to learn from at least some of these experiences. While this piece focuses on a specific article from McKinsey, many of these observations are more generalizable to the private sector and the way it is engaging with data more broadly.
Six lessons worth borrowing
1. Govern the fragment, not only the file, because AI dissolves the dataset.
This is perhaps the single most valuable idea to import. FAIR-R, like the FAIR tradition it extends, largely treats the dataset as the object that needs to be made ready. McKinsey’s article, on the other hand, observes that the dataset is no longer the operative unit within an AI pipeline. A single PDF or data set fragments into extracted text, parsed tables, image summaries, chunks, embeddings, and index entries, each a derived artifact capable of independently shaping an output. A document may be wholly accurate in its entirety yet yield an erroneous answer when an outdated or incomplete fragment is retrieved in isolation.
The implication for open data stewards is significant: provenance, sensitivity labels, and usage rules applied at the level of the dataset no longer survive the ingestion into an AI system. It therefore implies that every chunk carries its lineage, not merely every file.
2. Define “good enough” for each use case, and treat quality as a continuous process.
The article emphasizes that companies have learned to resist the instinct to perfect data before deployment. McKinsey’s framing calibrates quality to the use case and its attendant risk profile: the governance appropriate to an internal drafting tool differs markedly from that required for a regulated decision, where full lineage and auditability become essential (see illustration below).
For open data stewards, this reframing is methodologically valuable, in that it replaces the untenable mandate of universal data quality with a defensible, risk-tiered standard (while still countering the emerging assumption that data quality is of diminished importance in AI environments). The corollary is that quality assurance shifts from a one-time gate applied at publication to a continuous discipline exercised across extraction, retrieval, and generation.

From McKinsey Article: AI data readiness: The key to scaling impact June 23, 2026
In most cases, McKinsey’s metric of success is financial. For example, the report documents how, in the absence of shared infrastructure, teams often rebuild their own extraction logic and retrieval configurations, producing inconsistent outputs and duplicating cost. By contrast, McKinsey shows how a financial services firm converted reusable pipelines into an estimated $10–20 million in cost savings as use cases multiplied.
Translated into the open data ecosystem, this mechanism for cost savings is more than simply an argument for reusable infrastructure; it is an argument for reuse as an organizing principle, and it resonates with a commitment the open data world already holds, for example, in its championing of the data commons. The commons is not simply a shared asset to be drawn upon; appropriately designed, it can be a governed arrangement built deliberately to enable equitable, accountable reuse across many actors and purposes. While FAIR-R frames the commons primarily as a mechanism for equity, the corporate evidence demonstrates that the same architecture can also yield efficiency. For under-resourced public institutions, this dual justification is strategically consequential: shared, governed pipelines are both fairer and less costly, an argument that carries particular weight with those who control budgets.
4. Make traceability infrastructure.
When a single answer stitches together fragments from many documents, the ability to explain where the answer came from (i.e., the output of an AI system) collapses. McKinsey calls the result “indefensible” outputs, which are problematic in regulatory audits or legal discovery.
Public-interest deployments face an even sharper version of this problem: humanitarian, health, and civic uses require accountability to the people the data describes. According to McKinsey, corporations treat lineage as an essential characteristic with owners, versions, and retirement dates. Open data stewards should borrow that seriousness. A traceable chain of custody from source through every derived artifact to the final output is the deliverable, not a nice-to-have appended to it.
5. Expect the data steward’s role to broaden.
McKinsey describes an expanding mandate for chief data officers, from owning pipelines and warehouses to owning the standards and conditions. The article is explicit that this is a shift in kind rather than degree: the CDO becomes responsible not for the data alone but for the conditions under which data can be reused, traced, and governed consistently across the enterprise; linking structured and unstructured content, anchoring both to shared business entities, and ensuring that the tools and skills AI systems rely on behave predictably. As that mandate widens, McKinsey notes, the surrounding roles begin to blend, with data, engineering, product, and governance functions converging and organizations obliged to cultivate multifaceted skill sets rather than treating these as separate domains.
Our work on FAIR-R implies the same trajectory for public-interest data stewards. The lesson is to anticipate the shift deliberately: the data steward of the near future architects the conditions for systematic, sustainable, and responsible reuse rather than curating static holdings.
6. Extend readiness to unstructured and machine-generated data.
The FAIR tradition developed around structured, discrete datasets, yet the McKinsey report locates the AI-readiness problem primarily in unstructured content (documents, emails, transcripts, images, and video), which now drives a growing share of consequential decisions and which does not “stay whole” as it passes through extraction, chunking, and embedding. McKinsey also observes that AI systems do not merely consume data but generate it continuously: prompts, responses, summaries, and decisions accumulate and frequently flow back into core systems, where they seed further outputs and create feedback loops over time.
For open data stewards, this reality carries two implications. First, readiness cannot be confined to the tidy, tabular holdings the FAIR principles were first designed around; the harder and more valuable work increasingly concerns unstructured sources, whose provenance and sensitivity are far less legible. Second, the movement must begin to reckon with synthetic and machine-generated data as objects of stewardship in their own right (a matter we raised earlier), since data produced by AI systems can propagate error, especially if reused without lineage or quality controls. Readiness, in other words, is not a property to be established once at publication but a discipline that must follow data through its generation as well as its consumption.
What not to borrow
Borrowing is not copying. As we stated at the outset, we are interested in translation rather than repetition. Two differences in particular are important to highlight.
First, corporations govern risk to protect the institution: its outputs, its liability, its costs. Public-interest stewardship also governs risk, but to protect data subjects and the public. Our FAIR-R framework is more pronounced on this, insisting, for instance, that data be evaluated for bias and that a dataset that cannot be de-biased simply should not be used. McKinsey’s risk framing, managing exposure after AI accesses data, and applying human oversight proportional to business risk supply excellent mechanics, but the threshold for what counts as unacceptable risk must be set by public-interest values, not solely for-profit business ones.
Second, the corporate default is enclosure: governed and controlled, with data treated as a proprietary asset. The open data default, by contrast, is responsible openness. The reusable-foundation lesson is sound, but data stewards must resist letting “governed” quietly become fully “closed.” The goal is governed openness: commons that are equitable and accountable and that guard against extraction.
To conclude, McKinsey’s report can be of value to the open data world not as a mirror but as a field report from those who are rapidly experimenting with what AI-readiness means. Its transferable lessons provide valuable frontline experience for our FAIR-R framework and perhaps add some cautionary lessons as well.
Thanks to Adam Zable and Akash Kapur for editorial review.