The Data Foundation Beneath Enterprise AI

·Research
A vast, dimly lit data-center hall in cool graphite and silver, long parallel rows of dark server cabinets receding toward a hazy vanishing point.

Introduction

Most enterprise AI initiatives fail, and the model is rarely the reason. Independent studies put the failure rate somewhere between stark and staggering - 95% of generative-AI pilots deliver no measurable profit-and-loss impact; more than 80% of AI projects fail outright. What breaks is the foundation the model sits on. Effective AI needs four layers working together - computation, cloud infrastructure, governed data infrastructure, and a semantic ontology - and most organizations, especially in blue-collar and field-service industries, have only the bottom two.

  • The failure: AI initiatives collapse on data readiness, integration, and governance - not on the algorithm.
  • The stack: compute and cloud are commoditized and rentable; the scarce layers are unified, governed, modeled data.
  • The wrapper trap: a prompt over a model API with no data infrastructure hallucinates the moment it meets a real operational dataset.
  • The fix: retrieval and reasoning over a governed data foundation and an ontology - not a bigger model or a bigger context window.

The 95% problem

In 2025, MIT’s Project NANDA published a study of enterprise AI adoption with a number that landed like a verdict: despite $30–40 billion in enterprise spending, roughly 95% of generative-AI pilots were delivering no measurable impact on the bottom line. Only about 5% were capturing real value. The report’s own diagnosis was not that the models were weak - it was a learning gap: tools that never connect to a company’s data and workflows, and so never improve.

That figure is dramatic, but it is not an outlier. The RAND Corporation, working from interviews with sixty-five experienced practitioners, reports that more than 80% of AI projects fail- about twice the failure rate of IT projects that do not involve AI. S&P Global Market Intelligence found the share of organizations scrapping most of their AI initiatives jumped to 42% in 2025, up from just 17% a year earlier, with the average company abandoning nearly half of its proofs-of-concept before production. Gartner, from the analyst’s chair, predicted that at least 30% of generative-AI projects would be abandoned after proof of concept by the end of 2025 - naming poor data quality first among the causes.

Four independent methodologies - an MIT field study, RAND’s practitioner interviews, a thousand-respondent survey, and an analyst forecast - arrive at the same conclusion from different directions. The failure is real, it is common, and it is getting more common, not less. And every one of these sources locates the cause in the same place: not the intelligence of the model, but the state of the data and the systems around it.

Four layers, not one

It helps to be precise about what an enterprise AI system actually is, because “we’re doing AI” usually means investing in one layer of a stack that has four. From the bottom up:

  • Computation - the accelerators, power, and cooling that run the models. This layer is scarce and expensive, but it is a utility: you rent it. McKinsey projects the industry will spend on the order of $6.7 trillion on data-center capacity by 2030, and a single vendor frames a “$3–4 trillion” accelerator opportunity over five years. None of that is a differentiator for the business that buys it.
  • Cloud data infrastructure - the elastic plumbing: object storage separated from compute, the lakehouse pattern, data pipelines, feature stores, vector databases, and the machine- learning operations around them. This is now a well-documented, off-the-shelf blueprint anyone can assemble from the same handful of vendors.
  • Data infrastructure and governance - the work of ingesting scattered records, unifying them, resolving what an entity is, cleaning and validating, and enforcing lineage, quality, and access. This layer is specific to your organization. It cannot be rented.
  • The semantic layer, or ontology - a shared, machine-readable model of the business: its entities, the relationships between them, and the actions that can be taken on them. This is the layer that turns unified data into something an AI can reason over and act within. It also cannot be rented.

The economics of the stack are the whole argument. The bottom two layers are commoditizing fast - a16z reports that model inference costs are falling by roughly an order of magnitude every twelve months, and the storage-compute separation that made elastic data warehousing cheap is a solved, published design. You can buy parity at the bottom of the stack with a credit card. The top two layers are the opposite: they are proprietary, they encode how your specific business works, and no vendor can hand them to you. That inversion is why so much AI spend underperforms. Organizations pour money into the layers they can purchase and starve the layers they have to build - and then wonder why the model on top has nothing solid to stand on.

This is not a fringe view. Gartner predicts that through 2026, organizations will abandon 60% of AI projects that are not supported by “AI-ready” data, and reports that 63% of organizations either do not have, or are unsure they have, the data-management practices AI requires. The tax is being paid already: Gartner puts the cost of poor data quality at an average of $12.9 million a year - a bill organizations run up long before a single model is trained. The researcher Andrew Ng has spent years arguing for “data-centric AI” - that for most practical applications it is now more productive to hold the model fixed and improve the data. The bottleneck moved up the stack, and most budgets did not follow it.

The blue-collar reality

Nowhere is that gap wider than in the industries that run the physical economy - pest control, HVAC, lawn care, plumbing, the trades. Picture the data estate of a real field-service operator doing a few million dollars a year. The financial truth lives in accounting software. The jobs, routes, dispatch, chemical logs, and customer history live in a separate field-service platform - a FieldRoutes, a PestPac, an Aspire - built to send trucks out, not to be reasoned over. Around those two anchors orbit a call-recording system, a stack of spreadsheets, technician notes typed into free-text fields, inbox threads, and literal paper. The single view of the business that a useful AI would need does not exist anywhere except in the owner’s head.

Fragmentation at this scale is not a small-business quirk; it is the norm, and it worsens as you go down-market. Even well-resourced enterprises, in MuleSoft’s benchmark, run an average of 897 applications and integrate only 29% of them, with 90% reporting business obstacles from data silos and 80% naming data integration as their single biggest barrier to AI. Splunk’s research on “dark data” suggests more than half of what an organization collects is never analyzed at all - which, for a field-service operator, describes exactly the voice recordings, photos, and freeform notes where much of the real signal hides.

And yet adoption is racing ahead of readiness. The U.S. Chamber of Commerce found that small-business generative-AI use jumped to 58% in 2025, from 40% a year earlier, with adoption reaching well into construction and manufacturing. The problem is what that adoption is actually buying. A survey of roughly a thousand small businesses found that 51% are “Explorers” who have not yet seen enough value to commit, and that three-quarters of them say they need clearer proof of return before they will. McKinsey, surveying far larger firms, found that only about six percent capture real bottom-line value from AI. If six in a hundred well-capitalized companies get there, an operator whose data lives in five disconnected places starts the race further back - not ahead. These businesses do not lack one layer of the stack. They lack all four.

What most operators actually buy: a wrapper

So what gets deployed into that estate? Almost always the same thing: a wrapper. A wrapper is a product, or an internal tool, whose entire substance is a prompt and a user interface sitting on top of someone else’s foundation-model API. It owns no data infrastructure - nothing that ingests, unifies, governs, and retrieves the customer’s own records - and no way to measure whether its answers are right. It is the model, dressed up, and nothing underneath.

The venture firms that fund this category are blunt about its limits. Sequoia argues that the durable advantages in applied AI are not thin wrappers and not raw data, but workflows, systems of record, and user networks, and that winning companies “use foundation models as a piece of a more comprehensive solution rather than the entire solution.” a16z’s reference architecture describes the alternative concretely: a real system is a data-engineering pipelinethat chunks and embeds your proprietary data, stores it, and retrieves only the most relevant records into each model call - which, as a16z puts it, “reduces an AI problem to a data-engineering problem most companies already know how to solve.” The 2023 internal Google memo that declared “we have no moat” made the cultural version of the point: if the base model confers no lasting advantage, a product that is the base model has even less.

Here is the mechanism that matters for an operator. A wrapper’s value depends on your large, messy, private operational data - but it has no infrastructure to ingest, unify, or retrieve any of it. Its only path to “know” your business is to paste a slice of that data into the prompt. So the entire enterprise gets compressed to whatever happens to fit in a context window, and everything else is invisible. That is not a small limitation. It is the difference between a demo and a system - and it fails in a specific, predictable way.

A model wrapperA grounded system
Data it ownsNone - only what fits in the promptIngests and unifies your operational records
How it knows your businessYou paste context in by handRetrieves the relevant governed slice per query
Against a large datasetDegrades and hallucinatesGrounded retrieval with evidence and provenance
Meaning of a metricGuessed - joins and definitions vary per answerDefined once in a semantic layer, resolved consistently
What it can doSuggests textActs within a governed operating model
Durable advantageNone - a markup on someone else’s APIYour unified, governed, modeled data

The context window is not a database

The seductive rebuttal is that context windows keep getting bigger - a hundred thousand tokens, a million, more - so why not just paste the whole dataset in? The research is unusually clear that this does not work, and understanding why is the technical heart of the argument.

First, the usable window is far smaller than the advertised one. NVIDIA’s RULER benchmarkfound that of models all claiming to handle 32,000 tokens or more, only a handful actually held their reasoning quality at that length; models ace the simple “needle in a haystack” retrieval demo and then crumble on anything multi-step well before their stated limit. Second, models do not attend to a long context evenly. The peer-reviewed “Lost in the Middle” study showed a U-shaped curve: accuracy is highest when the needed fact sits at the very start or end of the input and sags in the middle - and in the harder settings, feeding the model twenty or thirty documents scored worse than giving it no documents at all.

This is not a quirk of one old model. Chroma’s 2025 “Context Rot” study tested eighteen current frontier models and found performance declining steadily as input grew - even on trivial tasks, and even when the surrounding text was a coherent document rather than random filler, which shows the degradation is caused by length itself, not by difficulty. And when a query requires inferring a connection rather than matching a keyword - which is how a real operational question reads - long-context recall collapses: on the NoLiMa benchmark, a leading model fell from 99.3% at short context to 69.7% at 32,000 tokens, and eleven models dropped below half of their short-context score.

Now combine that with what a model does when it cannot find the answer. Hallucination - fluent output unsupported by the source - is a well-documented, systematic failure mode. A model that cannot reliably attend to the relevant fact buried in a giant context does not say “I couldn’t find it.” It fills the gap with a confident, plausible, wrong answer drawn from its general training. For a marketing chatbot that is an annoyance. For a system asked which accounts are past due, or which technician’s re-service rate is drifting, it is the worst possible failure: a manufactured number that looks exactly like a real one. The context window is a workspace, not a database - and treating it like a database is precisely the wrapper’s fatal move.

The fix is grounding, not a bigger model

The resolution is not a larger window but a smaller, better-curated one: retrieve the relevant, structured slice of the data and ground the model in it. This is the idea behind retrieval-augmented generation, introduced in a 2020 paper that paired a model’s internal memory with a non-parametric memory of retrieved evidence to improve factual accuracy. Chroma’s own results make the case almost absurdly concrete: a focused prompt of a few hundred relevant tokens outperformed handing the model a hundred-thousand-token context that contained the same answer - better answers, at a fraction of the cost and latency.

But retrieval is only as trustworthy as the data underneath it. Point a naive retrieval system at duplicated, deprecated, draft, and contradictory documents and it treats every source as equally authoritative and blends them into confident nonsense. Retrieval inherits the governance of whatever it retrieves - which is why it raises the value of a governed data layer rather than removing the need for one. Structure helps even more than clean text. Microsoft Research’s GraphRAGretrieves over a knowledge graph built from the source data rather than over loose text chunks, and answers whole-dataset questions that flat retrieval cannot - “populating the context window with higher-relevance content” and carrying provenance for every claim. In production, a knowledge-graph-grounded system at LinkedIn cut median issue-resolution time by 28.6% by preserving the relationships in the data instead of flattening them into prose. Grounding, not window size, is what makes AI reliable - and grounding has to happen in infrastructure, not in the prompt.

The layer that turns data into decisions

Unified, governed data is necessary, but there is one more layer above it that is easy to miss and decisive to have: the ontology. In its formal sense the term is old - the computer scientist Tom Gruber defined an ontology in 1993 as an explicit specification of a conceptualization, later refined to a formal, shared vocabulary of the classes, relationships, and rules of a domain. In plainer terms: it is a machine-readable model of what your business is made of - customers, sites, routes, jobs, technicians, invoices - how those things relate, and what can be done with them.

Why does an AI need this? A governed data warehouse gives a model something to read; an ontology gives it an operating model to act within. dbt Labs makes the practical version of the argument: without a governed semantic layer, an AI “guesses at joins, filters, and time grains, and each model guesses differently” - so the same question returns different numbers. Define each metric and entity once, and every query resolves the same way. The most complete commercial articulation of the idea, Palantir’s Ontology, goes a step further by pairing a semantic layer (the objects and their relationships) with a kinetic one (the governed actions that change real systems, under inherited permissions and an audit trail). An agent grounded that way does not just answer questions about the business - it traverses the same objects a human operator would and takes the same governed actions. That is the difference between an AI that can query your data and one that can operate your business.

The objections - and why they don’t let you skip the foundation

A fair argument has to meet its strongest counterpoints, and there are three good ones.

“Foundation models and RAG mean you don’t need governed data anymore - just point a model at your documents.” Retrieval genuinely lets teams query unstructured knowledge without heavy upfront modeling, and that is real progress. But, as above, retrieval inherits the governance of what it retrieves, so pointed at an ungoverned pile it produces confident errors. And documents structurally cannot answer the operational, quantitative questions a business actually runs on - past-due balances, route density this week, a drifting re-service rate - which require governed structured data and a metric defined once. RAG is a technique layered on top of a governed foundation; it raises the value of that foundation rather than removing it.

“Just buy a vertical AI product and skip the infrastructure.” This is the strongest objection, and the evidence partly supports it: the same MIT NANDA study found that purchased tools succeed far more often than internal builds. But the reason they win is instructive - the vendor already did the integration and iteration work that internal teams skip, not that the vendor’s model is smarter. A vertical product with no access to your unified operational data is still a demo. Buying is often the right call; it simply relocates who builds part of the stack. It does not let you skip the governed data foundation and the ontology that maps a tool into your own operations.

“A semantic layer is heavy, top-down data modeling - the old enterprise data model that already failed.” The best critic here is the data-mesh movement, whose entire premise is that the monolithic, centralized, model-everything-up-front data platform became the bottleneck every request queued behind. That critique is correct - but the failure was the monolith and the boil-the-ocean delivery, not the existence of shared semantics. The modern answer decentralizes the same idea into domain-owned data products and incremental, operationally scoped ontologies tied to real actions. The layer is still required; the way it used to be built is what failed.

From foundation to operations

The reason so much AI underdelivers is not a shortage of intelligence. The models are, for most business problems, more than good enough. What is missing is everything the model needs to be useful: data pulled out of its silos and unified, governed so it can be trusted, modeled so it carries meaning, and retrieved so the model reasons over a real substrate instead of a prompt-sized guess. Harvard’s Iansiti and Lakhani made the strategic version of this point years ago - the durable core of a modern firm is a decision engine built on curated data, not the algorithm that reads it.

This is precisely the discipline Ardenus brings to the operators who run the physical economy. Rather than bolt a wrapper onto a fragmented data estate, Ardenus sits on top of the systems a business already runs - unifying the data, governing it, modeling it into an ontology, and acting on it. It is the intelligence layer the four-layer stack is missing, built for the industries most likely to be sold a chatbot and least equipped to make one work. You can read more of our research and writing on the Ardenus articles hub, or see the platform itself on the technology page.

Sources and methodology

  1. The GenAI Divide: State of AI in Business 2025, MIT Project NANDA (2025), as reported by Fortune - the ~95% of enterprise generative-AI pilots with no measurable P&L impact.
  2. The Root Causes of Failure for Artificial Intelligence Projects, RAND Corporation (Ryseff, De Bruhl & Newberry, 2024) - the >80% AI-project failure rate and its organizational root causes.
  3. AI project failure rates are on the rise, CIO Dive, reporting S&P Global Market Intelligence (2025) - 42% of firms abandoning most AI initiatives, up from 17%.
  4. Gartner Predicts 30% of Generative AI Projects Will Be Abandoned After Proof of Concept (2024) and Lack of AI-Ready Data Puts AI Projects at Risk (2025).
  5. The state of AI: How organizations are rewiring to capture value, McKinsey & Company / QuantumBlack (2025), and The cost of compute (2025).
  6. 2025 Connectivity Benchmark Report, MuleSoft / Salesforce - 897 applications per enterprise, 29% integrated; and Splunk’s research on dark data.
  7. Empowering Small Business: The Impact of Technology on U.S. Small Business, U.S. Chamber of Commerce (2025), and the Reimagine Main Street AI and Small Business Survey (2025).
  8. Lost in the Middle: How Language Models Use Long Contexts (Liu et al., TACL 2024); RULER (NVIDIA, 2024); NoLiMa (ICML 2025); and Chroma’s Context Rot (2025).
  9. Survey of Hallucination in Natural Language Generation (Ji et al., ACM Computing Surveys, 2023).
  10. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (Lewis et al., 2020); From Local to Global: A Graph RAG Approach (Microsoft Research, 2024); and Retrieval-Augmented Generation with Knowledge Graphs for Customer Service (SIGIR 2024).
  11. A Translation Approach to Portable Ontology Specifications (Gruber, 1993); the Palantir Ontology documentation; and dbt Labs on governed metrics for trustworthy AI.
  12. Emerging Architectures for LLM Applications and How 100 Enterprise CIOs Are Building AI in 2025, Andreessen Horowitz; Generative AI’s Act Two, Sequoia Capital; Andrew Ng on data-centric AI (IEEE Spectrum); and Competing in the Age of AI (Harvard Business Review).