The 5-Rung Data Maturity Ladder, Explained (With a Self-Score)
A data maturity model is a framework that scores how an organization collects, manages, and uses its data — from manual spreadsheets to a governed, AI-ready platform. Most models use four or five stages; imPROVE’s version has five rungs: Manual, Visible, Automated, Governed, and AI-Ready, each describing a concrete operational reality, not an abstract score.
What is a data maturity model?
A data maturity model is a framework that describes how an organization collects, manages, and uses its data, from manual spreadsheets to a governed, AI-ready platform. Most models use four or five stages. imPROVE’s version has five rungs — Manual, Visible, Automated, Governed, and AI-Ready — each describing a concrete operational reality, not an abstract score.
The ladder matters because “we need better data” is not a plan, it’s a feeling. Every operation we’ve worked with — waste haulers, insurers, manufacturers, life sciences companies, e-commerce brands, nonprofits — has data problems that look different on the surface and identical underneath: nobody agrees on the numbers, someone spends their week rebuilding the same report, and the people making decisions are working from a spreadsheet that’s already out of date. The ladder gives that feeling a name and a next step. It’s also the foundation this site is built on — see the parent framework at AI & Analytics for data-rich operations for how the ladder connects to AI readiness specifically.
Two things the ladder is not. It’s not a maturity contest — most healthy, profitable companies sit at L2 or L3 and that’s fine for what they’re doing today. And it’s not industry-specific — a waste hauler exporting route data from TRUX and an insurer reconciling claims in three different systems are climbing the same five rungs, just with different source systems.
What are the 5 rungs of the data maturity ladder?
The five rungs are Manual (L1), Visible (L2), Automated (L3), Governed (L4), and AI-Ready (L5). Each describes how data moves through a business day to day — who touches it and how much anyone trusts the number on screen. Moving up one rung means removing a manual step or adding a shared definition, not buying a bigger tool.
Below we walk each rung with what it actually looks like on the ground, what breaks at that level, and what’s required to move up. Two of the rungs use a waste-industry example, since that’s the vertical we know best, but the pattern holds regardless of what your business does.
What does Level 1 (Manual) look like?
At Level 1, data lives in spreadsheets that people build by hand, on a schedule, from memory of where last month’s file went. There’s no shared source of truth — there’s a source of truth per person, and they don’t always agree.
In a waste and recycling operation, L1 looks like this: a dispatcher exports route and stop data from TRUX every Friday into a spreadsheet, finance pulls accounts-receivable aging into a separate spreadsheet from the billing system, and neither file talks to the other. Nobody has written down what “active account” means, so the two departments quietly use two different definitions. The same pattern shows up everywhere else it happens — a claims adjuster’s spreadsheet that doesn’t match the underwriting system’s numbers, a plant floor tally sheet that doesn’t match what shipped.
What breaks at L1: reporting takes days instead of minutes, the same question gets a different answer depending on who you ask, and decisions get made on whoever’s spreadsheet is open, not on what’s true. This is also where one KPI showing three different numbers across three departments starts — it’s baked in from the first spreadsheet.
Moving up requires agreeing on a small number of shared definitions and getting the raw exports into one place people can actually see together — even before anything is automated.
What does Level 2 (Visible) look like?
At Level 2, a dashboard exists — usually Power BI, Tableau, or Looker — but it’s fed by hand. Someone still exports a file and uploads it, on a schedule that depends on that person being at their desk. The data is visible for the first time, but the pipeline behind it is still a person.
The tell at L2 is that the dashboard is only as current as the last manual refresh, and it breaks the moment someone renames a column or changes an export format upstream. One employee usually becomes the load-bearing wall for the whole reporting function, and everyone knows it, including that employee.
What breaks: refresh is inconsistent, so leadership starts asking “is this current?” before every meeting, and the org quietly builds a habit of not fully trusting the dashboard it just paid for. Moving up requires replacing the manual export-and-upload step with an actual pipeline — a scheduled job that moves data without a human in the loop — plus a named owner for each data source.
What does Level 3 (Automated) look like?
At Level 3, data moves on its own. Pipelines run on a schedule, dashboards refresh overnight without anyone touching them, and the copy-paste step from L1 and L2 is gone. This is a real milestone — most of the manual labor is out of the system.
The problem at L3 is quieter than at L1, and more dangerous for it: automation doesn’t fix definition drift, it hides it. Two departments can still mean different things by “revenue” or “diverted tonnage,” but now both numbers show up on a polished, automatically-refreshed dashboard that looks authoritative. People stop questioning numbers precisely because the pipeline looks solid — that’s the failure mode definition drift across departments describes in detail.
What breaks: trust erodes slowly instead of all at once, because nobody can point to “the export that broke” anymore — the automation itself gets blamed instead of the missing governance underneath it. Moving up requires a governance layer on top of the automation: written definitions, assigned data owners, and access and quality controls that don’t depend on any one person noticing a problem.
What does Level 4 (Governed) look like?
At Level 4, the organization has done the unglamorous work: core metrics have written definitions, someone owns each data source, quality checks catch bad records early, and access controls put the right data in front of the right people. Two departments pulling “active customers” get the same number, because there’s one definition instead of two guesses.
This is a genuinely strong position — most companies never get here, and 70% of companies are still running the business on spreadsheets in some form (BPM Partners/Vena), so L4 already puts an organization ahead of most of its peers. But governed data built for human-readable dashboards isn’t automatically ready for a model to consume.
What breaks: nothing, yet — which is exactly the risk. Teams at L4 sometimes assume the hard part is done and jump straight into an AI pilot on data that’s clean for a person to read but not structured, deep, or consistent enough for a model to train on. Moving up requires preparing the governed data specifically for machine consumption — consistent historical depth, feature-level structure, and documentation a model pipeline can use the same way a person uses a data dictionary.
What does Level 5 (AI-Ready) look like?
At Level 5, the governed platform feeds models directly — forecasting, anomaly detection, an internal copilot answering questions against real data — and the organization trusts the output enough to act on it, because the governance underneath it hasn’t changed, only the consumer has.
In waste and recycling, L5 looks like this: the same route, tonnage, and AR data that used to live in a Friday spreadsheet now feeds a governed model that forecasts route demand and flags accounts likely to churn, weeks before a human would have caught the pattern by hand. The data didn’t get smarter — it got trustworthy enough to hand to something that never gets tired of checking it.
What “breaks” at L5 is really what still needs tending: AI-ready is an operating discipline, not a finish line. Models drift, source systems change, and governance has to keep being enforced or the platform quietly slides back toward L3. Gartner projects that 60% of AI projects will be abandoned through 2026 due to a lack of AI-ready data — L5 is the rung that stat is describing the absence of. Before investing here, it’s worth reading what an AI readiness assessment actually checks and whether you need one first.
How do the 5 rungs compare at a glance?
The table below is a fast recap of all five rungs side by side — what each looks like day to day and the one thing standing between it and the next rung up. It’s meant to help you place your organization roughly, not precisely; a real placement takes an actual assessment, not a table.
| Level | What it looks like day to day | What’s missing to advance |
|---|---|---|
| L1 — Manual | Hand-built spreadsheets, exports from source systems, tribal knowledge of where the “real” numbers live | Shared definitions and a single place raw data lands |
| L2 — Visible | A dashboard exists but is fed by manual export-and-upload; one person is the pipeline | An automated pipeline and a named data owner |
| L3 — Automated | Data refreshes on a schedule with no human touch; definitions can still quietly drift underneath it | Written governance: definitions, ownership, quality checks |
| L4 — Governed | Documented definitions, assigned ownership, quality and access controls; numbers agree across departments | Data structured and deep enough for a model to consume |
| L5 — AI-Ready | Governed data feeds forecasting, anomaly detection, or a copilot directly; output is trusted enough to act on | Ongoing monitoring so governance doesn’t erode as models and systems change |
This table describes the general framework, which applies whatever industry you’re in. The linked self-assessment is imPROVE’s waste-and-recycling-specific version; if you’re outside that industry, the five rungs above still apply, just against your own source systems.
How do you score your own organization’s data maturity?
You can place your organization on the ladder honestly by answering five questions about how data actually behaves in your business today, not how you’d like it to behave. Answer them in order — most organizations find their honest answers cluster around one rung more than the others.
- Speed and source: When someone asks for last month’s numbers, how long does it take, and does the answer come from one place or several different spreadsheets?
- Agreement: Have two departments ever reported different numbers for the same metric — revenue, tonnage, active accounts — without being able to explain why?
- Bus factor: If the one person who builds your reports left tomorrow, would your dashboards keep updating on their own?
- Ownership: Can you name who owns each core data definition, and is that definition written down anywhere besides that person’s memory?
- AI-readiness: Could you hand your data to a forecasting or AI model today and trust what came back, or would someone need to clean it up first?
Rough guide to the pattern: mostly “no” and “I’m not sure” answers put you at L1 or L2. A working dashboard with occasional disagreements and a bus-factor problem points to L2 or L3. Written definitions and consistent answers across departments put you at L4. Confidently handing data to a model today, with monitoring in place, is L5. For a scored, numbers-based version of this same exercise rather than a directional read, the Data Foundation ROI Estimator puts a dollar figure on what climbing a rung is worth to your specific operation — most engagements we run see manual reporting effort drop 40-70% after automating the layer between L2 and L4.
Why does data maturity matter before you invest in AI?
Data maturity matters before an AI investment because AI amplifies whatever is underneath it — a governed L4 or L5 data foundation gets a genuinely useful model, while an L1 or L2 foundation gets a fast, confident, wrong answer. The model doesn’t know your spreadsheet was wrong; it just learns the wrong number faster than a person would have.
This is not a hypothetical risk. Gartner projects that 60% of AI projects will be abandoned through 2026 specifically because the data underneath them isn’t AI-ready — not because the model was bad, but because nobody checked the foundation first. Employees also lose over 20 hours a month per person working around spreadsheet limitations (ThoughtSpot), time that an AI initiative built on the same spreadsheets won’t get back, it’ll just automate the workaround.
If you’re weighing an AI pilot right now, the honest first move is usually not the pilot — it’s figuring out which rung you’re actually standing on. What is an AI readiness assessment, and do you need one first? walks through that check in more depth, and it’s worth doing before, not after, you pick a vendor.