Two weeks, six seats, one page. The runbook the framework post summarizes.
By Rizwan Yousuf, VP of Data & AI, Blue Orange Digital
The test I run on a self-scored AI readiness number is one question: did anyone open the last two incident post-mortems before the score was written down? When the answer is no, the pattern I've seen on our engagements this year is the same each time. Infrastructure comes in at least one tier high and data ownership comes in at least one tier low. That's one practitioner's observation across a handful of companies, not a study, and it's the reason this post exists.
The request usually lands as a single line in the 100-day plan or a QBR pre-read: "AI readiness score by end of quarter." The sponsor wants a number to put next to the other companies in the fund. The team that owns the pipelines wants its own number first, produced honestly, before an outside party produces one for them. This is the runbook for producing it. It runs on the AI readiness assessment framework our hub post lays out, so it doesn't re-explain the five dimensions. It shows you how to score them.
why the data team runs it first
The incentive that changed is the sponsor's. In our engagements a readiness score used to be a diligence artifact, asked for once, at the deal. Now it's asked for at the 100-day plan and again at every QBR, because the value creation plan has an AI line in it and the operating partner has to report against that line. That moves the question from the deal team's desk to yours. If the first score a sponsor sees comes from an outside assessor, the data team spends a quarter arguing with a number. If the first score is yours, backed by evidence you pulled, the outside assessment becomes a calibration step instead of a verdict. So the question data leads bring us is how to run an internal AI readiness assessment that holds up when the sponsor's assessor arrives, or an internal AI readiness audit if that's the word in the 100-day plan. We tell every one of them the same thing.
Score yourself before anyone scores you.
scope it to two weeks and six seats
Two weeks, one owner, one room. Longer than that and the assessment becomes a project with its own status meetings. The owner is the head of data or the lead data engineer, not the CIO and not a project manager. The owner writes the score.
The room is six people, and the seats matter more than the names.
- •The data owner: whoever answers when the revenue number in the board deck is wrong.
- •The platform lead: whoever owns the warehouse, the orchestrator, and the cloud bill.
- •One business owner of the first use case: the person whose team would act on the model's output, so the assessment has a real workflow to test against.
- •Governance or compliance: whoever signs off on access and retention today, even if that's the general counsel one afternoon a week.
- •One engineer who would build the first use case: the person who knows which tables are trusted and which are trusted on paper.
- •Finance: whoever owns the P&L line the first use case is meant to move. The score won't say whether that use case is worth doing, and the person who'll be asked that question should hear the evidence firsthand.
Here's what goes wrong when IT scores alone, and we've watched it happen more than once. Infrastructure gets over-scored because the people in the room built it and it works for them. Data and ownership get under-scored, or skipped, because nobody in the room is on the hook for whether the customer table agrees with the billing system. A readiness assessment run by the platform team without the data owner present produces a score for the platform, not for the company.
pull the evidence before anyone scores
The rule we hold to: score what operates today, not what's planned, funded, or in a pilot. Roadmap items go in a separate column and don't count. Week one is evidence, and the evidence list is the artifact of this section. Assign each item to a seat and give it three working days.
- •Data quality reports for the tables the business runs on, in whatever form they exist today. "We don't have any" is itself evidence and goes on the list.
- •Catalog coverage: how many production tables have a named owner and a description in whatever catalog exists (a catalog is the inventory of your datasets, with owners and definitions attached), as a count and a percentage.
- •Lineage for the three revenue-critical tables, traced from source system to board report. Lineage is the record of where a number came from and every transformation it passed through. If a person has to explain the trace, write that down.
- •The pipeline inventory with an on-call owner per pipeline. A pipeline with no named on-call is a Foundational signal however well it's built.
- •The access and permission map: who can read customer data, who can write to production, and the date it was last reviewed.
- •The model and vendor inventory: every model, agent, or AI vendor integration running in production, with its owner and what it costs a month. This list is longer than anyone in the room expects.
- •The last two incident post-mortems, or the two most recent outages if no post-mortems were written.
A week is enough because you're not fixing anything. You're photographing it. The gap between the photograph and the way the team describes itself is most of what the score will tell you.
score each dimension with a 1 to 4 rubric
Score each dimension from 1 to 4 on four named tiers: 1 Foundational, 2 Developing, 3 Scaling, 4 Leading. Place yourself at the highest tier where every tell is true today, with an item from the evidence list behind it. One person proposes, the room challenges, the owner records. Where you can't agree inside ten minutes, take the lower score and write down why. Keep the hub's weights, 25 percent each for data foundation and operational analytics, 20 each for production readiness and modernization, 10 for org and talent, so your composite is comparable to the benchmark later.
One term runs through all five, and it's worth naming before the rubrics. The data contract is who owns, funds, and is accountable for the data a company's AI runs on; it's the decision that is not reversible when the model decision is. You can swap a model vendor in a quarter. You can't swap who's accountable for the customer table without a reorg. Every dimension below has a data contract question hiding in it.
data foundation
Foundational: the definition of "revenue" or "active customer" differs across systems and the room can't agree which one is right. Developing: definitions are documented, and a named person reconciles them by hand each month. Scaling: the three revenue-critical tables have lineage, tests, and a named owner, and the tests fail loudly when they fail. Leading: quality thresholds are monitored on every production table and the owner is paged before an analyst notices. The data contract question here is the simplest one: who is paged?
operational analytics
Foundational: last quarter's numbers get challenged in the meeting and the challenge takes days to resolve. Developing: the standard dashboards are trusted, and anything new is a ticket to the data team. Scaling: business users answer new questions themselves, and an alert exists on at least two operational metrics that triggers an action when it fires. Leading: the business can name decisions it made differently last quarter because of a data signal, as a tracked practice rather than an anecdote. The tell in this dimension is latency to a trusted answer, not the number of dashboards.
AI and ML production readiness
Foundational: models live in notebooks and one person refreshes them. Developing: one model runs in production, deployed by hand, and a human verifies every output. Scaling: models deploy through a repeatable pipeline with drift monitoring (drift is when the data the model sees in production stops resembling the data it was trained on) and someone is on call for them. Leading: model outputs trigger an action without a human in the loop for at least one workflow, with an audit trail behind it. The question we ask first, every time: who gets paged at 2 a.m. when the model starts returning bad output? "Nobody" is a Foundational answer whatever the demo looked like.
technology modernization
Foundational: every pipeline is hand-wired and diagnosis needs the person who built it. Developing: core pipelines are orchestrated and alert on failure, but transformation logic lives in three places. Scaling: a second team could deploy a model the way the first one did, from the runbook, without a call. Leading: schema drift is caught before it breaks anything downstream, and you know what each pipeline costs to run. This is the dimension IT over-scores. Ask the engineer seat, not the platform seat, whether Scaling is true.
org and talent readiness
Foundational: data scientists with no data engineers, or the reverse. Developing: the skills exist but sit in one person, and that person is the on-call for everything. Scaling: a named bridge role exists between the model and the workflow, whether you call it a forward-deployed engineer or something else, and business owners can read a confidence score without help. Leading: the first use case's business owner can explain what the model did last week without calling engineering. Score this one last. It's the fastest gap to close and the easiest to over-score.
calibrate against an external benchmark
Self-scores drift high. Ours would too. Every internal scorecard I've reviewed has landed at or above the assessed score in at least one dimension, and I haven't seen one land below across the board. That's one practitioner's observation, not a study. The mechanism is easy to see, though: the people scoring know the workarounds, and a workaround doesn't show up in an evidence list unless someone writes "a person does this by hand" next to the item.
The calibration step is to run the same inputs through a benchmark you didn't build. For us that's the Blueprint assessment, run on the same evidence list inside the same two weeks. It compares your dimension scores against the companies Blue Orange has assessed and returns the gap per dimension. Its limit is the population: the benchmark is Blue Orange's assessed companies, PE-backed and mid-market, not the market at large, and we don't publish a count for that population because we haven't reconciled the completions data to a figure we'd stand behind. Treat the comparison as a second opinion from a specific population, not a percentile of the economy.
Where the two scores disagree by a full tier, that dimension is where the interview happens. Don't split the difference. Go back to the evidence and find which of the two scores it supports.
The wider context, for the sponsor who reads this page too: McKinsey's State of AI survey, published on 2026-08-25, had 1,719 respondents. In it, 37 percent attributed at least some EBIT impact to AI, and only 6 percent cleared McKinsey's high-performer bar, meaning they attribute 5 percent or more of EBIT to AI and call the impact significant. About 80 percent of respondents who use AI in their own work reported individual productivity gains. Those are self-reported answers from respondents at all levels, not filed figures, and the 80 percent counts AI users only, not every respondent. The distance between a company that sees some impact and one that clears the high bar is the distance this assessment measures before the sponsor's diligence does.
the one-page output
Score per dimension, the two largest gaps, the first three actions with an owner and a date. That's the whole page, and it fits on one. The values below are an illustration, not an engagement.
textReadiness scorecard · <Company> · Q4 2026 · Owner: Head of Data Dimension Weight Self Benchmark Gap Evidence Data foundation 25% 2 2 0 lineage on 3 tables, 1 traced by hand Operational analytics 25% 3 2 +1 2 alerts live; adoption 14 of 60 users AI/ML production readiness 20% 2 1 +1 1 model in prod, hand-deployed, no on-call Technology modernization 20% 3 3 0 dbt + Airflow; 4 flat-file feeds unorchestrated Org and talent readiness 10% 2 2 0 1 engineer covers all on-call Two largest gaps 1. Production readiness: no on-call, no drift monitor for the one live model. 2. Operational analytics: alerts exist but 46 of 60 licensed users have never logged in. First three actions Owner Date 1. Name an on-call and add drift alerts to the live model Platform lead 2026-11-14 2. Assign owners to the 4 unorchestrated feeds Data owner 2026-11-21 3. Write the data contract for the customer table Head of Data 2026-12-05
The evidence column is the part the sponsor's assessor will ask for on day one. Fill it in and the calibration conversation is short.
what the score doesn't tell you
It doesn't tell you whether the first use case is worth doing. That's a value-creation question: does the workflow the model would touch move a line the sponsor underwrites, and by how much? A Scaling company with a use case that saves eleven hours a month is less ready, in the way that matters to the operating partner, than a Developing company with a use case sitting in claims processing. The score is a statement about the foundation. It's silent on what you build on it, and it's a snapshot, so it says nothing about how fast you're moving. Run it again in two quarters and the delta is the number worth talking about.
the evidence list is the deliverable
A sponsor can get a score from anyone. The evidence list with a name next to every item is the thing only your team can produce, and it's the document the outside assessor asks for on day one regardless. Build that first and the score follows.
If your sponsor has asked for a score and your evidence list has more blanks than entries, that's the conversation to have: run the Blueprint assessment.
FAQ
Who should own an internal AI readiness assessment?
The head of data or the lead data engineer. Not the CIO, because infrastructure gets over-scored, and not a project manager, because the owner has to be able to read the evidence. The owner records the score, and the room challenges it.
What does an AI readiness scoring rubric look like?
Four tiers per dimension, Foundational through Leading, each with one concrete tell you can check against an evidence item. Place yourself at the highest tier where every tell is true today, not planned. Take the lower score when the room can't agree in ten minutes.
How often should the data team re-run the assessment?
Twice a year, and before every QBR where the sponsor has asked for the score. Use the same evidence list each time so the delta per dimension is real, and keep the roadmap column separate from the operating-today column.
