Concept Note v1.0 · Practice Toolkit v1.0 · Published for comment
Grounded
A framework for testing whether an organisation is ready for the AI investment in front of it
Most AI investments do not fail because the technology is weak. They fail because the foundation underneath them — the data, the economics at scale, the people whose behaviour must change, and the alignment of the leadership team itself — was never tested before the capital was committed.
Grounded is an open diagnostic for running that test: ten blocks, thirty observable evidence questions, three gates that cannot be averaged away, answers collected privately by role, and an output that is a verdict with a priced register of gaps — not a score.
It is not only an argument. Everything needed to run an assessment ships with it — the engagement letter, the private intake forms, the scoring workbook that computes the verdict, the session runsheet, the board one-pager. Eight artifacts, editable, free. If you came to run an assessment rather than read about one, go straight to the practice toolkit.
01
The problem — mispriced transformation
Every major technology wave has failed enterprises in the same way. ERP programmes in the 1990s did not fail because the software couldn’t post a journal entry. Cloud migrations in the 2010s did not fail because the infrastructure was unreliable. The pattern is stable across three decades: the technology is rarely the thing that breaks. The foundation is. AI now repeats the pattern faster and more expensively — and with a hazard the earlier waves did not have. An AI system can fail silently, degrading as data drifts while every dashboard stays green.
- The data is connected but not decision-grade.
- Sensors report, dashboards render, the volume is enormous — and yet nobody can reconstruct a single moment from it. Which batch was running, which material lot was loaded, whether the stoppage at 2am was planned. A model trained on data that cannot answer those questions learns noise, and predicts noise with great confidence.
- The economics were priced at pilot scale.
- The pilot's return was real. But integration cost, change management, data preparation and per-inference cost all multiply at rollout, and no one in finance ever signed their name to the number at ten times the usage. The business case that reached the board was, in the precise sense of the word, mispriced.
- Nobody's behaviour was going to change.
- The people whose daily judgement the system is meant to inform were not consulted in its design, cannot interpret its output in terms they trust, and — after the first false positive — quietly stop looking at it. The system keeps running. Nobody uses it. The dashboards, again, stay green.
- The leadership team never agreed on what was true.
- The CFO believed the data was ready. The IT lead knew it wasn't. Both sat in the approval meeting; only one of them spoke, and it was the one furthest from the evidence.
The organisation pays to close its foundation gaps whether or not it acknowledges them. The only choice is whether that cost appears in the business case before the capital is committed, or arrives eighteen months later disguised as “the AI didn’t work here.”
02
Why existing tools don't catch it
Canvases and maturity models were built for different questions, and they share two structural flaws. They run in reflection mode: they ask people to assess themselves, which works only if you already know what you don’t know. And they run in consensus mode: they are filled in by groups, where the highest-paid person’s opinion quietly becomes everyone else’s answer.
A canvas completed in a workshop records agreement with authority, not truth.
Reflection mode
Assumes the person filling it in already knows what they don’t know. Asked “is your data decision-grade?”, a VP who has never built an ML system will honestly answer “we have lots of data” — and score themselves a four.
Diagnostic mode
Demands evidence you can point at. Instead: “can you pull, right now, thirty days of the exact data this model needs — timestamped, no gaps, no IT ticket?” The question carries the expertise, so the respondent doesn’t have to.
03
Ten blocks, three of them gates
Ten blocks, thirty evidence questions. For seven of them, weakness degrades the return and can be traded off against strength elsewhere. For three — unit economics, data, and adoption — failure does not degrade the return. It voids it. Those three are gates, and you cannot average your way past a foundation that isn’t there.
- 01
Problem & opportunity
Business
- 02
Value proposition
Business
- 03Gate
Value capture & unit economics
Market
- 04
Build / buy / orchestrate
Technical
- 05Gate
Data & context
Technical
- 06
Moat & strategy
Business
- 07Gate
Adoption & change
Market
- 08
Governance & risk
Risk
- 09
Evaluation & quality
Quality
- 10
Sustainability & ops
Quality
A gate block. Two or more No answers fail the gate and cap the verdict — no amount of excellence elsewhere can average it away.
What each block exists to test
- 01Problem & opportunity
- What decision is slow, expensive, or inconsistent — and is AI load-bearing here, or decorative?
- 02Value proposition
- What will a specific user be able to do that no non-AI alternative allows — not faster, but at all?
- 03Value capture & unit economicsGate
- How does value reach the P&L — and does the margin survive at ten times the usage?
- 04Build / buy / orchestrate
- What is owned versus rented — and what breaks when a rented component is discontinued?
- 05Data & contextGate
- Is the data decision-grade — or merely connected? Who owns its quality, by name?
- 06Moat & strategy
- What sustains advantage when every competitor can call the same API?
- 07Adoption & changeGate
- Whose behaviour must change, why would they, and what happens the first time the AI is wrong?
- 08Governance & risk
- Who is liable when it's wrong — and what is the worst plausible failure?
- 09Evaluation & quality
- How is silent degradation caught in production — and by whom?
- 10Sustainability & ops
- Who owns the system when it degrades — and what happens when the model is deprecated?
04
The Evidence Rule
“Yes” is not an answer. “Yes” plus the name of the thing is an answer.
Every question is answered Yes Partially or No — under one rule. A Yes only counts if you can name the artifact that proves it: the document, the dashboard, the named person, the date. A Yes with no named artifact is recorded as Partially. No exceptions, including for the most senior person in the room.
The rule works because it changes what kind of claim an answer is. You can be optimistic about an opinion. Optimism about a document that does not exist is simply a false statement — and people are far more reluctant to make one.
- Q:
- “Is there a named person in finance who has signed off the running cost of this system at ten times pilot scale?”
- A:
- “Yes.”
- Q:
- “Who — and when did they sign?”
- A:
- “…I’d have to check.”
Recorded: Partially. Nothing accusatory has happened — but the answer is now honest.
The intake form that enforces this rule — a mandatory artifact field against every Yes — is one of the eight artifacts in the practice toolkit. You can take it without finishing this page.
05
The method — from one sentence to a verdict
It begins with one sentence, completed by the sponsor before anyone answers anything: “We are investing — so that — can —, measured by —.” If the leadership team cannot agree that sentence, that is finding number one: a scoping failure discovered in five minutes rather than eighteen months.
Then each role answers alone, before any meeting. That independence is not a procedural nicety. It is the entire defence against the dynamics of the room. The shift supervisor who knows the downtime codes are entered from memory will say so on a private form. They will not say so across a table from the plant manager.
- 1
Statement
one sentence, agreed
- 2
Private intake
each role, alone
- 3
Fault lines
where roles differ
- 4
Verdict
grounded / not / why
- 5
Debt register
gaps, owned, dated
06
Fault lines — disagreement as diagnostic
Two organisations can return the same score and face opposite prospects. The score does not predict failure. The disagreement does.
When two roles answer the same evidence question differently, only two things can be true: either the artifact does not exist and someone believes it does, or it exists and someone who needs it does not know. Both are foundation failures. Both are invisible in an ordinary meeting.
| 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | |
|---|---|---|---|---|---|---|---|---|---|---|
| CFO | Block 1: Yes — evidence named | Block 2: Yes — evidence named | Block 3: Yes — evidence named | Block 4: Partially | Block 5: Yes — evidence named | Block 6: Yes — evidence named | Block 7: Yes — evidence named | Block 8: Partially | Block 9: Yes — evidence named | Block 10: Yes — evidence named |
| Plant manager | Block 1: Yes — evidence named | Block 2: Partially | Block 3: Yes — evidence named | Block 4: Yes — evidence named | Block 5: Yes — evidence named | Block 6: Yes — evidence named | Block 7: Partially | Block 8: Yes — evidence named | Block 9: Yes — evidence named | Block 10: Partially |
| IT / OT lead | Block 1: Yes — evidence named | Block 2: Yes — evidence named | Block 3: Partially | Block 4: Yes — evidence named | Block 5: No | Block 6: Yes — evidence named | Block 7: Yes — evidence named | Block 8: Yes — evidence named | Block 9: Partially | Block 10: Yes — evidence named |
| Shift supervisor | Block 1: Yes — evidence named | Block 2: Yes — evidence named | Block 3: Yes — evidence named | Block 4: Partially | Block 5: No | Block 6: Yes — evidence named | Block 7: Partially | Block 8: Yes — evidence named | Block 9: Yes — evidence named | Block 10: Yes — evidence named |
YesPartiallyNo
Block 5 — a critical fault line. The roles closest to the data say No. The roles who approved the budget say Yes. The average score barely moves. The prognosis moves entirely.
An honest note on that claim
“The disagreement predicts failure better than the score” is the boldest statement in the framework, and today it is a design hypothesis — not a validated statistical finding. Field assessments will test it, and I will publish what they show, including if it is wrong. Note what does not depend on it: private intake is justified on its own merits, and a critical fault line is worth surfacing even if its predictive power turns out to be weaker than I believe.
07
A verdict, not a score
Scores get rounded up on the way to the board — a 34 out of 50 becomes “we scored well on readiness” by the third retelling. A verdict cannot be rounded. It is also more honest about what the diagnostic knows: the difference between 41 and 44 is noise; the difference between a passed and a failed data gate is everything.
- Grounded
Score ≥ 48 · all gates pass · no critical split
- Conditionally Grounded
Any failed gate or critical split, or a score of 30–47
- Ungrounded
Two failed gates, or a score under 30
Every No becomes a debt item — a piece of foundation the investment is borrowing against — with a cost to close, a named owner and a deadline. Boards do not intuit “readiness gaps.” They understand debt intimately, including the part where unacknowledged debt compounds.
And grounding expires. Every verdict carries a ninety-day validity, because vendors deprecate models, accountable people rotate, and pipelines rot. A readiness assessment with no expiry date becomes actively misleading within months.
08
The toolkit — everything the method needs to run
A framework that stops at the argument leaves the hard part — the actual running of it — as an exercise for the reader, and quietly becomes something you can only buy from its author. So everything is here: eight artifacts, from the engagement letter to the board one-pager. Six Word templates and two Excel workbooks. The scoring workbook is the engine — transcribe the intake forms and the block scores, gates, fault lines, verdict, debt register and expiry date compute themselves.
- 01Engagement Letter & Statement of WorkWord
- 02Investment Statement WorksheetWord
- 03Private Intake QuestionnaireWord
- 04Scoring & Aggregation WorkbookExcel
- 05Facilitation Guide & Session RunsheetWord
- 06Assessment Report & Board One-PagerWord
- 07Grounded Scan · ProductExcel
- 08Product Edition — Question MapWord
Each artifact, what it carries, and how to run it. Free and editable.
What is not on this page: the worked example — a €600K predictive maintenance investment that scores 51 out of 60 and still fails its data gate — and the section listing every objection I currently consider serious, including the ones I cannot yet answer.
Both are in the note.
Version 1.0 · Published for comment
Take both
The argument and the kit that runs it. Neither is gated — a framework that claims to be open and then puts a form in front of its own templates is not open.
The concept note
Fourteen pages. Every question, every scoring rule, the worked example, and a section devoted entirely to where the framework is most exposed to challenge.
Concept note (PDF)The practice toolkit
All eight artifacts — six Word templates, two Excel workbooks. The scoring engine computes the gates, the fault lines, the verdict and the expiry on its own.
All eight files (ZIP)No email required. No form in the way. Individual files are linked against each artifact above.
Optional
Leave an email only if you want to hear what the field data says when it comes in, or if you intend to tell me where the method creaked. It reaches me directly, at a personal address. No list, no sequence, no sales call — and if you would rather just take the PDF, take it. The framework is open either way.
Know what you’re building on before you build.