
An AI implementation plan is a working document that answers six questions at once. Which problems AI will solve, who owns each initiative, what data is needed, what guardrails apply, how success is measured, and what the sequence is. Miss any of the six and you have a wish list. The evidence for taking this seriously is blunt: RAND found more than 80% of AI projects fail, twice the failure rate of IT projects that do not involve AI (RAND, 2024).
An AI implementation plan contains six components, and a plan missing any one of them will stall. It is an operating document that ties business goals, data, governance and delivery together, not a slide deck and not a downloaded template.
The six are a ranked use-case inventory, a data readiness assessment, and a scoped pilot brief with acceptance criteria set in advance. Then a governance package, a production monitoring plan, and a measurement model with a baseline locked before anything is built.
Two of those are the ones teams skip. Acceptance criteria written after results are visible are not criteria, and a governance package written after launch is a document rather than a control.
| Deliverable | What it answers | Who signs it off |
|---|---|---|
| Use-case inventory | Where AI saves real time or money in this specific business, ranked | Function leads plus the executive sponsor |
| Data readiness assessment | What data exists, what is missing, what must be fixed first | Data or systems owner |
| Pilot brief | One use case, one user group, a fixed window, a pass mark set in advance | Named initiative owner |
| Governance package | Accountability, risk review, escalation, records, human oversight | Executive sponsor, with legal and security input |
| Production monitoring plan | How you find out that performance has degraded, before a customer does | Named model owner |
| Measurement model | The baseline, the KPI, and how value converts to money | Finance plus the initiative owner |
Most AI projects fail for organisational reasons rather than technical ones. RAND interviewed 65 experienced data scientists and engineers and found more than 80% of AI projects fail, twice the rate of IT projects that do not involve AI. The leading root cause was leadership misunderstanding the problem the project was meant to solve (RAND, 2024).
Data is the second cause, and RAND puts it memorably: 80% of AI is the dirty work of data engineering. Gartner's list matches, naming poor data quality, inadequate risk controls, escalating costs and unclear business value as the reasons at least 30% of generative AI projects would be abandoned after proof of concept (Gartner, 2024).
There is a competing explanation worth knowing about. MIT NANDA's The GenAI Divide (2025) argues the core barrier to scaling is not infrastructure, regulation or talent but learning. Tools fail, it says, because they do not learn, adapt or integrate with how the work is actually done. That report is a working paper rather than peer-reviewed work, and its headline 95% figure has been widely contested, so attribute it precisely rather than treating it as settled.
The gap between activity and outcome is the clearest number in the field. 88% of organisations report regular AI use in at least one business function, but only 39% attribute any level of EBIT impact to it. Nearly two-thirds have not begun scaling AI across the enterprise (McKinsey, 2025, from 1,993 respondents across 105 countries).
You audit by watching the work, not by reading a process document. The audit maps current processes, finds where AI genuinely saves time, and finds where it would add complexity for no gain. That means interviewing the people doing the task, observing the actual steps, and recording the friction points that cost time or create rework.
The output is a prioritisation matrix on two axes. Impact means time saved, cost reduced or quality improved. Feasibility means data availability, integration complexity and organisational readiness.
A strong first use case has four traits: it is high-volume and repetitive, the outcome is measurable, the data already exists, and one named person can be accountable for the result. Avoid anything where the output is highly subjective, the data is thin, or an error carries serious consequences downstream. Those are second-wave initiatives.
Quick wins matter for a reason that is not about the saving. A team that sees AI produce a measurable result inside a short pilot is far more willing to fund the twelve-month roadmap.
Data readiness means your data can support the system in production, not just in a demo. Five criteria decide it, and a gap analysis has to happen before pilot design, not during it.
Access control belongs here too, not in a later phase. Australia's Guidance for AI Adoption includes testing and monitoring, and sharing essential information, among its six essential practices, and both depend on knowing where your data came from and who can touch it.
Data drift is the failure that arrives quietly. When live data starts to diverge from what a model was built on, performance degrades without anything breaking, which is exactly why the monitoring plan is written before launch rather than after it.
A pilot proves something when it tests one use case, with a defined user group, over a fixed window, against a pass mark set before any results are visible. The pilot is not the deployment. It is the test that decides whether you deploy.
Scope creep is the most common way a pilot stops being a test. Adding use cases mid-run, changing the user group, or redefining success once the numbers are in all invalidate the result, and each one is tempting for the same reason: it makes the pilot look better.
Lock the baseline before you start. Without a pre-AI measurement you are comparing an outcome to a feeling.
| KPI category | Examples | What it tells you |
|---|---|---|
| Business impact | Time saved per task, cost per output | Whether the case for the investment holds |
| Process quality | Accuracy rate, rework rate, error frequency | Whether the speed is real or is being paid for in corrections |
| System reliability | Latency, failure rate, uptime | Whether it will hold at production volume |
| Adoption | Usage rate, reviewer confidence, escalation volume | Whether anyone is actually using it |
Pick one primary KPI tied directly to the business problem, and track the rest as supporting evidence. Deloitte's 2026 research found only 25% of organisations have moved more than 40% of their AI pilots into production, which suggests most pilots are not producing a decision either way.
The close is a formal moment with three possible outcomes: scale, revise and retest, or stop. Stopping a pilot that missed its criteria is not a failure. It is the reason you ran a controlled test instead of committing production budget on a hunch.
The governance layer needs named owners, documented decisions, and a record someone else could audit. IBM's guidance on AI governance implementation names four things. Model owners who own the full lifecycle from build to production. Model factsheets documenting intent and datasets. Continuous monitoring and audit trails to catch drift and policy violations. And clear reporting lines that define escalation authority (IBM, 2026).
For Australian organisations, the reference point changed recently. The Voluntary AI Safety Standard's ten guardrails were superseded on 21 October 2025 by the Guidance for AI Adoption. It condenses them into six essential practices: decide who is accountable, understand impacts and plan accordingly, measure and manage risks, share essential information, test and monitor, and maintain human control. If your governance document still cites the ten guardrails, it is out of date.
Shadow AI is the exposure most plans underweight. One in five breached organisations reported a breach involving shadow AI, adding an average US$670,000 to the cost. Of breached organisations, 63% either had no AI governance policy or were still developing one (IBM Cost of a Data Breach, 2025). The 2026 edition found the share of security incidents involving shadow AI more than doubled year on year, to 43%.
For directors, that is a liability question rather than an efficiency question. A sanctioned framework that brings informal AI use inside the tent is cheaper than the alternative.
You calculate ROI by locking a baseline, measuring the change, converting it to money at a stated rate, and subtracting the full programme cost. The sequence matters more than the formula.
Lock the pre-AI baseline for your primary KPI before the pilot starts. Measure the change after a defined production period. Convert the improvement using a unit cost rate: improvement per task, multiplied by monthly volume, multiplied by the fully loaded labour cost per unit. Subtract total programme cost, including build, run, oversight and change management. Net benefit divided by total cost gives ROI.
A worked example, with the assumptions stated so you can substitute your own.
| Step | Input | Result |
|---|---|---|
| Time saved per asset | 12 minutes | - |
| Monthly volume | 500 assets | 6,000 minutes saved |
| Fully loaded rate | US$1.20 per minute (US$72 per hour) | US$7,200 gross monthly value |
| Monthly run cost | US$2,000 | US$5,200 net monthly benefit |
| Build cost | US$30,000 | Payback in 5.8 months |
Three caveats belong with any calculation like this one. A US$1.20 per minute rate implies a fully loaded cost of about US$72 an hour, which is high for content review, so check it against your own. Saved minutes only become cash if headcount or contracted hours actually change. And simple payback ignores the cost of the pilot that came before the build.
Payback period is the second number leadership responds to, and it belongs beside ROI in the reporting. A quarterly spreadsheet will not catch performance drift, so the tracking needs to be continuous.
There is no credible published benchmark for mid-market AI pilot or deployment costs, and the ranges circulating online do not trace back to any primary source. The figures widely quoted, between US$20,000 and US$500,000 for a pilot and US$1 million to US$5 million for enterprise rollout, come from vendor and agency blogs with no disclosed sample or method.
The one authoritative figure we could verify sits at a different scale entirely. Gartner puts generative AI deployments that transform the business model at US$5 million to US$20 million, and is explicit that this excludes proofs of concept (Gartner, 2024).
What can be said honestly is about shape rather than price. A pilot is short enough that the market does not move under it and long enough to gather a real sample. Production deployment adds integration, monitoring and change management. Scale adds the governance and enablement work across functions. Price each stage against your own labour rates and vendor quotes, and treat any article quoting a confident dollar band with suspicion.
A strategy states which business outcomes AI should move. An implementation plan states who owns each initiative, what data it needs, what guardrails apply, how success is measured and in what sequence work happens. The strategy is the intent. The plan is what makes it auditable and fundable.
Long enough to gather a meaningful sample of real work and short enough that conditions do not change underneath it. The length matters less than fixing it in advance, alongside the user group and the pass mark. Extending a pilot because results are disappointing invalidates the test.
Shadow AI is employees using personal or unapproved AI accounts for work. It belongs in the plan because it creates data exposure with no controls. IBM found one in five breached organisations reported a breach involving shadow AI, adding an average US$670,000 to the cost.
The Guidance for AI Adoption, published by the Department of Industry, Science and Resources and the National AI Centre in October 2025. It superseded the Voluntary AI Safety Standard's ten guardrails with six essential practices. Many organisations also map to NIST's AI Risk Management Framework or ISO/IEC 42001.
One named person accountable for the outcome, not a committee and not the vendor. IBM's governance guidance calls for model owners who own the full lifecycle from build to production, with clear reporting lines defining escalation authority. Shared ownership reliably becomes no ownership at the first difficult decision.
Stop it or revise and retest, and record why. A pilot that misses its pass mark has done its job by preventing a larger commitment. The failure mode to avoid is redefining success after the results are visible, which converts a controlled test into an expensive opinion.
No, but you do need to know what data exists, where it lives, who can access it and how current it is. Data readiness is about completeness, accuracy, consistency, representativeness and freshness for the specific use case, not about a particular platform. Scope the assessment to the pilot.
If you'd like an AI implementation plan your leadership can sign off and your team can actually run, book a call.
A note on legal matters: this article is general in nature and is not legal advice. Involve your own legal, privacy and security teams before deploying AI on customer or employee data.
