Building an AI implementation plan that actually works

An AI implementation plan is a working document that answers six questions at once. Which problems AI will solve, who owns each initiative, what data is needed, what guardrails apply, how success is measured, and what the sequence is. Miss any of the six and you have a wish list. The evidence for taking this seriously is blunt: RAND found more than 80% of AI projects fail, twice the failure rate of IT projects that do not involve AI (RAND, 2024).

Key takeaways

  • More than 80% of AI projects fail, twice the rate of non-AI IT projects, from interviews with 65 experienced data scientists and engineers (RAND, 2024).
  • 88% of organisations report regular AI use in at least one function, but only 39% attribute any EBIT impact to it (McKinsey, 2025).
  • Gartner predicted at least 30% of generative AI projects would be abandoned after proof of concept, citing poor data quality, weak risk controls, escalating costs and unclear business value (Gartner, 2024).
  • One in five breached organisations reported a breach involving shadow AI, adding an average US$670,000 to the cost (IBM Cost of a Data Breach, 2025).
  • Only 21% of organisations have a mature governance model for agentic AI, and only 25% have moved more than 40% of pilots into production (Deloitte, 2026).
  • Australia's Voluntary AI Safety Standard was superseded in October 2025 by the Guidance for AI Adoption, which sets out six essential practices.

What does an AI implementation plan actually contain?

An AI implementation plan contains six components, and a plan missing any one of them will stall. It is an operating document that ties business goals, data, governance and delivery together, not a slide deck and not a downloaded template.

The six are a ranked use-case inventory, a data readiness assessment, and a scoped pilot brief with acceptance criteria set in advance. Then a governance package, a production monitoring plan, and a measurement model with a baseline locked before anything is built.

Two of those are the ones teams skip. Acceptance criteria written after results are visible are not criteria, and a governance package written after launch is a document rather than a control.

DeliverableWhat it answersWho signs it off
Use-case inventoryWhere AI saves real time or money in this specific business, rankedFunction leads plus the executive sponsor
Data readiness assessmentWhat data exists, what is missing, what must be fixed firstData or systems owner
Pilot briefOne use case, one user group, a fixed window, a pass mark set in advanceNamed initiative owner
Governance packageAccountability, risk review, escalation, records, human oversightExecutive sponsor, with legal and security input
Production monitoring planHow you find out that performance has degraded, before a customer doesNamed model owner
Measurement modelThe baseline, the KPI, and how value converts to moneyFinance plus the initiative owner

Why do most AI projects fail?

Most AI projects fail for organisational reasons rather than technical ones. RAND interviewed 65 experienced data scientists and engineers and found more than 80% of AI projects fail, twice the rate of IT projects that do not involve AI. The leading root cause was leadership misunderstanding the problem the project was meant to solve (RAND, 2024).

Data is the second cause, and RAND puts it memorably: 80% of AI is the dirty work of data engineering. Gartner's list matches, naming poor data quality, inadequate risk controls, escalating costs and unclear business value as the reasons at least 30% of generative AI projects would be abandoned after proof of concept (Gartner, 2024).

There is a competing explanation worth knowing about. MIT NANDA's The GenAI Divide (2025) argues the core barrier to scaling is not infrastructure, regulation or talent but learning. Tools fail, it says, because they do not learn, adapt or integrate with how the work is actually done. That report is a working paper rather than peer-reviewed work, and its headline 95% figure has been widely contested, so attribute it precisely rather than treating it as settled.

The gap between activity and outcome is the clearest number in the field. 88% of organisations report regular AI use in at least one business function, but only 39% attribute any level of EBIT impact to it. Nearly two-thirds have not begun scaling AI across the enterprise (McKinsey, 2025, from 1,993 respondents across 105 countries).

How do you audit your workflows before building anything?

You audit by watching the work, not by reading a process document. The audit maps current processes, finds where AI genuinely saves time, and finds where it would add complexity for no gain. That means interviewing the people doing the task, observing the actual steps, and recording the friction points that cost time or create rework.

The output is a prioritisation matrix on two axes. Impact means time saved, cost reduced or quality improved. Feasibility means data availability, integration complexity and organisational readiness.

A strong first use case has four traits: it is high-volume and repetitive, the outcome is measurable, the data already exists, and one named person can be accountable for the result. Avoid anything where the output is highly subjective, the data is thin, or an error carries serious consequences downstream. Those are second-wave initiatives.

Quick wins matter for a reason that is not about the saving. A team that sees AI produce a measurable result inside a short pilot is far more willing to fund the twelve-month roadmap.

What does data readiness mean in practice?

Data readiness means your data can support the system in production, not just in a demo. Five criteria decide it, and a gap analysis has to happen before pilot design, not during it.

  • Completeness: critical fields are present, and the gaps that remain are understood and handled deliberately.
  • Accuracy: values match a trusted source rather than simply being populated.
  • Consistency: the same entity does not carry conflicting values in different systems.
  • Representativeness: the dataset covers the real range of cases the system will meet, not just the straightforward ones.
  • Freshness: the pipeline delivers current data at the refresh rate the use case needs.

Access control belongs here too, not in a later phase. Australia's Guidance for AI Adoption includes testing and monitoring, and sharing essential information, among its six essential practices, and both depend on knowing where your data came from and who can touch it.

Data drift is the failure that arrives quietly. When live data starts to diverge from what a model was built on, performance degrades without anything breaking, which is exactly why the monitoring plan is written before launch rather than after it.

How do you design a pilot that proves something?

A pilot proves something when it tests one use case, with a defined user group, over a fixed window, against a pass mark set before any results are visible. The pilot is not the deployment. It is the test that decides whether you deploy.

Scope creep is the most common way a pilot stops being a test. Adding use cases mid-run, changing the user group, or redefining success once the numbers are in all invalidate the result, and each one is tempting for the same reason: it makes the pilot look better.

Lock the baseline before you start. Without a pre-AI measurement you are comparing an outcome to a feeling.

KPI categoryExamplesWhat it tells you
Business impactTime saved per task, cost per outputWhether the case for the investment holds
Process qualityAccuracy rate, rework rate, error frequencyWhether the speed is real or is being paid for in corrections
System reliabilityLatency, failure rate, uptimeWhether it will hold at production volume
AdoptionUsage rate, reviewer confidence, escalation volumeWhether anyone is actually using it

Pick one primary KPI tied directly to the business problem, and track the rest as supporting evidence. Deloitte's 2026 research found only 25% of organisations have moved more than 40% of their AI pilots into production, which suggests most pilots are not producing a decision either way.

The close is a formal moment with three possible outcomes: scale, revise and retest, or stop. Stopping a pilot that missed its criteria is not a failure. It is the reason you ran a controlled test instead of committing production budget on a hunch.

What belongs in the governance layer before you scale?

The governance layer needs named owners, documented decisions, and a record someone else could audit. IBM's guidance on AI governance implementation names four things. Model owners who own the full lifecycle from build to production. Model factsheets documenting intent and datasets. Continuous monitoring and audit trails to catch drift and policy violations. And clear reporting lines that define escalation authority (IBM, 2026).

For Australian organisations, the reference point changed recently. The Voluntary AI Safety Standard's ten guardrails were superseded on 21 October 2025 by the Guidance for AI Adoption. It condenses them into six essential practices: decide who is accountable, understand impacts and plan accordingly, measure and manage risks, share essential information, test and monitor, and maintain human control. If your governance document still cites the ten guardrails, it is out of date.

Shadow AI is the exposure most plans underweight. One in five breached organisations reported a breach involving shadow AI, adding an average US$670,000 to the cost. Of breached organisations, 63% either had no AI governance policy or were still developing one (IBM Cost of a Data Breach, 2025). The 2026 edition found the share of security incidents involving shadow AI more than doubled year on year, to 43%.

For directors, that is a liability question rather than an efficiency question. A sanctioned framework that brings informal AI use inside the tent is cheaper than the alternative.

How do you calculate ROI on an AI programme?

You calculate ROI by locking a baseline, measuring the change, converting it to money at a stated rate, and subtracting the full programme cost. The sequence matters more than the formula.

Lock the pre-AI baseline for your primary KPI before the pilot starts. Measure the change after a defined production period. Convert the improvement using a unit cost rate: improvement per task, multiplied by monthly volume, multiplied by the fully loaded labour cost per unit. Subtract total programme cost, including build, run, oversight and change management. Net benefit divided by total cost gives ROI.

A worked example, with the assumptions stated so you can substitute your own.

StepInputResult
Time saved per asset12 minutes-
Monthly volume500 assets6,000 minutes saved
Fully loaded rateUS$1.20 per minute (US$72 per hour)US$7,200 gross monthly value
Monthly run costUS$2,000US$5,200 net monthly benefit
Build costUS$30,000Payback in 5.8 months

Three caveats belong with any calculation like this one. A US$1.20 per minute rate implies a fully loaded cost of about US$72 an hour, which is high for content review, so check it against your own. Saved minutes only become cash if headcount or contracted hours actually change. And simple payback ignores the cost of the pilot that came before the build.

Payback period is the second number leadership responds to, and it belongs beside ROI in the reporting. A quarterly spreadsheet will not catch performance drift, so the tracking needs to be continuous.

What does an AI implementation plan cost and how long does it take?

There is no credible published benchmark for mid-market AI pilot or deployment costs, and the ranges circulating online do not trace back to any primary source. The figures widely quoted, between US$20,000 and US$500,000 for a pilot and US$1 million to US$5 million for enterprise rollout, come from vendor and agency blogs with no disclosed sample or method.

The one authoritative figure we could verify sits at a different scale entirely. Gartner puts generative AI deployments that transform the business model at US$5 million to US$20 million, and is explicit that this excludes proofs of concept (Gartner, 2024).

What can be said honestly is about shape rather than price. A pilot is short enough that the market does not move under it and long enough to gather a real sample. Production deployment adds integration, monitoring and change management. Scale adds the governance and enablement work across functions. Price each stage against your own labour rates and vendor quotes, and treat any article quoting a confident dollar band with suspicion.

Frequently asked questions

What is the difference between an AI strategy and an AI implementation plan?

A strategy states which business outcomes AI should move. An implementation plan states who owns each initiative, what data it needs, what guardrails apply, how success is measured and in what sequence work happens. The strategy is the intent. The plan is what makes it auditable and fundable.

How long should an AI pilot run?

Long enough to gather a meaningful sample of real work and short enough that conditions do not change underneath it. The length matters less than fixing it in advance, alongside the user group and the pass mark. Extending a pilot because results are disappointing invalidates the test.

What is shadow AI and why does it belong in an implementation plan?

Shadow AI is employees using personal or unapproved AI accounts for work. It belongs in the plan because it creates data exposure with no controls. IBM found one in five breached organisations reported a breach involving shadow AI, adding an average US$670,000 to the cost.

Which Australian AI governance guidance should we follow?

The Guidance for AI Adoption, published by the Department of Industry, Science and Resources and the National AI Centre in October 2025. It superseded the Voluntary AI Safety Standard's ten guardrails with six essential practices. Many organisations also map to NIST's AI Risk Management Framework or ISO/IEC 42001.

Who should own an AI initiative internally?

One named person accountable for the outcome, not a committee and not the vendor. IBM's governance guidance calls for model owners who own the full lifecycle from build to production, with clear reporting lines defining escalation authority. Shared ownership reliably becomes no ownership at the first difficult decision.

What should we do if the pilot fails its acceptance criteria?

Stop it or revise and retest, and record why. A pilot that misses its pass mark has done its job by preventing a larger commitment. The failure mode to avoid is redefining success after the results are visible, which converts a controlled test into an expensive opinion.

Do we need a data warehouse before we start?

No, but you do need to know what data exists, where it lives, who can access it and how current it is. Data readiness is about completeness, accuracy, consistency, representativeness and freshness for the specific use case, not about a particular platform. Scope the assessment to the pilot.

If you'd like an AI implementation plan your leadership can sign off and your team can actually run, book a call.

A note on legal matters: this article is general in nature and is not legal advice. Involve your own legal, privacy and security teams before deploying AI on customer or employee data.

Category
Insights
AI strategy
Written by
Harper
Editor
blogs and articles

Latest insights and trends

Insights

How to build a winning AI marketing strategy in 2026

An AI marketing strategy is a sequence, not a stack. Outcomes, audit, tools, pilot, governance, with the 2026 evidence behind each step.
Insights

AI content vs human content: the gap is editing, not AI

AI content vs human content: what the 2026 ranking, trust and cost data shows, and why the editing layer decides the outcome. Book a call.
Insights

The AI capability layer: work through AI, don't learn every tool

An AI capability layer lets a connected model operate your software for you, but only context makes the work worth shipping.
Line illustration of document cards standing behind a low barrier while dashed speech bubbles dissolve into dots.
Insights

Brand voice for AI: you build it, you don't just describe it

To make AI sound like your brand, feed it real gold-standard passages, not adjectives, because patterns beat labels every time.
Insights

Your business' best AI process is trapped in a personal login

AI as business IP turns your best AI processes into an owned asset that compounds, rather than a login you rent.
Let's talk

Ready to use AI well?

Tell us what's slowing your business down. We'll show you the shortest path through it.