An AI pilot project is supposed to answer one question: is this worth paying for at full scale? Too many pilots never answer it. They run for months, produce an impressive demo and then stall, because nobody agreed at the start what "worth it" would look like.
The fix is a tight scope written before any work begins. Below is a one-page scope you can fill in with your team, line by line, plus the warning signs that a pilot has grown too big to prove anything.
Why most AI pilot projects fail to prove anything
The numbers are sobering. Fortune's report on MIT's NANDA research found that only about 5% of generative AI pilots were delivering rapid revenue gains; the rest stalled with little measurable effect on profit and loss. The researchers pointed to poor fit with daily work rather than weak AI models.
Analysts saw the same pattern earlier. Gartner predicted in 2024 that at least 30% of generative AI projects would be abandoned after the proof of concept, citing poor data quality, weak risk controls, rising costs and unclear business value. A year later it predicted that over 40% of agentic AI projects would be canceled by the end of 2027 for similar reasons.
Notice what is missing from those reasons: the technology not working. Pilots usually fail on scope, data and measurement. Those are all things you can settle on paper first.
The one-page AI pilot scope
Each of the eight lines below should fit in a sentence or two. If a line takes a page to explain, the pilot is too big.
1. The job
Name one routine task, done by one team, with a clear start and end. "Matching supplier invoices to purchase orders in accounts payable" is a job. "Using AI in finance" is not.
Write down: the task, the team and how often it happens each week.
2. Today's baseline
You cannot show improvement without a starting point. Measure the task as it runs now, for at least two weeks, before the pilot starts.
Write down: hours spent per week, items handled, error or rework rate and how long each item waits.
3. The success measure
Pick one or two numbers that would make you say yes to a full rollout, and set the target before you see any results. Setting it afterward invites wishful thinking.
Write down: for example, "staff time on this task falls by a third with no rise in errors."
4. The data
List exactly which systems and documents the AI will read, and confirm you can get a realistic sample from each. Most delays in a pilot come from data access, not from building.
Write down: each source, who owns it and whether it can stay inside your own systems.
5. The human checkpoint
Decide where a person approves the AI's work before anything is sent, changed or paid. In a pilot, that checkpoint should cover everything that leaves the building. It also gives you a clean record of where the AI was right and where it was corrected.
Write down: who approves, what they approve and what happens when they reject a draft.
6. The risks and limits
Agree what the pilot is not allowed to touch: certain customers, payment instructions, personal data, regulated filings. The NIST AI Risk Management Framework is a useful free reference if your board or regulator will ask how risk was handled; its four functions (govern, map, measure, manage) translate well into a short pilot risk note.
Write down: what is out of bounds and who signs off on the risk note.
7. The time box and cost
Set a fixed length and a fixed budget. Eight to twelve weeks is long enough to see real cases and short enough to keep attention. Open-ended pilots drift.
Write down: start date, end date, total cost and who pays for what.
8. The decision at the end
Agree now what happens when the pilot ends, for each possible result. This is the line most teams skip, and it is why pilots linger.
Write down: who decides, on what date, and the next step for go, adjust and stop.
What a filled-in scope looks like
Here is an illustrative example. The firm and figures are made up to show the format.
Job: Check incoming supplier invoices against purchase orders and goods received, for the accounts payable team (about 600 invoices a week).
Baseline: Two clerks, about 30 hours a week combined; roughly one invoice in eight needs a query to the supplier.
Success measure: Clerk time on matching falls by at least a third, with no increase in wrongly approved invoices.
Data: ERP purchase orders and receipts; invoice PDFs from the AP inbox. Both stay in the company's cloud account.
Human checkpoint: The AI drafts a match and a supplier query; a clerk approves every one before it is posted or sent.
Limits: No payment runs, no changes to supplier bank details.
Time box: Ten weeks, fixed price.
Decision: Finance director decides in week eleven. Go means rollout to all suppliers; adjust means four more weeks on the failing invoice types; stop means the work ends with the findings written up.
Warning signs your pilot is scoped wrong
More than one team is involved. Each extra team adds meetings, data sources and opinions about success.
The success measure is "people like it". Satisfaction matters, but it won't justify a budget on its own.
Nobody has measured the current process. Without a baseline, every result is an anecdote.
The data request is still open. If IT has not confirmed access by week one, the timeline is already slipping.
The pilot runs on clean sample data only. Real documents are messy. Test on them, or the result won't hold.
There is no end date. With no end date, nobody ever has to decide, so the pilot just keeps going.
The vendor owns the result. Code, prompts and data should be yours, so the pilot's value doesn't leave with the vendor.
How to choose the right first task
The research above gives a useful hint. The MIT work, as Fortune reported it, found the largest returns in back-office automation, even though more than half of generative AI budgets were going to sales and marketing tools. For a first pilot, look for work that is:
High in volume and repetitive
Based on documents or data you already hold
Easy to check, so a person can tell quickly whether the AI got it right
Painful enough that the team wants it fixed
Invoice matching, contract review against a checklist, routine customer questions and month-end reconciliations often fit. A task that needs judgment on every item, or where a single mistake is very costly, is a poor first choice even if it is the most exciting one.
Some of these tasks need no AI at all. If the rules are fixed and the data is structured, plain workflow automation may do the job more cheaply. A good scope says so.
Where an outside view helps with an AI pilot project
Filling in the eight lines is easy for some tasks and surprisingly hard for others. The hard part is usually lines 2 and 4: measuring the baseline and confirming the data. That is exactly what our free 5-day AI readiness audit covers for one department.
We sit with the team, measure the work, check the data and rank the tasks where AI would save the most time, ending with a fixed price for the first one. If the pilot then goes ahead, it follows the same rules as all our AI agents and automation: your data stays in your systems and a person on your team approves anything that goes out.
If you have a pilot in mind, or one that has stalled, you're welcome to bring the scope to a 30-minute call with a founder. We will tell you honestly whether it is tight enough to prove its value.








