Milan Bogojevic Blog

AI transformation pilots fail at the handover, not the demo

6 min read updated 20 September 2026

A nonprofit AI pilot usually fails at the handover, not at tool selection. The demo works because the person who built it holds all the knowledge: which prompts work, which files to use, what a bad output looks like. When the pilot ends and someone else inherits the tool, that knowledge is missing. The fix is a short written record of decisions, tested by having someone else run the tool for a week.

Table of Contents

  1. Why this matters for nonprofits
  2. The NGO reality
  3. What a handover actually needs
  4. The test I use: the AI Handover Test
  5. What to do next
  6. Conclusion

Why this matters for nonprofits

A commercial company can absorb a failed pilot as a lost quarter. A nonprofit with a small team usually cannot. Time spent on a pilot comes out of program delivery, donor reporting or grant writing, and there is rarely a technical colleague who can rescue a tool once its builder moves on.

Three constraints make handover risk sharper in the nonprofit sector:

  • Small teams. One person often builds, tests and champions the pilot. That is a single point of failure.
  • Donor accountability. If AI output feeds a donor report or a grant application, an unchecked error becomes the organization's error.
  • Staff turnover and project cycles. Pilots often coincide with a grant cycle. When the cycle ends, so does the person's attention.

Frameworks such as the NIST AI Risk Management Framework (2023) also treat clearly assigned roles and documented accountability as a core part of managing AI. Small organizations rarely need that formality, but they do need its basic idea: someone owns the outcome.

The NGO reality

Consider an organization with eight employees and four donor-funded projects. The communications officer builds a workflow that turns project data into first drafts of donor updates. In the demo it is impressive. The director sees a draft in minutes that used to take half a day.

Here is what the demo hides. The officer knows to strip beneficiary names before uploading data. She knows the tool invents figures when a column is empty. She knows one donor's template needs a different tone. None of this is written anywhere, because she never needed to write it down.

The pilot ends. The tool passes to the program coordinator, who has capacity but no context. Within a few weeks, one of two things happens. The tool is abandoned, or it keeps producing drafts that nobody checks properly. The second outcome is worse.

Problem: knowledge lives in one head. Wrong approach: a handover meeting and a training session. Better approach: a written record and a real test. Typical result, in my experience: the tool either survives or fails visibly, instead of decaying quietly.

What a handover actually needs

A handover needs a written record of four decisions, not training. Training teaches people to click. The record explains why the tool is set up the way it is.

Record Question it answers Example of a good entry
Choice Why this tool and not the alternative? "Chosen because it works inside our existing document storage and does not retain uploaded files."
Boundary What is it allowed to touch? "Anonymized program data only. Never beneficiary records or safeguarding files."
Failure What does a bad output look like? One real example of a wrong figure or an invented claim, with the correction.
Owner Who do we ask when it breaks? A named person and a named backup, not a team or a role.

The person who ran the pilot knows all four. The person inheriting it knows none, and none of them are stored in the tool itself.

Each record connects to something you may already have. The Choice record is the short version of the questions in The Hidden Risk in Buying AI Tools, which covers data control, exit options and internal ownership before purchase. The Boundary record is the operational form of an AI policy, and What a governance policy for AI should not try to cover explains why a short one works better than a thirty-clause template.

The failure record matters most and is skipped most often. A single annotated example of a bad output teaches a new user more than a page of general guidance. For a fuller picture of how a wrong output reaches a donor report, see When AI Gets It Wrong.

The test I use: the AI Handover Test

The AI Handover Test is simple: before the pilot ends, the person who built it takes a week off, and someone else runs the tool. It is not a handover meeting. It is an actual week of real work.

Whatever breaks during that week is exactly what your documentation is missing. It is uncomfortable, and it is the cheapest quality check you will run.

Step Action Signal that you passed
1. Record Builder writes the four records above Each fits on one page
2. Swap Builder is unavailable for five working days Builder is genuinely not consulted
3. Break Successor runs real tasks and logs every stall Stalls are written down, not remembered
4. Fix Builder updates the records using the stall log Same stall does not recur
5. Sign off Director confirms a named owner and a review date Ownership exists in writing

My rule is simple: if nobody except the builder can run the tool for a week, the pilot has not finished. It has only been demonstrated.

I would not treat this as a technology decision. For most nonprofits, it is an operational decision first, and the real bottleneck is rarely the AI model. It is ownership of the process after the pilot ends.

What to do next

  1. Pick one live pilot and ask its builder to write the four records this week. Limit each to one page.
  2. Add one annotated bad output to the failure record. Use a real one from testing.
  3. Schedule the swap week before the pilot's official end date, and put it in the calendar now.
  4. Name an owner and a backup in writing, and set a review date 60 to 90 days out.
  5. Decide in advance what happens if the swap fails. Extend the pilot, simplify the workflow, or stop. Stopping is a legitimate outcome.

Conclusion

Most organizations treat the end of a pilot as a formality: the demo worked, so the tool moves into daily use. The real risk sits in the gap between the person who built it and the person who inherits it.

A pilot that cannot survive one week without its builder is not ready to become a process. Write down the decisions, test them with a real swap, and name an owner. That is a small amount of work, and it decides whether your AI investment lasts past the demo.

Next in AI Transformation