Short answer
The order matters more than the list. Name one process before you look at any tool. Agree on the number that proves it worked, and the date you will check it, before anyone builds anything. Decide what data is allowed to reach a third-party model before it does, not after a client asks. Pick the smallest stack for that one process, pilot on a slice small enough that a bad week costs a week and not a quarter, build it with the people who will actually use it, then review the number at 30, 60, and 90 days and make a real decision — scale it, fix it, or kill it.
The checklist that starts with a tool is already wrong
Most guides sold as an AI implementation checklist start at step one with picking a platform or a model. That is a vendor checklist wearing an implementation label. Vendor selection is a short decision once everything before it is done properly — skipping straight to it is why so many pilots that looked fine in a demo quietly die within a quarter. None of the steps below are exotic; what kills implementations is doing them out of order.
If you have not already found and scored your candidate process, that step is covered in detail in our guide on deciding what to automate first. This checklist assumes you already have one process in hand and picks up from there.
Scope one process, not a category of work
Name a single, specific process before you evaluate any tool. Not a department, not "customer service," not "reporting" — a process with a name a person in the building would recognize. "Draft the first reply to a support ticket in under two minutes" is scoped. "Improve customer service with AI" is not, and it is the version nobody can ever agree is finished.
A process is scoped enough to move forward once you can answer these without checking with anyone else:
- What it is called — a name someone outside the project would recognize, not an internal codename.
- Who owns it today — the person who currently does it or is accountable when it goes wrong.
- How often it runs — daily beats monthly, because you find the edge cases within a week instead of a quarter.
- Roughly how many hours a week it costs — a working estimate is enough; precision comes later.
If any of those four takes more than a guess to answer, the scoping is not finished — go back to observing before you go anywhere near a tool.
Set the number and the deadline before you build
Agree on the one number that will prove this worked, and the date you will check it, before anyone writes a line of anything. Hours saved per week, error rate, cycle time, cost per case — pick whichever number the process itself already produces, not a generic productivity metric that sounds impressive in a slide. Write it down with a target value and a check-in date attached.
This comes before the stack decision because a target chosen after something already exists tends to get quietly reverse-engineered to look successful. The method for keeping that measurement honest is covered in our guide to measuring the ROI of AI automation honestly.
Decide what data is allowed to leave your systems
Decide, in writing, what data is allowed to reach a third-party model before the first prototype touches anything real. This is a business decision before it is a legal one — the kind of decision the NIST AI Risk Management Framework exists to help structure: if nobody can state in one sentence what a vendor's model can see, what it retains, and who can revoke that access, the project is not ready for production data — no matter how good the demo looked.
- What specific fields or documents leave your system — named explicitly, not "customer data" as a category.
- Where the vendor processes and stores them, and for how long after the contract ends.
- Whether the contract says the vendor will not train on your data — silence on this point is not a yes.
None of these need a lawyer to ask on day one. They need someone to ask them before the pilot starts collecting real customer information, which is the point at which most companies discover nobody actually checked.
Choose the smallest stack that solves the one process
Pick the smallest stack that solves the process you scoped, not the platform that could handle everything the company might ever automate. Oversizing the stack is the most expensive mistake available at this stage — the bill and the maintenance burden scale with the platform, not with the one process you are actually running.
Most small-business processes need three pieces at most: a no-code layer to connect the systems involved, a model for the part that genuinely needs judgment, and a place to store what the workflow produces. Whether that middle piece needs a model at all, or a fixed rule does the job for less money and fewer failure modes, is what the line between automation and judgment is for. If a no-code layer is new territory, what holds up in production versus what quietly breaks is covered in our guide to business process automation without developers.
If any of that output reaches people in the EU, one more question belongs on the checklist before launch: whether the EU AI Act reaches you, and which of its dates now apply after the July 2026 amendment — the EU AI Act for US and UK companies.
Run the pilot on a slice small enough to fail cheaply
Run the first version on a slice small enough that a bad week costs a week, not a quarter — one team, one ticket queue, one region, never the whole company on day one. A pilot's job is not to prove the idea works in general; it is to find the cases nobody anticipated while the blast radius is still small enough to fix by hand.
- An isolated, clearly bounded slice — a service line, a document type, a single location.
- A human safety net — someone reviews the output until trust is earned, not assumed.
- A hard decision date at the end — go, no-go, or one more iteration, never a pilot that quietly becomes permanent because nobody scheduled the conversation to end it.
Build it with the people who will use it
Build it with the two or three people who will use the result every day, starting in week one — not with a committee reviewing a finished demo at the end. Implementations fail on the adoption side more often than the technical side, and adoption is usually decided by whether the daily users had any say before launch.
- Train on their actual cases, with their real files, not a slide deck.
- Ask what they would need to see to trust the output before they stop double-checking it by hand, and build that visibility in from the start.
- Name one internal point of contact — not IT, not the vendor, someone who sits near the users and can answer "why did it do that" the same day.
Review at 30, 60, and 90 days — then decide
Put dates on the calendar for 30, 60, and 90 days before the project starts, not after — and treat a number that has not moved as information, not a verdict. At each check-in, compare the number against the target set at the start, and check whether the daily users are still using it without being reminded. Both conditions matter: a metric that improved but gets quietly worked around is not a success, and neither is a tool used out of habit while the number never moves.
Most gaps at the 30-day mark trace back to an earlier step rather than the model itself — a process less predictable than it looked during scoping, a target that was never the right metric, or a pilot slice too broad to learn from cleanly. Fix the step that actually broke, and reserve killing the project for cases where two iterations — a threshold from LYVIA's own engagements, not a fixed rule — have not closed the gap.
Frequently asked questions
What should be on an AI implementation checklist for a small business?
One scoped process with a named owner, a single measurable target with a check-in date, a written answer to what data a vendor can see, the smallest stack that solves that one process, a pilot small enough to fail cheaply, and a review date already on the calendar before the project starts. Skip any one of those and the project tends to drift even when the model itself works fine.
How long does a first AI implementation actually take?
For a well-scoped process, plan on roughly four to eight weeks from kickoff to the pilot going live — a range LYVIA sees across its own engagements rather than a published industry figure. Within that range — again a pattern from LYVIA's engagements, not an industry standard — scoping and the data decision usually take the first one to two weeks, the build another two to three weeks, and the rest is the pilot running long enough to trust the number. In LYVIA's experience, projects that stretch past three months are almost always ones that were never properly scoped to begin with.
What is a realistic first budget?
The build cost for one well-scoped process is usually the smaller line item. The one people underestimate is the recurring cost of model or API usage once volume is real, plus the time spent training and supporting the small group of people who use it daily. Get any vendor to quote the recurring cost at your actual expected volume, not a demo volume, before you sign anything.
Do we need a data scientist or AI specialist in-house to get started?
No, not for a single well-scoped process. Modern orchestration tools and model APIs remove most of the need for a dedicated data team on a first project. What matters more is someone who understands the process in enough detail to write the one-sentence answer on what data leaves your systems and to pick the right target — that is usually the process owner, not a data hire.
How do we know if an AI implementation actually worked?
It worked if the number you set at the start moved, and the people who were supposed to use it still are, thirty days after launch, without anyone reminding them. Either condition alone is a false positive: a tool that hits the metric but gets quietly worked around is not a success, and neither is a tool people use out of habit while the number never moves.
What is the most common reason an AI implementation fails after a working pilot?
Usually adoption, not the model. A pilot that ran well in a demo but never got a scheduled review date, or that launched without the daily users involved before week one, tends to get quietly abandoned even when it technically works — the tool never became trusted or habitual. That is why the review date and the user involvement belong before the launch, not after it.
If you want this run by someone outside the process — the scoping, the data decision, and a pilot sized to fail cheaply if it has to — that is where we usually start with a new client. Book a call.
