Skip to content
Built by SamirTalk to Samir →

Free guide

The Practical Guide to AI for Small Business

What AI can actually do for an operation your size, what it costs, how to find the work worth automating, and when the answer is to leave it alone.

No vendor is paying for this, which is why it can tell you not to buy anything.

Chapter 01

What AI can realistically do today

The five task shapes this technology handles well right now, the three ways it fails, and why a demo tests none of it.

Every other chapter here assumes you have a rough working picture of what this technology can and can’t do. So that’s what this one is for.

Two things get in the way of that picture. One is the pitch, where software runs your operation while you sleep. The other is the thing that falls apart the moment somebody asks it a second question. Both of those are real.

It’s tempting to ask what AI can do for a dental practice, or a plumbing company, or a firm of accountants. That question has no good answer, and not because nobody’s bothered to write it down. Work doesn’t sort by industry. It sorts by shape. The supplier paperwork in a workshop and the intake forms in an office are the same job wearing different clothes, so here are the five shapes, and then the honest other half.

The five shapes that work right now

Turning a mess into fields. An invoice, a scanned form, a voicemail, an email that buries the actual order somewhere in the third paragraph. Out the other end you get supplier, date, amount, reference, in the same slots every time. This is the one that changed most. The old approach needed a template per sender and it broke whenever somebody redesigned their stationery, and that particular failure is mostly behind us. A lot of workflows became worth a second look purely because of it.

Producing a first draft for somebody to edit. A reply, a product description, a job advert, the opening pass at a policy nobody wants to write. You’re paying for the blank page to go away, not for a finished piece. So if nobody is editing the output, what you’ve actually bought is publishing, and you’d want to know that.

Sorting and routing. Is this a new enquiry, a reschedule, a complaint, or a supplier chasing money? Pick one, send it to the right queue. What makes routing safe is that being wrong is loud. Whoever opens that queue can see in a second that the thing doesn’t belong there, and moving it costs nothing.

Summarizing, as long as the original is still sitting right there. Fine for the long attachment nobody was going to read to the end. Not fine once the summary is the only version that survives, because then there’s nothing left to check it against, and you won’t know what got dropped.

Flagging the one that doesn’t fit. Four hundred rows look alike and one of them doesn’t, so that one gets raised for a person to go and look at. Notice it isn’t fixing anything, which is exactly why it’s a safe place to start.

Three ways it fails, all of which you can design around

Leave it alone with something irreversible and you’ll regret it eventually. Payments leaving your account. A message that reaches a customer. A deleted record. A cancelled appointment. Which has nothing to do with how good the software is. It’s about what a mistake costs when there’s no way to walk it back.

It has no idea what’s absent. Give it a contract that’s missing page three and nothing in the reply will mention a gap. What comes back is smooth and confident, assembled out of pages one, two and four. There is no hole in the output where the hole in the input was, which is the whole problem.

And it’s likeliest to be wrong precisely where you can least afford it. It’s right on the ordinary items, which are also the ones where a single run saves you very little. And the strange item, the one your most experienced person would have spotted in a glance, is where the confident wrong answer comes from. So the value and the danger live in different places, and you have to build for both at once.

Which hands you the shape to aim for, and it’s the same shape in every good project in this guide. Software takes the ordinary case. Anything strange goes to a person. And nothing you can’t reverse happens without somebody releasing it.

A demo is a rehearsal

Rehearsed inputs, chosen beforehand, and the driver is the one person alive who knows that software best. Fair enough, I’d run mine the same way. But not one of those three conditions survives contact with your building on a Tuesday, and everything that goes wrong lives in that difference.

So here’s what’s in the gap. Four things, and none of them is about how clever the model is.

Your inputs come from people who don’t know this software exists. Customers and suppliers, plus the order that shows up as a photo of a form, cropped so the last two rows are missing. Nobody briefed them and nobody’s going to. In the demo, somebody typed the input already knowing the shape it had to arrive in.

Whoever drives it on the day is not the person who built it. It’s the one covering reception this morning, who sat through one training session for this back in March, and who has two people waiting right now. Software that argues with them loses, and honestly it should, because the customers in front of them are real.

Nothing in a demo has to deal with an item nobody can classify. Your operation gets those every week. Something arrives that fits no category, and there are two endings available. One is a queue with an owner who empties it before lunch. The other is a quiet corner where it sits until somebody notices in six weeks.

And then a week arrives when your one genuine expert on it is on holiday. That week always turns up. Not a fair test of anything, obviously, and still the week that tells you whether you own this thing or you’re renting somebody’s attention.

Four things to ask for instead

So make these four requests of whoever is selling it, in your words or in mine.

Show me the pile. Not an explanation of how errors get handled. The screen itself, with items sitting in it, and tell me who empties that screen and how often.

Show me what somebody joining on Monday gets handed. Nothing written is survivable, except then writing it becomes your job, and that’s a thing to find out in month one rather than month four.

Let one of my people drive it for ten minutes with everybody watching, on something that arrived this week rather than the run you’ve rehearsed.

Tell me what my operation does on a day this is down, because everything goes down eventually and nobody should be working that out at the time.

Building it in-house, with no vendor to interrogate, changes none of this. You put the same four demands to yourself in an empty room, which is the harder version, and skipping them is how a promising prototype becomes a project that never quite lands.

One habit to keep, ahead of any tool evaluation. Name which of the five shapes your task falls into, out loud, to another person. If it won’t go into one of them, find that out at a whiteboard now, rather than after somebody has quoted you for it.

Chapter 02

AI vs automation vs agents

Three words that get used as if they mean the same thing, what each one actually is, and a single invoice walked end to end to show you how little of the work needs AI at all.

These three words get thrown around as though they’re interchangeable, and they’re not. The difference decides who has to check the output, how much it costs you when the thing is wrong, and how much of what you’re being quoted for is ordinary software with a better name on the invoice.

So here they are, one at a time, and then one worked example with all three in it.

Automation is fixed rules on a trigger

Something happens, and then a sequence you wrote runs. A form gets submitted, so a record is created, a notification goes out, a field changes, a row is added to a report. Nothing in there is deciding anything.

The defining property is that the same input produces the same output every time, and you can read the rules and say in advance what it’s going to do. That’s worth more than it sounds. It means you can test it, it means it behaves in December the way it behaved in August, and it means when it goes wrong you can find the step that went wrong.

Most of what gets sold as an AI project is mostly this. That is good news, not a scandal, as long as you know that’s what you’re buying.

AI is the judgment step inside it

AI is the one step in that sequence that couldn’t have been written as rules. Reading a document somebody else formatted. Working out which of four queues this belongs in. Turning six bullet points into a paragraph a person will then edit.

It’s a component, and the system around it is ordinary software. It also comes with a property the rest of that software doesn’t have, which is that the same input doesn’t guarantee the same output twice. So it belongs in places where the output can be checked cheaply, and where anything you can’t reverse waits for a person. That arrangement is the only one where you find out you were wrong while being wrong is still cheap.

An agent chooses its own next step

An agent is AI that decides the sequence instead of following one. You hand it a goal and some tools (a mailbox, a database, a browser, the ability to send something to somebody) and it works out what to do first, what to do after that, and when it’s finished.

The risk changes shape here. With a judgment step, you check one output against one input, and you can sit down and read a month of them. With an agent, the thing you’d have to check is a route through your systems that you didn’t specify, and it may take a different route next time on the same request. A wrong field is a wrong field. A wrong sequence of actions is an incident, and it’s already happened by the time you’re reading about it.

That doesn’t make agents useless. It narrows where they belong: work that’s reversible, inside a boundary you drew on purpose, where the whole route it took gets written down and somebody actually reads it, and where nothing irreversible happens without a person releasing it. If what’s being sold to you is software that pays, sends, cancels or deletes with no person in the way, the thing being removed from your operation is your confirmation step. Ask what you’re getting in exchange for it.

Which is why the first project probably shouldn’t be an assistant over your documents

An internal assistant is the judgment step with no automation around it. Nothing but a person can start it, and it reads whatever it was pointed at. Its output is prose, so checking the answer means already knowing the answer. And it writes nothing back anywhere, so nobody can go and find out later how often it was wrong.

It also changes how people ask for the work rather than changing the work. Before, somebody opened three systems and read until the answer turned up. Now they ask a question and read one paragraph, which is less reading and no other change at all. Those three systems haven’t changed. They hold what they always held, scattered and half up to date. What’s new is that something reads that material for you and repeats it in one confident sentence, so an out-of-date answer now arrives sounding certain, which the folder never managed.

A project lands here when nobody could name a workflow, and that’s the actual diagnosis. Worth saying out loud in the meeting where it comes up.

What to look for instead comes down to four properties, none of which concern the model at all. It gets started by something other than a person, so a clock ticking over, a file landing, a status changing. The records it reads belong to you and turn up in a shape you set. Checking the output takes a colleague a few seconds and never involves producing it a second time. And it leaves a written trace, so a month later you can count the times it got something wrong. Run a monthly sweep of your renewal dates and you’ve got all four, and after one month you can say plainly whether it worked.

One task, walked end to end

So take a supplier invoice landing in your inbox, and walk the whole thing out loud, one step at a time. I’ll do this one, and then you do the same for whatever tool you were about to sign for.

What starts it. The invoice arrives at an address that receives nothing else. That matters more than it looks. Nobody has to notice it, and nobody has to remember to begin.

What it reads. Four records. The invoice itself. The matching purchase order, found by its reference in your purchasing system. The delivery record of what actually turned up. And the terms you hold for that supplier.

The decision. One decision, which is really three comparisons. Is this a supplier you have on file? Does every item billed match an item that was ordered, at the price that was ordered? And does the billed quantity match what the delivery record says arrived?

What gets written back. A draft entry in payables, coded, carrying the order reference, plus a note naming whichever of the three comparisons failed.

Who confirms. A person releases the payment, because you don’t get money back once it’s gone.

Now count the AI in that. One step. Turning a document somebody else designed into fields. The three comparisons are arithmetic against records you already own, and writing the draft entry is ordinary database work, and between them they are most of the build. All worth knowing before you accept a quote for an AI project, because most of what you’d be paying for is plumbing.

Notice also that nothing in that sequence needed an agent. The order of the steps was known before it ran, and when the order is knowable in advance, writing it down beats paying software to rediscover it every morning.

Now do it with your own task. Five things to name: the trigger, the inputs, the decision, what gets written back, who confirms. Any of the five you can’t fill in is one a vendor won’t be able to fill in either, and a proposal that leaves them blank is a number attached to work nobody has defined.

Chapter 03

How to find AI opportunities inside your company

Eight questions in a fixed order, each with the answer that kills the candidate, and two worked examples that land on different verdicts.

You don’t find these by studying AI. You look at work you already do, one workflow at a time, because the answers live in how your operation works, not in a demo.

Be strict about what counts as one: you can say what triggers it, what comes out at the end, and whose desk it lands on. Miss one and you’re holding a department, and the answers come back mush.

Eight questions, in this order. Each has an answer that kills the candidate, and each of those is either permanent or repairable, which is how you get three verdicts instead of a score. Cheapest to answer first, and among the cheap ones, the questions that can end it before the ones that only postpone it. So “is the work repetitive and rule-based,” where the familiar checklist opens, sits at 7 and 8, because it costs the most to answer honestly and kills the fewest candidates.

The first three, answerable in a conversation

1. If this stopped for a month, who would come asking? A person or a system, not a department. What kills it: nobody would come, and nothing is being prevented either. Stop it and see who complains. Permanent.

But hold one category open, or it eats your best candidates. Checks whose whole value is that nothing happens: verified backups, expiring certificates, an access list checked against who works here. A check counts, and what it protects is the customer. A working check produces silence, and so does an abandoned one, right up until the failure arrives. And never run that stop-and-see experiment on anything preventive, because the only way it reports back is by letting the failure through.

2. Does this work exist only because something upstream is broken? What kills it: the work is a workaround. Two systems don’t exchange data, so somebody retypes between them. When the upstream fix costs less than automating around it, the verdict is don’t, and that same answer hands you the real project. Automate the workaround and the defect becomes cheap to live with, which is worse than annoying, because annoying gets fixed. Permanent. One exemption: a check that verifies another process is not a workaround for it, so question 1’s controls survive here.

3. Could somebody who didn’t produce the output tell whether it’s right, faster than producing it again? What kills it: the only real check is redoing the work. Automation then moves the labor from producing to reviewing and you’ve bought nothing.

Also kills it: nobody can say what a correct output looks like. Two reasons there, two verdicts. Never-written criteria are repairable, so go and write them. An output that’s a judgment the person making it can’t reduce to criteria is permanent, and it’s the worse of the two, because an output nobody can verify keeps supplying evidence in its own favor. Criteria nobody has written can be written. A judgment can’t be.

The next three, which cost a bit more

4. How often did it run last month, and what starts it? Write your volume threshold down before you count. What kills it: runs times minutes lands under that threshold and being wrong costs nothing. Note the and. Volume and error cost substitute for each other, so twice a month still passes when a mistake is expensive, and question 5 settles that.

Then the trigger. “Somebody notices and starts it” gives software nothing to listen for. That’s repairable when you can instrument the trigger, and permanent when the trigger is itself a judgment, as in “when the case looks complicated,” because then that judgment is the candidate and it has to pass question 3.

5. What’s the cost of a wrong output, and who notices? This one mostly doesn’t vote. It sets a requirement on the build instead.

The usual answer: the error is expensive, hard to reverse, and reaches a customer, an employee, a regulator or your own money before anybody notices. That candidate isn’t disqualified, it’s constrained. Software drafts, a person releases, and that goes in the specification.

An error nobody catches argues for building rather than against it, because an automated sequence leaves a record and the version you have now leaves none.

One branch does end it. Same harm, and nowhere in the sequence could a person approve the irreversible step. An approval given by the same person the automation is acting for is not an approval. A one-person back office where software releases outbound payments has no approval point to insert, because the only available approver is the hand that prepared the payment. The verdict is don’t, and it stays don’t until a second person exists. Permanent.

6. Who owns this today? One name. What kills it: nobody, or three people who own a piece each and describe the steps differently. Always repairable, never skippable, but not deferrable: every repair gets a name and a date, or it’s a don’t in better clothes.

The two expensive ones, which need a month of real work

7. Do the inputs arrive in the same shape every time? What kills it: inputs turn up however the sender felt like sending them, and getting them usable is the actual work. Fixable when you control the intake, which means one channel with required fields. Permanent when the sender is outside your control, because then normalizing the input is the judgment.

8. Could a competent stranger run this from written instructions alone, and how many runs would need a ruling those instructions don’t give?

First failure: they can’t be written, because the procedure lives in one person’s memory. Repairable. Second failure: they exist and too many runs need a ruling they don’t cover. Pick your share and write it first. One in five is my rule of thumb. Repairable by splitting the workflow, so the majority path comes back through these questions and the exception path stays with a person on purpose. Permanent only when nobody can characterize the exceptions at all, which means somebody is deciding an unwritten branch by feel.

Reading the answers

Go with the worst answer, not an average. One permanent failure decides it no matter how the other seven read. And don’t total them out of eight, because a checklist that averages says yes.

Don’t means at least one permanent failure. Write the verdict as one sentence with the question number attached, and keep it, because this workflow comes back in six months anyway.

Not yet means no permanent failures and one or more repairable ones. I’d expect that verdict most often, and it’s the one that pays for the afternoon, because those repairs make a workflow quicker before any software exists.

Automate now means nothing ended it, so the candidate is worth pricing. Do that arithmetic before talking to anybody, because it kills candidates that survived all eight. Chapter 07 is how.

Two worked examples

Both are workflow shapes an operation your size tends to have.

A monthly access check. Question 1 looks fatal, because nobody comes asking when a check quietly works. It isn’t: this is a check, so the carve-out covers it. After that it’s two lists and one comparison, checkable in a minute. The error it catches, access outliving somebody’s departure, is one nobody finds out about today, which argues for building it. Revoking is the irreversible half, so it proposes and a person releases. Automate now, with revocations gated on a person. One warning: if the same failure turns up monthly, offboarding is missing a step, and that’s the project.

Onboarding a new hire, end to end. Questions 1 to 5 pass, and 5 carries it, because a handful of hires a month is under any sensible threshold. A missing account on day one is embarrassing and fixed by eleven, but access that outlives a departure is serious and nobody finds out. Then 6 finds nobody owning the sequence, 7 finds requests arriving by message and in person and as a forwarded letter, 8 finds nothing written down. All three repairable. Not yet. Name one owner, have them run the next two hires from a checklist they correct as they go, require role and start date on one intake request, and put a date on the re-test.

It costs a notebook and an afternoon. If you run it across your workflows and decide to fix three processes and buy nothing, it worked.

Chapter 04

What to automate first

You have a list of candidate workflows and the budget for one of them, so here is the arithmetic that puts them in order, including the term almost everybody leaves out.

Say you’ve gone looking and you’re holding a list. Six or seven pieces of your operation that software could plausibly take over, all of them real, none of them obviously the place to begin. You can fund one this quarter. Which one?

The way this usually gets settled is that somebody in the room feels strongly, and whatever annoyed the most people last Tuesday goes first. That’s not a stupid instinct, and I’d rather work with someone who has one. It’s just that it picks wrong in a way you can see coming with a pen and four numbers per candidate.

The arithmetic, and the term everyone drops

Count how many times each candidate happened last month. Then work out what one run of it is worth to you, and multiply the two.

The second half is where this goes sideways, because the obvious way to work out what a run is worth is to time it. A run is worth more than the time it takes, and also less. More, because some share of runs come out wrong today and wrong has a price. Less, because once software is doing it, a human still has to look at the output and decide whether to trust it.

So one run has three pieces to it. The effort of doing it now. Plus the damage a bad one does, averaged across every run. Minus the effort of checking one output afterwards.

That third piece gets subtracted, and it sits inside the multiplication, which is the entire reason your ranking moves. Build effort mostly lands once and then it’s done. Checking lands on every run, forever, for as long as the thing is switched on. So volume scales up the payoff and the checking bill at exactly the same rate, and if a check costs anything close to what the work costs, high volume stops being an argument in favour and starts being an argument against.

Minutes or money, but pick one

Four numbers per candidate: runs last month, how long a run takes, how long a check takes, and what a bad run does to you. Three of those arrive as time. The fourth arrives as money. You can’t add them together in that state, and a spreadsheet will let you try, which is how this gets shipped wrong.

I’d work in minutes, since three of the four are already there. To bring the fourth across, take what a bad run costs you in money, divide by what one hour of the relevant person costs the business all in, and multiply by sixty. A mispriced quote that eats a day of somebody’s pay is 480 minutes. Then weight it, because a failure is only worth counting as often as it actually happens. One bad run in twenty, at 480 minutes each, is 24 minutes a run.

If you’d rather see money, go the other way and price every minute at that same all-in hourly cost. All-in meaning what the hour really costs you, not the figure on the offer letter. Either version gives you the same ranking, because you’re applying one scale to every candidate, and it’s worth knowing that so nobody wastes an afternoon arguing about which version is the honest one.

Three candidates, ranked twice

Invented numbers. The shapes are ones you’ll recognize.

Coding card transactions. 600 a month. Four minutes each to open it, read it and pick a category. Checking one means opening the receipt and reading it yourself, so three and a half minutes. Call it one in a hundred landing in the wrong category, costing an hour to find and unpick at month end, which is 0.6 minutes a run.

Checking supplier certificates. 120 a month. Nine minutes each, most of which is hunting for the document. Checking takes thirty seconds, because a date either agrees with the record or it doesn’t. One in twenty gets recorded wrong at two hours to sort out, so six minutes a run.

Pricing quotes off site-visit forms. Eight a month. Forty-five minutes each. Four minutes to check, since you’re re-adding a short list against a price file. One in eight comes out wrong, and a wrong one you’ve already sent costs you roughly a day, so sixty minutes a run.

Score them on effort plus damage, skip the subtraction, and you get the order anybody in the room would have guessed. Card transactions 2,760 minutes a month. Certificates 1,800. Quotes 840.

Now subtract the check. Card transactions drop to 1.1 minutes a run, certificates to 14.5, quotes to 101. Multiply back out by volume and the board reads certificates 1,740, quotes 808, card transactions 660.

Same three workflows, same month, and the one that was ahead by 960 minutes is now last by a distance. The middling-volume job where a check takes thirty seconds beats the six-hundred-a-month one by more than two and a half times.

Why the loudest candidate wins arguments it should lose

Six hundred a month is the figure that gets said out loud. It sounds like scale, and scale sounds like the answer.

But six hundred outputs a month that can only be verified by doing the work over again is six hundred reviews a month, and that’s a job now. Not a launch task, a standing one. And a review that costs nearly what the work costs gets quietly dropped somewhere around week three, which is the worst result on the menu. You’re still paying for the software, the reviews have stopped, and nobody ever decided that out loud.

So a candidate where the check costs about what the work costs has a design problem, and the arithmetic is only reporting it. Usually it’s telling you to go back and find the version of that workflow where being right is a comparison rather than a judgement. Does this date match that date. Does this total agree with that total. Is this the same name as that name. Volume only helps you once the check is cheap.

Then go and do it on four of them

Pen, paper, four candidates, four numbers each. Runs last month, minutes a run, minutes a check, and the cost of a bad one converted into minutes and weighted by how often it shows up.

Two things I’d expect. The order won’t match what you’d have said from memory, and the thing that made you angriest this week won’t sit at the top of it. Neither is grounds for overruling the result. If you want to overrule it, change a number and say which one, out loud, in front of whoever gave it to you.

Chapter 05

25 practical examples

Twenty-five workflow shapes software can already handle, split into the ones where a mistake stays inside the building and the ones where it doesn't, each with what starts it, what it reads and what it hands back.

Every example here is written to the same three-part shape, because a description of automatable work is useless without all three. What sets it off. What it reads. What it puts in front of a person. If somebody demos a thing at you and you can’t name those three afterwards, you haven’t been told what you’d be buying.

None of these is a case study. They’re shapes, and the reason there are twenty-five of them is that operations of a similar size tend to run on a similar set of chores. Read it like a parts catalogue. Most won’t apply. Three or four will land hard enough that you’ll picture the person whose Friday afternoon this currently is.

Back office first, where a mistake stays indoors. Then the customer half, where it doesn’t, which changes how every one of them has to be built.

Start with the money, where a check takes seconds

Your bank feed. It updates overnight, and that’s the trigger, so nobody has to remember to start anything. Each row gets compared against open invoices and card charges on amount, date and reference. Agreements get coded and posted. Everything else comes back as a short list with a sentence against each one saying what didn’t line up. The pile a human reads shrinks from four hundred rows to eleven.

A remittance advice. A customer settles nine invoices with one transfer and emails a PDF explaining how. It gets read, allocated back across the nine, and the one deduction they took off gets put in front of a person with the invoice it came from. The deduction is the whole point. Absorbed quietly, that’s the thing you discover in November.

Receipt photos. Staff take pictures of receipts and drop them in a group chat, because that is where staff will genuinely do it. Each one gets read for supplier, date, total and category, then paired with the card charge that already cleared. Only the ones with no matching charge need a human, which is a much shorter conversation.

The payroll comparison. Before the run submits, it gets laid beside the previous one, and anything that changed comes out as a list. New starter, leaver, changed hours, changed rate, a bonus. Somebody writes a reason next to each. One page, four minutes, and it catches the thing that would otherwise surface a quarter later.

Paperwork somebody else designed

The clever part of these is unglamorous. It’s reading a document whose format you had no say in.

  • Certificates of insurance. A supplier sends one, the cover dates get pulled, and the insured name gets compared to the name on their application. Either the record opens or a reply goes back naming which document is wrong.
  • Leases and supplier contracts. Read for the dates that can hurt you. When it renews, how much notice you owe, how the price escalates, each one into a table alongside the page it was found on, so anybody can check the reading in a click.
  • Scans in a shared drive. A file turns up called scan_0042.pdf, gets a name matching the record it belongs to, and lands in the right folder. Anything unrecognisable piles up in one place rather than twelve.

The ones a clock starts

Nothing here needs more than a date and a comparison.

  • Timesheets close on Friday and get held up against the roster. Two lists fall out. Shifts nobody has claimed, and claims for shifts nobody worked.
  • The monthly return your insurer or your regulator wants. Twelve figures, four systems, one template, every month. Software fills eleven of them. The twelfth needs somebody to make a call, so it comes back empty with a note saying what the call is.
  • Utility bills for every site, read and set against last month’s, with each movement flagged under the site it came from. Nobody opens forty bills by hand, which is precisely the argument for it.
  • A nightly pass for near-duplicate customer records. One phone number, three spellings of a name, presented as pairs for a person to merge or wave off. It merges nothing by itself.
  • A month-end draft note saying which figures moved and by how much. Which, not why. Why is a judgement and it stays yours.

The customer half, where nothing sends itself

Thirteen more, and every one stops at a person before anything leaves the building. Which isn’t me being precious about it. A wrong output on this side reaches somebody who doesn’t work for you, and there is no recalling it. So the pattern is draft, then release, without exception.

An enquiry form arrives. Whoever sent it gets looked up, and a draft reply comes back answering the three things they actually asked, carrying the account number and both dates. Somebody reads it and presses send. The job stops being a hunt across two systems and becomes reading a paragraph.

The shared inbox. Each message gets read for which site and which service it concerns, then dropped into the right queue with the record attached. The first hour of somebody’s day stops being sorting.

Voicemail left after you close. It gets turned into text during the night and split by whether that caller sits on your records already, so a callback list in that order is waiting before anybody has taken a coat off.

A summary after every call, written into the same fields in the record every single time, so the next person to open it sees what was said rather than “spoke to customer”.

Quoting, where being slow is visible to the buyer

The survey form comes back carrying measurements and photographs. Those get priced off the price file you keep up to date, and what lands in front of you is a quote with every assumption written underneath. You check the total and send it yourself.

Five days later, silence. A follow-up gets drafted that names the quote number, what’s on it, and the date it expires. Which is a different email from “just checking in”, and it’s the one that gets answered.

Scheduling, once a conversation is already open

  • A reminder goes out and the customer asks for Thursday. That reply gets read, compared against what’s genuinely free, and turned into a proposed change that a person confirms.
  • Somebody cancels and a slot opens up. Everyone who wanted that window hears about it, one at a time, until one of them says yes.
  • Two days out, one intake form is missing. It gets chased from the single person who hasn’t sent it, rather than everybody getting the same nagging reminder.

The chasing nobody gets around to

  • A public review lands. A reply gets drafted, and it’s allowed to say only what your own records back up. If the review describes a specific incident, nothing gets drafted at all and the whole thing goes straight to a person, because that one is a conversation rather than a reply.
  • Customers who’ve gone quiet since whatever date you pick get pulled into a call list that shows their last order and its date.
  • An invoice tips past terms. A chase gets written quoting the invoice number, the amount outstanding and the day it fell due. Somebody hits send on that one personally, because it’s arriving inside a relationship.
  • A list every morning of enquiries older than a day where nobody has written a word in the record. Cheap to build, and it turns something up in the first week.

What all twenty-five have in common

Not one of them decides anything. Every one produces a comparison, a draft or a list, and every one can be checked in seconds by somebody who didn’t produce it. The checkability is doing the work there, not the cleverness. Software can manage plenty more than anything on this list. What it can’t do is survive month six inside an operation where nobody can tell whether its output was right.

So take the one you recognized hardest, the one where you already know whose job it is, and go and find out how many times it ran last month. That count is the front end of the arithmetic in chapter four, and it decides whether any of this deserves your money.

Chapter 06

How much AI implementation costs

Market ranges for the three things you actually pay for, plus the five costs that never make it onto a proposal and can each come out bigger than the ones that do.

Cost is the first thing anybody wants to know and the hardest thing to find written down anywhere, so here is the shape of the spend, roughly in the order you’ll run into it.

Ranges, not precise figures. Anyone who quotes you to the dollar before looking at your systems is guessing, and a confident guess is worse than an honest range, because you can’t argue with it.

Per-seat tools, the one piece with a price you can look up

This is the chat assistant your staff keep open in a tab, plus the AI features your existing software vendors have started charging extra for. Reckon on $20 to $40 a head per month, with premium tiers and add-ons pushing toward $60. A team of twenty puts you around $500 a month, and if it isn’t earning that, you stop paying in April.

It’s the only piece of the bill that behaves like normal software. Published price, monthly term, and you can go and read it right now without booking a call. It’s also the easiest money on this page to waste, for the reason in the last section of this chapter.

Usage, which is far smaller than your gut insists

Build something that runs unattended and it reaches out to a model over an API on every run, billed by the token. A token is roughly three quarters of a word. Feeding text in costs cents to a few dollars a million tokens depending which model you pick, and what comes back is dearer, a few dollars up to twenty-odd for the same million.

Meaningless until you put your own workflow through it, so do that now. Say the thing fires 400 times in a month, and each run reads a couple of pages and writes half a page back. Round everything up hard and you land at maybe two million tokens across the month. That’s tens of dollars at the expensive end. Cheaper still, single digits, if a smaller model can handle the job, and for lifting four fields off a form, a smaller one usually can.

If that sounds too low, run it again with your own volumes. The compute is cheaper than the person who reviews its output, by a wide margin. Worth having in mind the next time a monthly platform fee turns up with “AI compute” printed beside it as the justification.

The build, which is where the money really goes

This is the one that arrives as a single number at the bottom of a page. Budget five figures for a first serious workflow. A four-figure quote usually means nobody has looked at your systems yet and you’re being sold a template. Six figures means what you’re buying is a platform, not a single workflow, and that’s a different decision made for different reasons, so don’t let it arrive dressed up as this one.

And hardly any of that spend is AI. The model call itself is a day of work, two at the outside. The rest is getting data out of a system built to keep it, reconciling records that almost match, working out what should happen when a required field is empty, and building the alarm that goes off at 2am when the whole thing stops. Plumbing, mostly.

Which hands you one good question for whoever wrote the quote. How many of these hours are integration and error handling, and how many are the model? If the answer comes back as “nearly all of them”, without a pause, you’re talking to somebody who has done this before. If it comes back as a pause, you’re the practice run.

Five costs that won’t be on the proposal

A proposal shows you two numbers. What the build costs, and what the month costs. Then there are these five, and any one of them can come out larger than either.

Checking, on every run, forever. Somebody has to read what came out and rule on whether it’s right, and that is not a launch task. Three minutes an output at 400 runs a month is twenty hours of somebody’s month, permanently. Price those hours at whatever that person costs you, and the yearly total is a number no proposal will ever show you. Stretch that across three years and it comfortably outgrows the build figure.

Your data, wrong in your own particular way. It got typed in by humans over a decade. The same customer exists twice since 2019. One supplier’s name is spelled three ways. An address sits in a notes field because the field meant for addresses wouldn’t accept it. Somebody has to clean that up before anything automated touches it, that’s usually weeks and not days, and it’s your own staff who have to do it, since nobody else can tell which of two near-identical customer records is the live one. Ask who does that work and whether it’s inside the price you were quoted.

The software you already own. The thing you need to plug into may turn out to have no API whatsoever. Or there is one, but your subscription doesn’t include it, and the level that does costs a few hundred more a month, permanently. Or a connector exists and it only reads, when writing back was the entire point. Find this out before anybody prices anything, because it’s what separates a two-week integration from a three-month one.

Someone whose job this is once it’s live. Not building it. Owning it. Your billing system pushes an update that quietly changes what one status code means, your automation starts producing confident nonsense, and it needs a name attached to noticing. An honest answer here is a couple of hours a month. Zero is not, and a proposal that doesn’t mention ownership at all is claiming zero.

The seats nobody opens. This is the expensive one, and the least dramatic. Twenty licences at $30 each is $600 a month, so call it $7,200 across a year, for software your team logged into twice in March and hasn’t touched since. Nothing broke. The vendor delivered exactly what was sold. The money goes out anyway, quarter after quarter, and nobody has to make a single decision for that to continue.

Which is why I’d put money into one workflow that runs whether or not anybody remembers it, ahead of twenty licences that need people to change a habit on a Tuesday. The workflow’s failure is loud. The licences fail silently, and silent failure never makes it into a review meeting.

Then total the ones that weren’t on the page

Add the five up, set them beside the two numbers you were given, and then ask whether you’d still sign. That’s the real figure, and it’s the one you’d have arrived at eventually anyway, just later and with less room to move.

A vendor who walks you through those five before you ask has done this more than once. One who hasn’t thought about them is going to find them out on your budget.

Chapter 07

How to calculate ROI

The hard part of proving a return isn't the arithmetic, it's that nobody wrote down what things looked like before you changed them.

Chapter 04 is about ranking your candidates, which is a forecast. You’re guessing, carefully, at what a build is worth before you do it. This chapter is the other question, the one that turns up six months later when somebody asks whether the thing you paid for actually paid you back. Same subject, different problem, and the second one is harder for a reason that has nothing to do with multiplication.

The problem is that nobody kept the before

A return is a comparison, and a comparison needs two sides. The new way you can measure any time you like, because it’s running in front of you. The old way is gone.

Six weeks after a change goes in, ask anybody in your building how long the old process used to take. You’ll get an answer. It will be delivered with confidence, and it will mostly be a feeling about the project rather than a measurement of anything. If they liked the change, the old way was a disaster. If they weren’t asked first, the old way was working fine, thanks. Neither person is lying to you. That’s just what memory does when there was never a record to check it against.

So you end up in one of two bad places. Either the return is unprovable, and the whole thing gets filed under “seems better,” or somebody reconstructs a before figure from vibes and now you have a business case built on a number that was chosen after the answer was known.

Which means the first job isn’t the build

It’s measuring how things work right now, this week, while right now is still available to measure.

This part is duller than it sounds and it’s the whole ballgame. Export whatever your software will give you and start counting. How often did this task happen last month. How many of those had to be redone because something was wrong. How long, in days, between a request arriving and somebody dealing with it. Three numbers, and most operations can get all three out of tools they already pay for.

If your systems won’t tell you, sample instead. Time thirty of them with a stopwatch, and have it done by whoever’s hands are actually on the work rather than by whoever supervises it, because a supervisor’s estimate is always the version where nothing goes wrong. Thirty is small enough to finish this week and big enough that nobody can wave it away in a meeting later. Ten is an anecdote. Three hundred is a project you’ll abandon.

Decide what would count as a win, out loud

Before anybody is emotionally attached to the answer, and definitely before anybody has spent money. There are three kinds of return worth keeping apart, because they aren’t interchangeable and people mix them constantly.

Hours that come back to you. Money that stops leaving the building. And errors that stop occurring, which is the category everybody forgets, because an error that never happened leaves no evidence anywhere. There’s no ticket, no complaint, no refund, nothing to point at. Name it up front or you’ll never be given credit for it.

Hours come with a condition attached. An hour is only a return if it lands on work you’d otherwise be paying somebody to do. So if the person who used to re-type addresses all morning now spends that hour doing the job you actually hired them for, put it in the case. If instead the afternoon is simply less frantic than it was, what you bought is slack. Slack is a genuinely good thing to buy. It’s just not a return, and calling it one is how a business case gets quietly discredited later by the one person in the room doing the subtraction.

Pick the window now as well, and make it long enough to include at least one bad month. Opening weeks flatter everything, because whoever pushed for the project is personally watching it, spotting problems before anybody else has to notice them. Nothing you build will ever be watched that closely again.

Working one example all the way through

Take the phone, because the shape of it will be familiar and the before figure is unusually easy to get hold of.

A receptionist can hold one conversation at a time. Training doesn’t change that, and neither does a nicer phone system or a better attitude. So at lunch, after you close, and any time two calls arrive together, somebody gets voicemail. A handful will leave a message. Everybody else rings whoever is next in the search results, and that revenue is gone before anyone in your building knew it existed. Nothing shows up in a report, because a caller who gave up isn’t an event any of your software cares about.

Sizing it takes five numbers and you have all five sitting around already. How many calls arrive on a normal day. What percentage of them nobody gets to. How many of those missed callers were strangers rather than people you already look after. How often a new inquiry turns into actual work. And what one customer is worth to you across their whole time with you. Multiply those together, scale it up to a year, and you have a figure for revenue that left because a phone rang out.

Two things about that number. First, the before side of it is genuinely available, which is rare. Your phone provider knows how many calls came in and how many nobody answered, and that record exists whether or not you ever build anything. Second, the number sizes the decision in both directions, which is the part I’d care about. If it comes out small, your staffing is doing its job, there’s nothing to build, and you’ve just saved yourself a project. If it comes out large, you’re looking at a systems problem, which means overflow routing, callback handling, or something that picks up when your people physically can’t. You can’t make either of those calls without the figure sitting in front of you.

Worth noticing while you’re there that hiring lifts the ceiling by about one simultaneous conversation, only while that person is on shift. And the shifts nobody wants to staff are exactly the ones where the calls go missing.

If you’d rather not do the arithmetic by hand, the calculator on this site will run it off those five inputs. It saves nothing and transmits nothing anywhere.

One page, dated, before the work starts

That’s the whole discipline. Today’s figures and where each one came from. What counts as a return, split into hours, money and mistakes. The date you’re going to look. Signed by whoever is going to be asked about it later.

It’s an afternoon of work and it’s boring and it will feel like paperwork standing between you and the interesting part. Do it anyway, because there’s no version of next spring where you can go back and get it. The alternative is a project everybody has an opinion about and nobody can settle.

Chapter 08

Build vs buy

There are three answers here, and the third one, don't, is the one that never makes it onto a slide because nobody in the room is paid to put it there.

Start with the disclosure, because it changes how you should read everything after it. I build software. Custom work is what I do, so when I tell you that buying is usually the better answer, I’m arguing against my own commercial interest. Weigh it accordingly. This isn’t neutral advice from nobody in particular.

Start from buying. Anything you build has to earn the exception.

Buying wins more often than it feels like it should

The test is whether the problem is a generic one. More of yours will turn out generic than they look when you’re standing in the middle of them. Payroll, for instance. Bookkeeping, very much so. Scheduling, invoicing, and most of what happens between a customer asking for something and you sending them a bill.

Someone has already solved those, and they’ll carry on solving them through the years you aren’t paying attention, which counts for a great deal more than any feature comparison you’ll build in a spreadsheet. Tax rules change. A payment processor deprecates something. A browser update breaks a form. All of that lands on somebody else’s roadmap and never reaches your inbox.

Accountability moves with it. When something you bought breaks at nine on a Monday there’s a company on the hook, somebody to call, and a contract that says who owes whom. With something you had built, the party responsible for fixing it is you, permanently, including the two weeks you were supposed to be on vacation.

Building earns its place in three situations

I’ve tried to think of a fourth and I haven’t managed it yet.

First, when the way you do this one thing is unusual on purpose, and that difference is part of why customers choose you instead of the place down the road. This comes up less often than it feels like it should, because every business sounds distinctive when its owner is describing it. The honest test is whether a customer would notice if you switched to the standard way of doing it. If they wouldn’t notice, what you have is a habit, not a differentiator.

Second, when the configuration adds up to building it anyway. You buy the platform, then you spend four months bending it into the shape of your operation with custom fields, rules, and a consultant. You’ve built something. The difference is that you built it inside a product somebody else owns, so you can’t read back what you made, can’t test it properly, and can’t take any of it with you the day the pricing changes.

There’s a second half to that bill nobody quotes. You also reshape how the work actually happens to suit what the platform expects, and your people pay that every week, forever. It never appears as a cost anywhere in the business, which is exactly why it never gets weighed against anything.

Third, when the connecting work is itself the product. You have five tools, each perfectly good at its own job, and what’s missing is whatever should sit in the middle of them. No vendor sells that, because its shape is the shape of your particular operation. Chapter 09 goes into what it takes.

Then there’s don’t

Don’t is a legitimate answer, and I’d guess it applies more often than the meeting discussing it will allow. Three versions.

The first is when the workflow itself isn’t settled yet. Maybe intake gets redesigned after the new year. Maybe one of your core systems is being replaced. Maybe a second location opens and will do half of this differently. Automating a process that’s already on its way out is paying to make something obsolete faster, and waiting costs nothing. You’ll also get a cheaper build later, because you’ll be building the version that’s going to exist instead of the one that’s leaving.

The second is when nobody will own it. Every automated thing rots, and something has to notice when it does. That something is a person with a name. If the best you can come up with is that somebody will pick it up eventually, the decision has already been made for you. This one kills more projects than technology ever does, and it kills them quietly, about eight months in, when the person who set it up has moved on and nobody remembers what the failure email means.

The third is when the thing in front of you is really a disagreement between people wearing a software costume. If two departments can’t agree who signs off on what, no product resolves that. Somebody with authority has to make the call, and then, possibly, software can hold everyone to it. Buy the tool first and what you’ve bought is an expensive record of the argument, updated daily.

Whichever way you land, write two things down

The first is ownership. Who holds the code, the accounts, the credentials and the documentation once the people who set this up have moved on. Answer it while everyone is friendly and nobody needs anything, because the day you need the answer is a day when somebody is annoyed.

The second is a single written sentence describing the circumstance under which this turns out to have been the wrong decision. Not a paragraph of caveats. One sentence, specific enough that you’d recognize the situation if you walked into it. “If we’re still fixing this by hand in three months, we bought the wrong thing.” If you can’t write that sentence, you haven’t made a decision yet. You’ve made a preference and dressed it up.

Chapter 09

Connecting AI to existing systems

The plumbing between two systems is the easy half. The hard half is deciding which one of them is allowed to be right.

Start by counting the systems your operation depends on. Scheduling is probably one of them. Billing is another. Then wherever customer records live, wherever the marketing tools sit, and the spreadsheet that one person keeps updated and that half the business leans on without anybody quite admitting it. You’ll probably get to five without trying. And nobody picked that number. They arrived one purchase at a time over years, and every one of them made sense on the day it was approved. None of those decisions had anything to do with the others.

Which is how a single customer’s phone number ends up living in three places, entered by three people weeks apart who never spoke to each other. Now you have three answers and no way to know which one is current. That’s where the cost of skipping integration actually sits, and because it never shows up as a line item, it never gets weighed against anything. You pay it in double entry, in delay, and in the mistakes that happen in the gaps between one tool and the next, which is also where a surprising amount of margin goes.

None of that is about AI yet, which is sort of the point. Anything clever you add later, an assistant that drafts, a model that reads documents, sits on top of this plumbing. Get the plumbing wrong and the clever part inherits the mess.

Measure it before you talk to anybody selling anything

This takes a week and costs nothing. Ask your people to keep a tally every time they retype information that a computer in the room already holds. Four columns is plenty. The person, the piece of information, where they read it from, and where they put it.

By Friday you have a ranked list of what to connect, in your own handwriting, with no vendor in the room. That list is worth more than any discovery call, because it’s about your operation rather than about what the person on the call happens to sell.

The connection may not be available at any price

Buyers get caught out by this one, so check it early rather than after the money has moved.

Software you already own and pay for may have no way in at all. Or it has one, but writing to it is only available on the tier above the one you’re on, so the quote you just approved is missing a subscription upgrade. Or there’s a connector and it can read but not write, which is fine for reporting and useless for anything that has to update a record. Or it can write, but slowly, with a limit on how many times an hour, which matters a lot if you were planning to sync a few thousand records a night.

And it’s somebody else’s software, so it changes when they decide it changes. You’re building against a moving target you have no say over.

None of that is a reason to give up. It’s a reason to ask two things in writing before anybody signs. Can this system be written to on the plan we’re paying for, and what happens to us the day the vendor changes how it works. Anybody who has done this before answers both in a sentence.

Then the actual hard part, which is ownership

Copying a record out of one place and into another has been a solved problem for twenty years. Deciding which of the two is allowed to be right about a given field hasn’t, and that’s where these projects live or die.

Give two tools equal authority over a customer’s phone number and they’ll overwrite one another forever, politely, every time a sync runs. Your staff will spend years correcting whichever copy they happened to open that morning, and every correction will be undone by the next sync, and everyone will conclude that the software is broken. The software is doing exactly what it was told.

So before anything gets connected, go through the fields that matter and name two things for each one. The system that owns it, and the direction information is allowed to move. Addresses are owned here and pushed there. Appointment times are owned over there and come back this way. Nothing gets to be two-way merely because two-way sounds more capable.

That document looks technical and it isn’t. What you’re really settling is which people are permitted to change which facts about a customer, and that’s an operating decision only you can take. Software does nothing more than enforce a choice you already made.

Failure design, because the other system will be down

Not might be. Will be. It will also time out halfway through a write, and somebody will click the button twice, and a request will land after the thing that was supposed to happen first.

A connection that fires off a request and assumes it arrived will lose information silently, for months, before anybody spots the pattern. And when somebody finally does, you can’t tell which records were affected, because nothing kept a record of what was attempted.

A queue is the version worth paying for. The message gets held, it gets retried on a schedule, and if it still can’t be delivered, something makes noise at a person. It’s more expensive to build. It’s also the only shape of this I’d put my name on. Ask any vendor what happens to a request when the receiving system is unavailable. The answer tells you which of the two you’re being sold.

Some actions can’t be taken back, and those get a gate

Publishing a listing. Submitting a claim. Sending money. Emailing a customer. Nearly everything else a computer does for you is either looking something up or preparing something, and preparation is forgiving, because you can throw a draft away and nobody outside your building ever knew. The small number of actions that leave your building aren’t forgiving, and they earn a different design from the rest.

One of my own projects, CardScope, is built to that standard, and I did it that way partly to learn first-hand what the careful version costs before telling anybody else to pay for it. It reads a trading card from a photograph and can put it up for sale on eBay. Getting from photograph to structured fields was an afternoon. The gate sitting in front of publishing took most of the build.

The shape of it is this. Nothing reaches eBay until a person has looked at a screen holding the precise values about to go across (the title, the asking price, the stated condition) and has clicked. And the values on that screen aren’t my own guesswork about what eBay accepts. They’re pulled live from eBay’s own category rules, so what a person approves is exactly what the far side receives. The model is also under instruction to leave a field empty rather than fill it with something believable, because a blank is obvious to whoever is reviewing the screen, and a confident invention isn’t.

Publishing also isn’t a single call. Under the hood it’s a short run of them, and the connection can die at any point in that run, leaving the job half done. So every step writes down where it got to, and if the whole thing has to run again it looks up what’s already sitting on eBay’s side before making anything. Double-click the button, lose your wifi mid-request, give up and try again the next morning, and the end state is a single listing. Not a duplicate, and not a partial one that reads as correct until a stranger pays for it.

All of that produces something which, from where a user stands, is indistinguishable from the naive version with a single button on it. You find out there was a difference on the day something breaks, and that’s the only day this kind of design ever gets judged.

The rule underneath all of that is the same one I’d apply to any system that touches the outside world. The machine drafts. A person releases anything that can’t be recalled. And the system checks what’s already there before it creates anything new.

The twenty-minute version

Write out each action your software is able to take that can’t be reversed afterwards. Then put a name beside each one, being whoever currently lays eyes on it first. If the answer is nobody, work out what it would cost you if that action fired twice on a bad Tuesday, and you’ll know immediately whether it needs a gate.

Chapter 10

Data privacy and security

What actually changes when your customer records start going through somebody else's software, and the three answers to get in writing before any of it touches real material.

You already hold information that would give you a very bad month if it turned up somewhere it shouldn’t. Case files, patient records, tax returns, payroll, whatever your version of it is. And you’ve held it for years without much drama, because the arrangements around it are old and boring and understood. There’s a system, there’s a set of logins, and somebody can tell you who has which one.

An AI tool disturbs those arrangements, mostly because of the way it arrives. Somebody finds a website on a Tuesday afternoon, pastes something in, gets a useful answer back in eight seconds, and no purchase order was raised and no review happened and nobody feels like they made a decision at all.

So the work here is to slow down at four questions before anything real goes through anything new. I’d take them in this order, because the first one is the cheapest to ask and the answer to it tells you the most about who you’re dealing with.

Where it goes, and how many companies that turns out to be

Your data leaves the building, which your email has been doing for years without anybody lying awake over it. The question worth your time is how many separate companies end up holding your material on the way to producing one answer.

Plenty of AI products are an interface sitting on top of a model that another company runs. Sometimes a third company does the storage and a fourth does the transcription. Each one of those is a place where your documents live, with its own staff, its own access rules, and its own bad week somewhere in the future.

So ask for the sub-processors by name, in writing. Not a paragraph about how seriously they take security. Names, and what each one does with what.

That request doubles as a competence test, which is really why it goes first. A vendor who has been asked this before sends the list the same afternoon, because it’s a document they already keep. A vendor who hasn’t will need three weeks, and the three weeks are spent finding out.

Whether it trains on what you put in

Ask it plainly, and ask it about the exact plan you’re on rather than the product in general.

One product’s free version and its paid version can carry completely different terms on exactly this point, and the free version is the one somebody in your office may already be using, with real names in it. So read the terms attached to the plan on your invoice, not the reassurance on the page sales sent over.

Processing and keeping are two different products

Retention is where this stops being abstract. There’s a real difference between a tool that reads your document, produces the output, hands it over and retains nothing, and a tool that files every request away so people can scroll back through last month’s work. Both of those are legitimate. Only one of them still has copies of your files in six months, sitting somewhere you’ve never looked, under access controls you didn’t write.

So work it as a set of five. What is the retention period, in days. Which country the copy sits in. Which roles inside the vendor can open it. Whether a deletion is an actual deletion or a flag on a row somewhere. And then the one almost nobody asks, which is what leaves with you the day you cancel and how you’d ever verify that it did. That last answer decides whether walking away is a clean break or a six-week project.

The gap between the agreement and the assumption

Pick the thing you’re most confident this vendor does with your data. Now go and find the sentence that says so in the contract you actually signed.

If what you’re relying on came out of a slide, a webinar, or somebody being helpful in a chat window, then it isn’t binding on anybody. It’s the memory of an impression, and it will not turn into a contractual obligation on the day you need one. That check takes twenty minutes and I’d run it early, because it tells you whether you’re evaluating a document or a feeling.

Your own logs are a retention policy nobody wrote down

The copies on your own side of the wire are as much of a problem as the ones on theirs, and they get talked about less.

Once a tool is wired into how you work, copies start piling up in your own systems. The transcript saved onto the customer record. The summary emailed to three people, which is now in three mailboxes and a backup. The request log the integration keeps, holding the full text of everything it sent, because that’s what request logs are for and nobody went back and turned it off after the thing went live.

That’s not the vendor’s retention policy. It’s yours, written by whoever configured the integration, on an afternoon when the goal was getting it working. So ask your own side the same questions you asked them. What does this write, where does it write it, and who can read it. Every awkward search-and-delete job later starts life as a default that looked harmless.

The tool nobody approved

All four questions above assume somebody is buying something. The likelier failure is that nobody buys anything. Someone’s behind, the thing is due at five, and a free summariser is one search away. That route skips every question in this chapter, and being annoyed about it afterwards changes nothing, because the pressure that caused it is still there next Thursday.

What helps is one page. Which tools are cleared for which kinds of material, which ones are cleared for nothing with a name in it, and who to ask when whatever’s in front of you isn’t covered. Then a person under time pressure has a fast answer that isn’t “paste it and hope.” Kept current, that single page protects you better than most of what you could win in a negotiation.

Where to start

You don’t need a lawyer to make a start on this, and you don’t need a policy document with numbered sections either. It needs three answers in writing before the tool touches anything real. Who else ends up holding your material. Whether anything gets trained on it. And how long a copy sticks around.

Whatever body you answer to, and whichever rules apply where you operate, what they expect from you sits on top of those three answers. And not one of the three gets easier to ask for after the first invoice has been paid.

Chapter 11

Evaluating AI vendors

One question to ask before any money moves, how to read the answer you get, and what a demo can and cannot tell you about the thing you're about to sign for.

Chapter 01 covers the gap between a demo and a Tuesday morning from the operating side, meaning why the gap exists and what it’s made of. This chapter is the same problem from the buying side of the table. You’re in a room, somebody is being paid to be persuasive and is doing it honestly, and you have to work out what you’re actually looking at.

Two things do most of the work here. A question, and a way of watching a demo. The question first, because it’s the cheaper of the two and you can ask it before anyone builds a proposal.

The question

“What result would bring you back in three months telling me to stop?”

Asking takes about four seconds. It’s the minute afterwards that’s difficult, because you then have to work out on the spot what you were just given, while staying polite, with an audience.

So ask it, and then don’t speak. The silence gets uncomfortable and somebody will move to fill it, and if you’re the one who fills it, whatever they had prepared stays unsaid. Count five before you say anything. The hard part of this is entirely the staying quiet.

Which direction the answer travels

Then listen for where the answer points. The common shape points forward and every branch of it is good news. Adoption. How the team’s getting on. A conversation at the end of the quarter. Play it back to yourself afterwards and look for any sentence in there that could one day show up as bad news. If you can’t find one, then what you were told is that this work carries on, warmly.

Take it as sincere, because it usually is. A person can be completely straight with you for an hour and still give you an answer that no future result could ever contradict. That’s the thing to listen for, and it isn’t dishonesty.

A genuine answer tends to be short. It gets concrete quickly, and it can land a bit flat in the room, because saying that sentence aloud commits them to a meeting nobody wants to sit in.

What a real gate looks like on a page

Auditing the model itself is out of reach for you, and it isn’t a good use of your time either. What you can inspect is whether a gate exists at all, and a real one shows up as three things on a page.

A number. A stated method for arriving at that number, including which system the data comes out of and whose job it is to pull it. And a deadline, a real dated one, printed in the document. “When we’ve seen enough to be confident” is not a deadline, and that kind of confidence never finishes arriving.

Which gives you a blunt test for the answer you got in the room. If it ran a whole paragraph and no number appeared anywhere in it, that wasn’t an answer. What you were handed is a way of narrating anything that happens next as forward motion.

The follow-up, which matters as much as the question

A written threshold can still come to nothing, for one structural reason. A vendor can have a number in the document, hold the pen that decides whether the number was met, and be paid identically either way. None of that requires anybody to act in bad faith. A gate is simply weaker when the person holding it has a preference about which way it swings.

So three more questions. Who records the verdict. On which date they record it. And what the agreement does if the figure comes in under.

Someone who has done this before answers all three in a sentence and doesn’t need to check with anybody. Someone who hasn’t will offer to come back to you, and the number of days that email takes to arrive tells you something on its own.

Two other moves are worth keeping in your pocket. Ask the engineer who’d be doing the work, rather than the person who turned up with the slides, because the smooth version of the answer often hasn’t reached the people building anything. And ask for it in writing, as an attachment to the proposal itself. An agreeable sentence costs nothing in a meeting. Typing one into a document that a contract points at is a different act, and people treat it differently.

A vendor with no answer at all isn’t automatically hiding something. They might simply not have a number yet. Then what your money is paying for is the work of finding one, and that’s a reasonable thing to pay for, provided the invoice describes it that way.

What a demo proves

It proves the product is real and that a trained person can operate it. Which is worth knowing, genuinely. It’s also the cheapest artefact in the whole evaluation to produce, and that’s the thought to keep hold of while you watch one.

What you’re being shown is a rehearsed happy path, on data chosen by the party who gets paid if you like it. There’s nothing wrong with that. Anybody would do the same. Three things are missing from that half hour, though, and they happen to be the three that decide whether you’re glad about this purchase next spring.

The first is how it behaves when something goes wrong. Something will, so that part isn’t in question. What matters is what the software does next. So ask to watch a live run on deliberately broken input, and bring the broken input with you. A record that’s in there twice, a required field left empty, a date written the other way round. Then watch whether the system stops and says it couldn’t handle this and why, or whether it hands over something confident and wrong and keeps going. Confident and wrong is worse than having no system at all, because from then on a person has to check everything it produces, forever.

The second is your own data. Theirs is clean because somebody cleaned it first. Yours has a customer name parked in an address field, the same supplier spelled four different ways, and records with a mandatory field left blank, because the person who typed them in left the company in 2023. So hand over a real extract, unedited, in whatever state it’s in, and ask for the same demonstration run against yours.

The third is month six. In month one everybody’s keen and there’s an engineer from the vendor sitting in your chat channel. By week ten that same engineer is setting up somebody else’s account. So ask who answers your questions in month six, how many days a typical fix takes, and what the plan is when one of your own systems quietly renames a field, which it will.

Put the error path at the top of your list, above the features. If nobody can show you one on the day, that doesn’t prove there isn’t a good one in there. It does mean you’re taking the costly half of the product on faith, and the costly half is the part your team ends up living inside.

A pilot doesn’t settle this on its own

The natural move when a demo leaves you unsure is to ask for a pilot, which is the right instinct. It only helps if the pilot is built so that it can come out as a no. Otherwise what you’ve bought is a demo that lasted three months, and by the final week everybody involved is tired and committed and the purchase happens on momentum. Chapter 12 is how to set one up so that can’t happen to you.

Chapter 12

How to run your first AI pilot

Four things have to be written down before a pilot starts, and four conditions have to hold while it runs, or what you've bought is a purchase with a longer approval process.

If a pilot can’t come back as a no, it isn’t really a pilot. Twelve weeks go by, people are broadly pleased, the contract gets signed, and hardly any of that turned on whether the software did the job.

It’s an easy trap to walk into, because a pilot feels like the careful option. It is the careful option, but only if it’s built so that “no, we’re not doing this” is a live possibility on the last day. That takes four things written down before it starts, and four conditions holding while it runs.

The number, and it’s a number about your business

Before anybody configures anything, one sentence has to exist in writing. The figure this thing has to reach. How that number gets produced. The date on which somebody checks it. And the name of the person doing the checking. Four parts, and with any one of them missing what you’re holding is an intention.

The number has to be about your operation rather than about the software, and the software number is the one you’ll be offered. “95% accuracy” is a software number. 95% of what, counted across which cases, chosen by whom? You can lose most of a morning to that one question, and the morning is worth losing.

What you want instead is the thing that was irritating you before anybody mentioned software. Invoices posted without anybody touching them. Days between a request arriving and an answer going out. Calls that reach a person. Pick something you’d have noticed anyway, because that’s the thing whose value you don’t have to argue about later.

The way of measuring has to already work today

This is the part that gets left out, and leaving it out is what quietly wrecks the review at the end.

Say you’ve settled on days from request to answer. Can somebody hand you last month’s version of that figure this afternoon, measured off the way the work happens today? If not, week twelve arrives with nothing to compare the pilot to, so it gets settled on impressions instead, and those belong to whoever ran the thing.

Two things to nail down here. Where the figure comes from, and who pulls it. If the new system is also the thing producing its own report card, you’ve built something that grades itself, and it will grade itself kindly. Write down today’s value as well, before anything changes. A few months in, you’ll find people genuinely disagree about where it started, and memory of the old way runs better or worse than the reality depending on who’s in the room.

A date, and a name

The date has to be an actual date. Not “once we’re confident we’ve seen enough,” because that day keeps moving another month out.

Then a name next to the date. One person, whose job on that day is to record whether the number was reached. And if the person recording it is also the person who built it, or the person who pushed to buy it, you can predict the review right now, and so can they.

What happens if it misses

Then write the half of the sentence that nobody writes. What happens when the figure lands under. Does the project stop there? Is there one extension, and exactly how long is it? Does the agreement simply end on that date?

Decide it now, while nobody is invested. On the day the number actually lands low, every option in front of you will look political, and the conversation will be about the people in the room rather than the result.

Then run it so the number can bite

Four conditions, and none of them are technical.

Run it against real work. Your busiest week, your most difficult customers, the inputs that turn up in a format nobody planned for. Feed a pilot a selection of tidy cases and the only thing you learn is how it copes with tidy cases.

Keep the current way running alongside. Yes, that’s annoying, and yes, you’re paying for two ways of working at once for a few weeks. It’s also the only thing that converts “people seemed happier” into a comparison with numbers beside it. And it makes switching off at the end into a choice instead of an outage.

Keep the cost of switching off low. There’s a week in most pilots where the thing stops being an experiment and starts holding weight, and after that, pulling it out breaks something on a Thursday and untangling it is a piece of work in its own right. Past that week the review date is decoration, because the cost of answering honestly quietly went up and nobody noticed it happening.

And don’t re-grade it every week. Do that and you’ll land on a good couple of weeks and start treating those as the answer. Look at the whole period, on the date you agreed, and not before.

A note on the writing-down, from someone who has had to do it

I hold myself to the same rule on a system that earns nothing and has exactly one user, which is where you find out whether a rule of yours is real at all.

I’ve got a prediction system of my own. One of the models inside it spent months producing forecasts that reached nothing at all. Same schedule as everything else, same live inputs, each forecast written into a table nothing was allowed to read, with the actual result filled in beside it once it was known. Running something that way, where its output can’t affect anything, is called shadow mode, and it costs almost nothing to set up. 125 of those accumulated. The model landed at 52.1%, short of a figure I’d committed to in writing long before it produced its first result, so it never got switched on. It’s still off.

The deciding part took seconds. The difficulty had all been paid for earlier, back when I set the figure with no real idea where the thing would eventually land. That’s the whole mechanism, and it’s the reason a threshold written after you’ve seen the result is worth nothing. Once you know the number, you can find a reading of it that lets the work continue, and you can do it completely sincerely.

There’s a longer version of that with the mistakes included, if you want it: The Threshold I Set Before I Knew the Answer.

Ask it out loud first

Before your pilot starts, say the question in the room. What result, measured on which date, would make us stop doing this?

If nobody can answer, the decision has already been made, and the twelve weeks in front of you are paperwork.

Chapter 13

When not to use AI

The most useful thing an assessment can produce is a decision not to build, which is also the one outcome nobody has an incentive to sell you.

Twelve chapters of how to do this well, and the last one is about not doing it. Which isn’t a contradiction. The decision not to build is where a lot of the money in this actually sits.

Every other outcome of an evaluation leads somewhere for the person recommending it. Build this leads to a build. Buy that leads to a purchase. “Leave this alone” leads nowhere at all, which is exactly why nobody is in a hurry to hand it to you, and why it can be worth more to you than either of the others.

Work out what a bad build actually costs you

The quote is the smallest part of it. Start there anyway, since the money is the part everybody counts, and then keep going.

Add the time your operations lead loses to a project instead of running the place, priced at whatever an hour of their attention is worth to the business. That figure surprises people.

Add the process you reshaped to suit the tool. That shape outlasts the tool by years, because restoring a process to the way it used to work is a job nobody ever gets round to.

Add the year that follows, in which the project that would actually have paid off doesn’t happen, because the money went into this one and so did everybody’s appetite for another go.

And then the item that never shows up on any page of any proposal. What it takes out of you to stand in front of everybody you persuaded and tell them it’s being switched off. That one doesn’t get settled in money, and it gets settled over months.

Put even a rough figure against each of those and the total you get is several times the quote. What the multiplier is for your operation, I have no way of knowing. It’s larger than one, though, which means the figure everybody in the room is haggling over is not the figure being decided.

The answers that should end it before a vendor call

Chapter 03 has the long version of this, eight questions asked in a particular order for a particular reason, and I’m not going to run through them again here. What’s worth pulling into one place are the answers that end a candidate outright, because they’ve been scattered across this guide and they’re cheap to check.

Any one of these should close the file before anybody books a demo.

This workflow is about to change anyway. A reorganization, a new practice management system next spring, a rule change you can already see coming, or the person whose job this is leaving in June. Automating something with a few months left in its current shape means paying for it twice.

Nobody owns it. Or three people each own a slice of it, and each of them describes the steps differently when you ask them separately. Then nobody can tell you what a correct output is, nobody receives the exceptions, and nobody notices when the automation and the work stop matching. This one is fixable, and it has to be fixed first rather than during.

The judgment can’t be written down. Not “hasn’t been written down yet,” which is normal and repairable. This is the case where the person doing the work genuinely can’t reduce the decision to criteria, even sitting down with you and trying. If nobody can state what right looks like, nobody can check the output cheaply, and all you’ve done is move the labour from producing the work to reviewing it.

It’s a process problem in a software costume. The task only exists because two systems don’t exchange data, or a form doesn’t ask for something it needs, or there’s no index of anything, so a person answers the same question all week. Automate around that and the underlying defect becomes permanent and comfortable, and comfortable is worse than irritating, because irritating is what eventually gets things fixed. The real project is upstream and it’s usually smaller.

The approval is the same person twice. Where the automation does something irreversible with money or with a customer, somebody has to release the action. If the only person available to release it is the person the automation is acting for, that isn’t an approval, it’s a formality with a button. The verdict stays no until a second pair of hands exists.

Judge a candidate on its worst answer rather than its overall feel. One of these settles the question no matter how well the rest of it went, and any checklist you can add up will find a way to say yes.

Spend an afternoon trying to disqualify your favourite candidate

Take the workflow you’re most excited about, the one you already picture working, and spend an afternoon arguing against it. Go looking for the five answers above rather than hoping you don’t find one. Ask the person who does the work whether they could hand you written instructions, and mean it.

That afternoon is the cheapest money in the entire project, and it’s the only part of any of this you can do without hiring anybody. A conversation is by far the cheapest place to lose a candidate. The most expensive place is month seven.

The answer I’d expect to give most often is “not yet”

Not “no.” Not yet, with a short list attached: one named owner, the procedure written down, and six weeks of actually running it that way before anybody looks at software again.

That’s an unsatisfying thing to be told when you were hoping to start. It’s also how the eventual build gets cheaper, because a good share of what makes an automation expensive is that nobody had defined the work before somebody went and automated it.

If you want a second opinion on which of these applies to your operation, that's the assessment.