The model is only one part of an AI solution
One of the biggest misconceptions around AI implementation is that choosing the right model is the main technical decision.
Choosing the model matters. Increasingly, though, the harder decisions sit around it: giving it the right information, controlling what it can do, checking its output and connecting it properly into the wider process.
You'll hear a few terms around this.
Context is the information given to the model so it can complete a task. That might be the contents of an email, customer information, a policy document, previous actions, or data retrieved from another system.
Guardrails are the rules and checks around the AI. They might limit what information it can access, constrain the actions it can take, validate the format of an output, or send certain decisions to a person for approval.
Orchestration is the plumbing. It's the logic deciding which system is called, which model is used, what information passes between steps and what happens when something goes wrong.
Evaluation is having a proper way of deciding whether the thing actually works.
That last one is easy to overlook. An AI demo producing an impressive answer is not the same as a reliable business process. You need to know what a good result looks like, how often you're getting it, what kinds of mistakes happen and what the system does when it's uncertain.
In a lot of real systems, those surrounding pieces matter at least as much as the model.
Start with the thing you want to improve
Instead of "we need to start using AI", try something like:
Someone spends six hours every week taking information out of PDFs and putting it into our system.
We receive hundreds of enquiries and somebody has to work out manually which team each one goes to.
Producing our weekly management report means collecting information from four different systems and a spreadsheet.
Those are problems you can work with. Once you have one, try to pin down three things.
What's the starting point? What information, event or action begins the process?
What should exist or have happened when it finishes?
And how will you know whether it worked? That might be time saved, accuracy, response time, manual steps removed or something else that actually matters to you.
Now you've got something you can work backwards from. You can ask what technology is genuinely required between the input and the result. Sometimes that's a good prompt in a chat tool. Sometimes it's several AI calls, APIs and a bespoke application. Sometimes it barely involves AI at all.
Your data may be the real starting point
This is where a lot of otherwise sensible AI ideas hit reality.
A business identifies a genuinely useful case and then discovers the information needed to make it work is spread across several systems, stored inconsistently, or simply not available in a usable format. At that point the clever AI bit isn't the first problem, and the first piece of work is connecting systems, cleaning up records or getting structured data out of documents. That's the subject of an earlier article in this series, so I won't repeat it here beyond saying it comes first more often than people expect.
You don't need to automate everything
There's a temptation to think a successful AI workflow has to remove the person from the process completely. It doesn't.
In plenty of situations, automating 90% of something and leaving the last 10% to a person is a much better outcome than spending significantly more time and money engineering every edge case out of the system.
Imagine a workflow that reads an incoming document, extracts the information, checks it against your internal system and prepares the next action. A person then spends twenty seconds reviewing the result and clicking approve. If the old process took somebody ten minutes, that's already a very successful implementation.
Trying to eliminate that last review step might require much more engineering, increase the risk and deliver very little in return.
The right balance depends on the process. Something low risk and easily reversible may suit a high degree of automation. Something involving money, sensitive data or an important decision may benefit from deliberate human oversight. The goal is a better process. Sometimes that means full automation, and quite often it doesn't.
Start small enough that you can learn something
Once you've identified a sensible problem, it's usually better to prove the workflow before building the ultimate version of it.
Take some real examples, build the smallest thing that can process them, and see where it works and where it falls over. Then start asking the useful questions. How accurate is it? What inputs cause problems? What happens when information is missing? Can the workflow tell when it's uncertain? How much does each run cost? Does a smaller model work just as well? Where does a person need to be involved?
This is also where model choice gets much easier. Rather than arguing in the abstract about which model is best, you can measure which one performs best on your actual task.
Often the answer is a mixture. One model may be excellent at extracting information from a particular type of document, another better at reasoning over it afterwards, and a conventional piece of code may still be the safest way to validate the final result. That's orchestration in practice: the right tool for each part of the job.
Know when to stop
This is one of the less glamorous but more important parts of using AI well.
Sometimes it doesn't work, or more accurately, sometimes it doesn't work reliably enough for the problem you're trying to solve.
You might get to 80% or 90% accuracy and find that getting significantly closer to perfect would cost far more than the additional value it creates. You might find the remaining cases genuinely need human judgement. You might discover a conventional software rule solves the problem more reliably. Or the models might just not be good enough for that particular use case yet.
None of those are failures. They're the result of having a measure in the first place, which is the whole point of starting small. A project that stops at 85% because you know the last 15% costs more than it's worth is a well-run project.
Some things are still better done by people. Some are better solved with traditional software. Some processes work extremely well with AI handling the repetitive part and a person dealing with the exceptions. And some good ideas are simply worth writing down and revisiting in six months, because the technology will have moved.
A sensible approach to AI should be able to say "not yet" as comfortably as it says "yes".
By Josh Gosselin, Founder, DEXM