First, understand roughly what you're working with

You don't need to become an AI researcher before you can find useful applications for it, but understanding a few of the basics makes it much easier to make sensible decisions.

When people talk about a model, they generally mean the underlying system doing the work. ChatGPT, Claude and similar products are applications that give you access to one or more models, usually alongside other functionality such as file handling, memory, web access or tools.

Different models behave differently. Some are particularly strong at coding, some at working with documents or images, some at more involved reasoning, and others are optimised for speed or cost. There often isn't a universally best one.

We've had projects where one model performed noticeably better for a particular task while another was a better fit somewhere else in the same workflow. The sensible way to choose is to take representative examples of the real task, test a few options and compare the results.

Cost varies too. A highly capable model may be worth paying for when a task genuinely needs the reasoning, but there's little point using it for every simple classification or extraction job if a smaller one can do the same thing reliably.

In a real system you might use a cheaper model for straightforward processing and only bring in a more capable one when the workflow needs it. Once something is running hundreds or thousands of times, those decisions start to matter.

Then work out where your data is going

This is one of the first things a business should think about.

It's very easy to open a consumer AI chat, paste something in and start experimenting. Once you move beyond harmless examples, though, you need to understand what happens to the information you provide.

There are really three separate questions.

Is the data used to train the model?

Training is how models learn patterns from large amounts of information. It doesn't simply mean a model stores your document and can later retrieve it word for word, but it's still an important distinction if you're dealing with confidential, commercially sensitive or personal information.

The consumer and business versions of the same product can work very differently here.

Data sent to the OpenAI API isn't used to train its models unless you explicitly opt in, and the same applies to its business products. Anthropic takes a similar position on its commercial products, including Claude for Work and the Anthropic API, which aren't used for training by default.

Consumer accounts are different. In August 2025, Anthropic announced changes to its consumer privacy terms giving users of Claude Free, Pro and Max a choice over whether their chats and coding sessions can be used to improve its models. If users opt in, eligible data may be retained for up to five years; if they opt out, Anthropic says its standard 30-day retention period continues to apply, subject to exceptions such as safety, legal and feedback-related retention.

It's the kind of choice that's easy to click past, which is the point worth knowing: someone in your team on a personal account may well have agreed to something they don't remember agreeing to. Don't assume a public chat interface behaves like a business API.

How long is the data kept?

Training and retention are separate questions.

Providers may keep prompts and outputs for operational, security or abuse-monitoring purposes. OpenAI currently states that API inputs and outputs are removed after 30 days unless it's legally required to keep them. Zero data retention is available, but it needs a qualifying use case and prior approval, and it only covers eligible endpoints, so it isn't something you can simply switch on.

Retention should be part of the design decision, particularly where a workflow handles personal or sensitive information.

Where is the data processed and stored?

This is a different question again, and for some organisations it matters a lot.

Data residency describes where data is stored at rest. Inference residency describes where the actual model processing happens. They aren't necessarily the same thing.

OpenAI offers data residency in a number of regions, including the UK. For ChatGPT, in-region GPU inference is currently available to eligible Enterprise and Edu customers in the US, Europe and the UAE, and it requires data residency to be enabled in the same region first. The UK is not currently an inference-residency region.

So "our data is stored in the UK" shouldn't be read as "every part of the processing happens in the UK". Even where inference residency is available, some non-GPU processing, such as authentication, routing and certain service operations, may still take place elsewhere.

If geographic processing is a hard requirement, the model and hosting architecture need to be chosen around it from the beginning.

One option is to run models through cloud platforms that give you more explicit regional deployment controls but this comes at a cost. Claude is available through Amazon Bedrock, for instance, which offers In-Region, Geo Cross-Region and Global Cross-Region inference options. Exactly where processing occurs depends on the model, endpoint and AWS region involved, so regional availability needs to be checked for the specific deployment.

The same principle applies to Azure-hosted model platforms. What matters is exactly which service, model, deployment type and region are involved. For example, Microsoft's Global and DataZone deployment types have different processing-location guarantees.

The broader point is that "they don't train on our data" is only one part of the privacy story. You also need to understand retention, processing location, access controls, what information is being sent to the model in the first place, and whether some of it should be excluded altogether.

A good AI implementation starts with understanding the data flow: where information comes from, where it goes, what touches it along the way, and what genuinely needs to be sent to a model at all.

None of this needs to become anyone's specialist subject, but the moment your team is working with anything other than test data, training, retention and residency stop being abstract concerns. Someone will paste a client document into a personal account at some point, usually with good intentions and a deadline.

A short set of AI guidelines is usually enough to prevent that. Which tools are approved, what can and can't go into them, when to check before using something new, and who to ask. It doesn't need to be a policy nobody reads. It needs to be clear enough that people know where the line is without having to guess, and understood well enough that they know why it's there.

Provider policies were checked in September 2026. They change often, so check the current terms before relying on them.

So, once you've got a handle on the basics, where do you actually start? That's what we look at in Part two: finding something worth building.

By Josh Gosselin, Founder, DEXM