Data security and GDPR

The question we hear most before a contract is signed is “where will our data go”. We answer it with the flow and the role split, not with an assurance that we are compliant.

Where the data goes

In the most common arrangement — a model provided as a service — “we implemented AI” means, in practice, that a fragment of your data leaves your system, reaches a model, and comes back as an answer. It is worth knowing by which route, because each of these four steps is a place where a decision gets made:

  1. Your system. A document, a ticket or a database record — data you already hold and already process on your own lawful basis.
  2. The intermediate layer we build. This is where it is decided what gets sent onward at all: which fields, in what detail, whether personal data is stripped or replaced, and what is written to logs. It is the point where minimisation can be enforced directly, and therefore the most important one.
  3. The model provider. Receives whatever the intermediate layer sent and returns an answer. Its region, its contract terms, and whether customer content is used for training are parameters of the project — see below.
  4. Back into your system. The answer lands where it should, together with enough of a record to reconstruct later why the system answered as it did.

With a model running in your own infrastructure there is no step three — the data never leaves your environment. Steps two and four remain, because limiting what is sent and being able to reconstruct a decision matter wherever the model runs.

The most effective safeguard is usually not the choice of provider but step two: data that was never sent needs no guarantees. That is why the set of fields leaving your system is agreed during the audit phase described in ourdelivery methodology, rather than during testing.

Who is controller and who is processor

This split comes from the GDPR itself, not from the contract — the contract only reflects it. It is confused regularly, and it has direct consequences:

  • You are the controller. Your data, your purposes. The lawful basis, the duty to inform the people concerned, and the record of processing activities all sit with you.
  • We are a processor — for the data that passes through the system we build. We act on your documented instructions and within the agreed scope, we do not determine the purposes of processing, and we do not use your data for our own. Separately, we process the contact details of people on your side, correspondence and accounting records as an independent controller; that is not processing on your behalf and rests on its own lawful basis.
  • The model provider is normally a sub-processor. That requires your authorisation for sub-processing, and it is one reason choosing a provider is not purely a technical decision.

Scope, deletion deadlines and the list of parties you entrust with processing belong in the data processing agreement for the specific project. We deliberately do not publish them here as a list: it would be out of date at the first project with a different profile, and what binds is what is in the agreement, not what is on a website. Processing carried out by this site itself — analytics and the contact form — is covered separately in theprivacy policy.

Decisions we take together at the start

The following four are settled before any building, because each is expensive to change afterwards:

  • Processing region. Whether the service runs in a European region, or whether you accept transfer outside the EEA and on what basis.
  • Use of content for training. The business tiers of the major providers switch this off as standard; consumer tiers are sometimes the reverse. It always has to be checked in the terms of the specific service rather than assumed.
  • What actually leaves your system. Which fields are necessary, what can be omitted, and what can be replaced by an identifier. Usually more can be removed than the first draft of a design assumes.
  • Your own model or a service. A model in your own infrastructure removes the transfer, but costs more to run and typically means a weaker model. It earns its place against a hard requirement, not as default caution.

The AI Act and GDPR are not the same thing

Both have to be satisfied in parallel and neither replaces the other. The difference comes down to what they govern:

  • GDPR governs personal data, whatever the technology. It applies the same in a spreadsheet as in an AI system.
  • The AI Act governs AI systems, including ones that process no personal data at all — a model predicting machine failures, for example.

Where they meet can mislead. The duty to tell a person they are talking to an AI system comes from the AI Act and applies whether or not personal data is mentioned in the conversation. Automated decision-making about a person — where it produces legal effects or similarly significantly affects them — is conversely GDPR territory, and applies whether the decision comes from a model or from a rule written twenty years ago. The technical side of preparing for the AI Act is covered in ourAI Act readiness audit; the legal assessment stays with your lawyer.

What to ask any supplier

Not just us. If a supplier cannot answer these specifically and in writing, that is information in itself:

  • Which fields exactly leave our system, and at what point?
  • What region does the service run in, and who are the sub-processors?
  • Can our content be used to train a model, and where is that written down?
  • What is logged, for how long, and who can read it?
  • What happens to the data and to the system if we end the engagement?
  • Who is answerable when the model produces an answer that causes harm?

We put the same questions when selecting a solution as part of anAI implementation. The answers come from the provider’s contract terms and documentation, and what we establish goes into the project documentation so it can be revisited.

Frequently asked questions

Will our data be used to train the model?

It should not be, and that is a matter of which provider you choose and what the contract says rather than of anyone’s assurance. The business tiers of the major model providers switch off training on customer content by default, while consumer tiers are sometimes the reverse. Which tier applies to your project is settled before we start and written down — and if the requirement is absolute, we look at running a model inside your own infrastructure.

Does data leave the European Union?

That depends on the provider and the region the service runs in. The major model providers now offer processing in European regions, but it is not the default setting in all of them. We treat it as a decision to be taken deliberately at the start, because changing region after an implementation usually amounts to a migration.

Who is the controller — us or you?

You remain the controller: they are your data and your purposes. For the data that passes through the system we build we act as a processor, meaning we work on your documented instructions and within the scope we agreed with you. The model provider is normally a sub-processor. The split matters in practice, because it is the controller who is answerable for the lawful basis and for informing the people whose data it is.

Does the AI Act replace GDPR?

No — they are two independent regimes that have to be satisfied in parallel. GDPR governs personal data whatever the technology; the AI Act governs AI systems, including ones that process no personal data at all. Complying with one does not discharge the other, and some duties look similar — telling a user they are talking to an AI, for instance — while resting on a different basis and reaching different things.

Can we implement AI without sending anything outside?

Yes, with a model running in your own infrastructure or a private cloud. It costs more to run and usually means a smaller model than the largest commercially available ones, so it earns its place where the requirement is hard — data under professional secrecy, for example. For most back-office processes, limiting what reaches the model at all turns out to be the better move.

What happens to the data when the project ends?

The principle is that data entrusted for processing returns to you or is deleted once the work ends. That does not cover records we have to retain on other grounds — accounting, or defending against claims — whose scope and retention period follow from law rather than from our choice. The specific deadlines and the way deletion is confirmed are governed by the data processing agreement, because that is what is enforceable, not a statement on a website.

This page describes a technical and organisational approach and is not legal advice. What binds a given project is the contract and the data processing agreement.

Ask about a specific scenario

Describe the process you want AI to support and we will tell you what data would have to leave it — and how much of that can be stripped out.