A quote, a client list, a payroll table - every day someone on the team drops something like that into an AI tool. Three places that data can end up, and nine questions worth asking a supplier before an incident rather than after one.
A scene now playing out in every other company: a colleague has to write a quote, so they paste a price list, last year's contract and a contact list into an AI chat to give the model some context. Thirty seconds later they have the text. Nobody asked where those three attachments went.
It isn't carelessness. It's the normal response to a tool that works. The question “where does this go” simply has nowhere to be asked.
The question isn't “AI or no AI”
Banning AI in a company doesn't work today. People will route around it using a personal account and you won't know - which is a worse state than the one you started from.
The useful question is different: which of three categories does the tool we're using fall into?
1. A public service on a free account. Sign up with an email, accept the terms, off you go. On free tiers it's common for the provider to be permitted to use inputs to improve the service. It isn't a scandal - it's written in the terms nobody reads.
2. A paid tool with no data processing agreement. You're paying, but your company hasn't signed anything. You don't know for certain where the data physically sits, how long it's kept, or who has access.
3. A platform with a data processing agreement. There's a signed document stating where the data is, for how long, who processes it and what may not be done with it.
The difference between the third and the first two isn't technology. It's that with the third you have something to point at.
Three places your data can end up
What makes the difference isn't model quality, it's paperwork. Either you know where the data sits and for how long, or you're assuming.
Nine questions for a supplier
This is a list worth working through for every tool company data flows into. The answers should be findable in documentation, not promised on a phone call.
About the data
- Are our inputs used to train models?
- How long is data retained after it's deleted from the interface?
- Can we export all our data at any time, and in what format?
About location
- Where do the servers physically run?
- Is data transferred outside the EU, and if so on what basis?
- Who are the subprocessors - which other companies take part in processing?
About contracts and access
- Is there a data processing agreement, and is it available without lengthy negotiation?
- Who on the supplier's side can see our data, and under what circumstances?
- How will we hear about a security incident?
Question one is the most useful
The answer is usually findable in the terms of service within two minutes, and it decides whether the tool is suitable for anything sensitive at all. If a supplier hasn't written it down clearly, that's an answer in itself.
What to do about it in practice
You don't need a ten-page policy. Three simple rules the team can remember will do:
Sort data into three categories. Public (marketing copy, the price list on your website), internal (quotes, processes) and sensitive (personal data, payroll, health information, trade secrets). For each category, say which tool may be used.
Name one approved tool. The most common cause of a leak isn't malice, it's that everyone found their own. When a team has one clear option that works, they stop looking.
Personal data gets its own regime. Names, contacts, ID numbers, clients' health information - only a tool you have a signed agreement with and know the location of.
Check the legal position
Rules for processing personal data and for deploying AI systems are being refined continuously in the EU, and they differ according to what you use the system for. Before you adopt internal rules, check the current position with your lawyer or data protection officer - this article is practical guidance, not legal advice.
Why none of this is a reason to avoid AI
Work through the nine questions and you notice something surprising: most of the risk has nothing to do with AI. It has to do with sending company data to an outside service - which you also do with cloud storage and email, you've just got used to it.
The only difference is that the AI category is young and plenty of suppliers are still building what's standard elsewhere. That's why the questions are worth asking. And “yes, we've sorted that, here's the agreement” is an answer that exists today.
How Apexloop handles backups, availability and responsibility for data is covered in Backups and data security. Who on the team sees which data is covered in Access rights and data visibility. Our data processing agreement is publicly available on the Data processing page.
Common questions about company data and AI
Is it safe to paste a client quote into AI?
It depends on the tool. With a service that has a data processing agreement and doesn't train on inputs, it's ordinary practice. On a free public account it means the contents of that quote leave your control - and if it contains contacts or prices, it's sensible not to.
How do I tell whether they train on our data?
Look in the terms for “training” or “improve our services”. Business tiers usually exclude training on customer data explicitly; free accounts often permit it.
Is it enough that the AI runs in the EU?
It's an important part of the answer, but not all of it. Alongside server location, look at retention, subprocessors and training - a service running in the EU can have unsatisfactory settings on all three.