AI Data Ownership: What Regulated Companies … | The AI Briefing
The AI Briefing Episode 37 July 8, 2026 · 5:56

AI Data Ownership: What Regulated Companies Must Know Before Uploading Data

0:00 / 0:00

What you'll learn

  • The moment you upload a spreadsheet of company or customer data into a public model like ChatGPT or Claude for a quick summary, without the right agreements in place, you can breach confidentiality clauses across many customer contracts and no longer own the data.

  • Once data leaves your environment it goes to the AI vendor, who may train their models off it. Vendors need more data to make money, so this is by design, not an accident.

  • How much protection you get depends entirely on your agreement with the vendor: different contracts give different levels of confidentiality and different commitments not to train on your data.

  • A three-step response: first read the contract properly; second, disable the vendor's train-on-your-data switch where one exists; third, use a framework that keeps the model under your direct control.

  • Constrained environments like Microsoft Foundry, AWS Bedrock, GCP Vertex and Databricks let you leverage external models without them training on your data, and you can optionally fine-tune on your own data as a process you own rather than value handed to a vendor for free.

By the end of this episode you should be able to explain the data-ownership and confidentiality risks of uploading regulated data to public AI models, and choose between reading vendor contracts, disabling training, and running models in a controlled environment like Bedrock, Foundry, Vertex or Databricks.

In this episode

  1. AI, confidentiality and who owns the data
  2. How uploading data breaches your contracts
  3. Why vendors want your data
  4. Step one and two: read the contract, switch off training
  5. Step three: keep the model under your control
  6. Foundry, Bedrock, Vertex and Databricks

RegTech expert Tom reveals critical risks of using AI tools in regulated environments. Learn why uploading company data to ChatGPT or Claude could breach confidentiality agreements and what solutions exist for FinTech and HealthTech companies.

AI Data Ownership in Regulated Environments

Key Topics Covered

The Data Ownership Problem

  • Why uploading company data to consumer AI tools is risky

  • How confidentiality agreements and customer contracts are impacted

  • What happens to your data when you use AI vendors

  • The model training issue: vendors using your data to improve their products

Three Solutions for Safe AI Use

1. Read Your Contracts Carefully

  • Understanding vendor terms and conditions

  • Identifying data ownership clauses

  • Recognizing training rights in agreements

2. Disable Data Training Features

  • Finding the opt-out switches in AI platforms

  • Limitations of relying on vendor settings

  • Internal compliance challenges

3. Use Enterprise-Grade Solutions

  • Microsoft Foundry

  • AWS Bedrock

  • GCP Vertex

  • Databricks

  • Benefits of constrained environments

  • Maintaining control over model training

Regulated Industries Affected

  • FinTech

  • HealthTech

  • Any organization with confidentiality agreements

  • Companies subject to data protection regulations

Action Items

  • Audit current AI tool usage in your organization

  • Review vendor agreements for data ownership clauses

  • Establish AI usage policies and procedures

  • Evaluate enterprise AI platforms for your needs

  • Train employees on safe AI practices

Host

Tom - RegTech specialist focusing on AI and digital transformation in regulated environments

Chapters

  • 0:02 - Introduction: AI in Regulated Environments

  • 0:48 - The Data Ownership Problem

  • 1:47 - Why AI Vendors Train on Your Data

  • 2:18 - Solution 1: Read Your Contracts

  • 2:36 - Solution 2: Disable Training Features

  • 3:25 - Solution 3: Enterprise AI Platforms

  • 4:53 - Final Recommendations and Action Items

Subscribe to our newsletter: https://newsletter.concepttocloud.com/

Want to apply AI to your engineering workflows? We build production ML pipelines, not demos.

Explore AI Services

Transcript

So today, uh, I do a lot of work in the reg tech field and, uh, you know, helping and advising companies when it comes to using AI or digital transformation inside of, um, you know, regulated environments, Fintech, Healthtech, all those types of things. Um, and so one thing that I wanted to be able to discuss with you today just briefly is the use of AI within side of your organization. Because, you know, there is a lot that we have to think about when it comes to leveraging AI on a regular basis, especially when it comes to confidentiality agreements and who owns the data. Now, of course, you know, when it comes to leveraging AI models, everyone would like to be able to think [chuckles] that they can just chuck a spreadsheet full of, uh, company data into your favorite ChatGPT, Claude, whatever, and, you know, ask for a quick summary. But as soon as you have done that, if you, uh, have...

don't have the right agreements in place, you've suddenly breached a whole bunch of laws and confidentiality clauses in potentially, you know, many different customer contracts because you've uploaded the data and you no longer own the data. It has gone to, you know, your favorite AI vendor, and at that point, you basically own a copy of it because your AI vendor now is gonna train, you know, their stuff from your model. This is not new. Of course, for these vendors to be able to make their money, they need more data. And so depending on your, uh, agreements with the, uh, AI vendor, you will get different levels of confidentiality and, uh, their commitment to not train off your data, um, yeah, as an organization.

So what can you do about this? Of course, there are many different things you can do about this, but you have to make sure that you have the right procedures and policies in place. One thing is, you know, when you're dealing with the, the bigger vendors, is making sure that you read the contract properly, um, because, you know, step number one is read the contract. Step number two can be like, make sure that if you're using a vendor but you don't have a sort of global team setup, is to make sure that if there is a switch to be able to switch off, uh, vendors training off your data, switch it off. Um, because, you know, last time I checked, that was a pretty good way of asking them to not do it, was by disabling it.

Um, but of course, again, you're still beholden to a vendor. That vendor may or may not train off your data, but also people inside your organization may not flip the switch, and then you're still at risk of, you know, capitulating to, um, you know, the laws and regulations. So the third option, um, is using a framework or a function that allows for you to either host a model or leverage a model that is un- more under your direct control. So when it comes to, you know, internal use of LLMs, it may not be as simple as signing up to, uh, anthropic. com or, you know, ChatGPT and opening an account, because what you might wanna be able to do is leverage those models but in a more constrained environment.

So we're talking about, um, you know, Microsoft Foundry, AWS Bedrock, um, GCP Vortex, uh, and, you know, Databricks, for example. If you're already a Databricks user, you can leverage those models inside of Databricks, and they're not gonna train data off of your... Not gonna train the model off of your data because they bring in the models from external providers. Of course, going back to what we were saying the other day, you can, of course, if you would so choose, then train models off of your data so that it can become more intelligent and answer more pertinent questions. But that becomes something that you own and a process that you manage, and it is not, uh, an organization training their own model off of your data for no additional value to yourself.

And so, you know, if you're working in regulated technology, just make sure before you start uploading customer information to your favorite AI model for a quick summary, a quick check, or any of that type of stuff, just what happens to the data that you upload. Because believe it or not, as soon as you upload it, you may not own it. So there we go. Bit of food for thought, bit of something to do, bit of homework, and go and check your own models and your agreements. Um, if you enjoyed this, I will be back tomorrow with another AI briefing.

Thank you very much for joining me. My name is Tom, and we'll see you all soon. Bye for now. Why hire when you can partner? Concept Cloud's leading engineers build your startup's prototype without the overhead.

Launch faster. Conceptcloud. com.

Subscribe to The AI Briefing