Skip to main content
Industry Insights
11 min read

How AI Companies Scrape and Sell Your Personal Data in 2026

A source-aware guide to public-web collection, AI product interactions, provider training controls, robots.txt boundaries, and personal-data requests.

Rahul Kandoriya
Written byRahul Kandoriya·Last updated August 25, 2026
How AI Companies Scrape and Sell Your Personal Data in 2026
How AI Companies Scrape and Sell Your Personal Data in 2026
Coverage scope: The OfflistMe catalog currently records 1,000+data-broker workflows. Paid access lets you select workflows at once; you review and send or submit the generated requests, while provider eligibility and outcomes remain outside OfflistMe's control.

AI services can encounter personal information through several different paths: public web pages, user prompts and uploads, connected applications, business data feeds, advertising systems, and service providers. Those paths are often confused. A public post is not automatically a training record, a chatbot answer is not proof that a data broker sold a profile to the model, and an account setting is not necessarily a deletion request.

This guide separates what can be verified from what is often inferred. It focuses on practical controls for individuals and website owners. Provider terms and settings change, so open the current first-party policy before submitting a request or sharing sensitive information.

Key takeaways

  • An AI company may process data to provide a response, prevent abuse, personalize a service, evaluate quality, or improve models. The applicable purpose depends on the product, account type, region, setting, and provider policy.
  • A personal ChatGPT workspace can turn off “Improve the model for everyone” so new conversations are not used to train ChatGPT; business, enterprise, education, and API offerings have different default terms. OpenAI documents these distinctions in its Data Controls FAQ.
  • Gemini’s Keep Activity, connected-app, temporary-chat, feedback, and account settings are separate controls. Google says that connected-app data can be used to improve services, including training generative AI models, when the relevant activity setting is on. See the Gemini privacy hub.
  • Website owners can use crawler controls such as `Google-Extended` for certain Gemini training and grounding uses. Google says that token does not affect Google Search inclusion or ranking. It is not a universal deletion mechanism for prior crawls or third-party datasets.
  • Removing a people-search profile may reduce one public source. It does not prove that an AI company, search index, data broker, customer, or model has a copy, and it does not erase every downstream use.
  • Do not put secrets, confidential client material, health details, identity documents, or another person’s personal information into a service unless you have a clear reason and permission to do so.

Three different data paths

1. Public-web collection and indexing

AI systems and search products may retrieve, index, license, or otherwise process public pages. Public availability does not establish that a specific model trained on a page. A website owner can change or remove the source, request search refreshes, and use provider-specific crawler controls where available.

The source remains the first place to address. If an old profile, forum post, image, or business page exposes personal information, ask the site owner to correct or remove it when appropriate. Then use a search engine’s current outdated-content or personal-information process if its requirements fit. Search-result removal does not automatically remove the source page.

Request Drafting

Tired of dealing with data exposure?

Choose relevant provider workflows, review the generated drafts in your browser, and send or submit each request yourself. Matching, eligibility, and provider requirements still need checking.

Review Removal Options Free for selected workflows · No opt-out profile stored · No card needed

2. Product interactions

Prompts, responses, uploads, voice recordings, screenshots, feedback, and connected-app data can be processed to deliver the feature and may be used for improvement under the product’s current policy. The same provider may have different rules for personal, business, education, API, and regional accounts.

For example, OpenAI says that personal ChatGPT workspaces can opt out of training for new conversations through Data Controls. It also says business, enterprise, education, and API offerings do not use inputs and outputs to train models by default, subject to their applicable terms. OpenAI separately notes that feedback can cause the associated conversation to be used for training. Read the current explanation of how data is used to improve model performance.

Google’s Gemini documentation distinguishes Keep Activity, connected apps, temporary chats, feedback, and work or school administrator controls. Google says that turning off Keep Activity stops future chats from being used to improve Google AI, but it retains chats for a limited period for service and safety purposes. It also warns that deleting data in a connected app does not necessarily delete the corresponding Gemini activity. Read Manage and delete Gemini Apps activity before treating a setting as erasure.

3. Commercial data and retrieval services

Some AI products may use licensed, customer-provided, public, or partner data for search, grounding, identity, fraud, marketing, or another business function. The existence of a data broker and an AI product does not prove that a specific people-search profile was licensed to a model or that an output came from that profile.

To make a defensible claim, identify the provider, product, notice, dataset, contract, public statement, or enforcement record that supports it. Otherwise describe the relationship as a possibility or an unresolved question, not as an established pipeline.

What an AI answer can and cannot show

If a chatbot produces your name, address, employment detail, or an incorrect profile, that output does not by itself reveal whether the information came from training data, a live search, a connected tool, a prompt, a public page, or a hallucination. Save the exact prompt and response only when necessary, avoid republishing the personal information, and use the provider’s current reporting or privacy route.

A model’s inability to answer does not prove that it never processed the information. A model’s confident answer does not prove that the information is true. Verify the underlying source and correct it at the source whenever possible.

Controls for personal accounts

OpenAI and ChatGPT

For a personal ChatGPT account, open Profile → Settings → Data Controls and turn off Improve the model for everyone. OpenAI says that new conversations will not be used to train ChatGPT after the setting is turned off, while chats can remain in history. Temporary Chat has separate retention and training behavior described in OpenAI’s help materials.

If you use ChatGPT Business, Enterprise, Edu, or the API, read the applicable business or API privacy terms. Do not copy personal-account steps into a business environment or assume that deleting a chat is the same as a data-subject erasure request. For a personal-data request, use OpenAI’s Privacy Portal and keep the confirmation.

Google Gemini

Review Gemini Apps Activity and the Keep Activity setting. Google says that future chats with Keep Activity off are not used to improve Google AI when the relevant conditions are met, while a limited retention period can still apply for service and safety. Connected apps, Gemini Live audio or video, screenshares, uploads, feedback, and Workspace accounts can have additional rules.

Do not connect Gmail, Drive, Photos, Calendar, contacts, or other sources containing confidential information merely to make an answer more convenient. Google’s connected-app explanation says that summaries, excerpts, generated media, and inferences can be used to improve services, including model training, when the applicable activity setting is on.

Other providers

For any other AI service, look for the current privacy notice and settings for:

  • model-improvement or training use;
  • prompts, uploads, voice, screenshots, and feedback;
  • retention and deletion;
  • human review and safety monitoring;
  • connected applications and third-party tools;
  • business, education, or API accounts; and
  • regional rights, appeal, and complaint routes.

Do not infer a provider’s policy from a competitor’s setting or from a generic “AI privacy” article.

Controls for website owners

Website owners have a different set of controls. Google documents that the `Google-Extended` user-agent token in `robots.txt` can manage whether content Google crawls may be used for training future Gemini models and for certain grounding uses. Google also says `Google-Extended` does not affect Google Search or other products. See Google’s crawler documentation.

Crawler controls are prospective and product-specific. They do not recall a page already downloaded, remove a third-party copy, delete a model’s parameters, or prevent a human from reading a public page. Use access controls, authentication, removal, and data minimization when content should not be public at all.

Before adding a robots rule:

  1. Confirm the provider token and the product scope in current documentation.
  2. Test that the rule does not block the search crawling or integrations you need.
  3. Keep public and private content on separate paths with appropriate authentication.
  4. Remove unnecessary personal information from the source page rather than relying only on crawler preferences.

Removing personal information from an AI service

Start by classifying the exposure:

ExposureFirst controlBoundary
Your public page or profileCorrect or remove it at the sourceSearch and AI indexes may update separately
A chatbot conversation or uploadDelete activity, change training setting, or use the provider privacy routeRetention, safety, feedback, and legal exceptions may apply
An AI search answer quoting a sourceReport the answer and correct the sourceA response report does not erase the source
A people-search profileUse the provider’s current opt-out or deletion routeIt does not prove what an AI provider received
A customer or employer’s AI systemAsk the organization about its controller, processor, retention, and rights routeA model vendor may not control the customer’s copy

Use a specific URL, account identifier, prompt date, or record reference when the provider asks for one. Supply the minimum verification information necessary. Never upload a government ID to a generic “AI removal” form without confirming the recipient, legal basis, storage, and alternatives.

What a data-broker opt-out can do for AI privacy

A people-search opt-out may reduce a public lookup path used by a person, an advertiser, an application, or an AI-powered search feature. It may also prevent a future provider copy if the provider honors the request. It does not establish that a broker supplied an AI model, recall data already licensed to another party, remove public records, or change model behavior.

If you use OfflistMe’s directory, its selected provider workflows can help you review current routes and prepare browser-local request drafts. You review and send or submit each request yourself. The catalog is not an AI-training dataset inventory and does not contact AI companies, request model retraining, or guarantee removal from an AI output.

Review selected provider opt-out workflows →

Legal boundaries

Privacy rights vary by jurisdiction, controller, account, data type, processing purpose, and exception. GDPR Article 17 can provide an erasure route in appropriate circumstances, but it is not an automatic order that every model, backup, search index, or lawful record be erased. U.S. state laws may provide access, deletion, correction, opt-out, or sensitive-data rights subject to scope and exceptions. A provider’s voluntary training control is different from a statutory erasure request.

The answer also depends on the organization’s role. A model vendor, an application developer, a customer using an API, a data broker, and a website owner may each control different copies or purposes. Ask which organization receives the request and whether it can act on the specific data path.

Practical reduction plan

  1. Search for the actual source page or output; do not assume a broker-to-model pipeline.
  2. Remove or correct unnecessary personal information at the source you control.
  3. Review the AI product’s training, retention, feedback, connected-app, and deletion settings.
  4. Use the provider’s current privacy request or appeal route for a specific account or output.
  5. For an AI search result, report the result and separately address the source page.
  6. Review public people-search profiles as a separate privacy task.
  7. Use unique passwords, multifactor authentication, device controls, and least-privilege app connections.
  8. Re-check only the sources that matter to your risk; no universal schedule proves that every copy is gone.

Frequently asked questions

Can I make an AI company delete my personal information from a model?

You can submit a provider-specific privacy or erasure request where a route and legal basis apply, but the result depends on the provider, data path, jurisdiction, technical process, and exceptions. A user cannot assume that deleting a conversation removes information from every training artifact or model parameter.

Does ChatGPT know personal information about me because it was trained on the web?

Not necessarily. A response may come from a public source, a connected tool, a prompt, or a model error. OpenAI provides account data controls and a privacy portal, but an output alone does not identify the source or prove that your specific profile was in training data.

Does turning off training delete old conversations?

No. OpenAI describes the training toggle and deletion as separate controls. Google likewise distinguishes Keep Activity, deletion, and temporary chats. Read the current provider instructions and use the control that matches your goal.

Can robots.txt stop AI training?

It can express a website owner’s preference to a provider that documents support for the token, but scope and enforcement are provider-specific. Google says `Google-Extended` does not affect Google Search; it is not a universal deletion or recall mechanism.

Will removing my data from people-search sites remove it from AI systems?

Not necessarily. It can reduce one source or public lookup path. It does not prove that the information was used by an AI provider or erase independent copies, public records, or prior processing.

AI privacy checklist

  • [ ] Identify the actual source, account, output, or provider before making a causal claim.
  • [ ] Remove unnecessary personal information from public pages you control.
  • [ ] Review training, retention, feedback, connected-app, voice, and upload controls.
  • [ ] Use the provider’s current privacy or deletion route for a specific data path.
  • [ ] Keep business, API, school, and personal-account policies separate.
  • [ ] Treat robots.txt as a crawler preference, not complete deletion.
  • [ ] Submit people-search opt-outs separately and retain source evidence.
  • [ ] Avoid putting confidential or sensitive information into an unverified AI service.

Related guides

Take back your privacy today

Review provider-specific routes, prepare your requests locally, and send or submit each one yourself.

Review Provider Routes

Free to review provider routes · Optional one-time unlock from $9.00 · No subscription