ai privacy risks

AI Privacy Risks: What the Sources Actually Document

AI privacy risk is not one problem. This guide separates it into the failure modes the cited sources actually describe — the data appetite of the systems, what users hand to AI assistants, exposure of stored data, what training datasets absorb, and downstream misuse such as targeted fraud — then covers the research framing, the legal layers as described by IBM's vendor-published explainer, the consent default, and a single-vendor worked method for reading an AI provider's own privacy documentation.

This article was researched with AI assistance and independently reviewed by multiple AI models before publication.

Key takeaways

  • The Office of the Victorian Information Commissioner attributes the recent period of rapid AI development to three factors — improved algorithms, increased networked computing power, and an increased ability to capture and store an unprecedented amount of data — and, by this guide's reading, only the third directly involves personal information (a gloss OVIC's list itself does not state).
  • Yale's Privacy Office, in a notice to the Yale community, states that AI assistants are designed to collect, process and analyze vast amounts of data, and that this raises significant privacy concerns when the data collected includes sensitive personal information.
  • A student-authored CSULB College of Business legal resource article states that for AI to function correctly it needs, alongside high-functioning algorithms and processing power, a large quantity of personal data — a framing from that article, stated without distinguishing between systems in the captured extract.
  • A post on the Qualys security blog states that AI systems can collect and store large amounts of personal data, making them attractive targets for cybercriminals, and that many store data over long periods, sometimes without clear limits.
  • Stanford HAI's March 2024 coverage of a report on AI risks states that the data it discusses helps enable spear-phishing — the deliberate targeting of people for purposes of identity theft or fraud.
  • IBM's vendor-published Think explainer describes a rule set spanning the EU's GDPR, the California Consumer Privacy Act, the Texas Data Privacy and Security Act, Utah's March 2024 Artificial Intelligence and Policy Act, and China's 2023 Interim Measures for generative AI services.
  • A 2026 review by Voloch and Hirschprung systematically analyzes 94 research papers and describes the AI–privacy intersection as presenting both significant challenges and opportunities.

AI privacy risk is not one problem. This guide separates it into the distinct failure modes the cited sources actually describe: how the systems consume data, what users hand to AI assistants, what happens to data once it is stored, what training datasets absorb, and how already-collected data gets reused against people. It then covers how the research literature frames the field, the legal layers, the consent default and where it came from, and a checklist you can run against an AI provider's own privacy documentation.

Two rules govern everything below. Every factual statement is attributed to the specific source that makes it, linked at the point the claim is made — and where a source is a vendor, a student author, or a body writing to its own community, that is said plainly so you can weight it. Separately, this article keeps apart what a source states from what can be read into it: nearly every sentence here is the former, and the few consequential interpretations are marked where they appear. One further limit applies throughout: the material cited is a captured excerpt of each page, not the page in full, so every attribution means this extract states X — never a claim about everything a source does or does not cover.

Why AI concentrates privacy risk in the first place

The Office of the Victorian Information Commissioner (OVIC) attributes the recent period of rapid AI development to three factors: improved algorithms, increased networked computing power, and an increased ability to capture and store an unprecedented amount of data. Reading that list, only the third of the three directly implicates personal information — but note what that is: this guide's gloss on OVIC's list, not OVIC's own wording, which states the three factors without drawing that distinction. Taken as a reading rather than as a finding, it is the reason privacy risk has tracked AI capability rather than sitting beside it.

A frequently surfaced framing of the data question comes from a student-authored piece: "Artificial Intelligence: Privacy Concerns", written by Chloe Alexander, a candidate for business management, for the CSULB College of Business Legal Resource Center in 2024. It states that for AI to function correctly, it requires not only high-functioning algorithms and processing power but a large quantity of personal data. Weight that for what it is: a legal-resource article by a business management candidate, and the captured extract states the point about AI generally, without distinguishing between systems or applications. Treat it as that article's framing of why data appetite matters, not as a demonstrated property of every AI system.

The same CSULB article also sets out what is at stake — it states that the right to privacy allows people to manage their information, protect their reputation, and avoid unwanted attention — and, still in that article, frames the evolution of AI in society as having increased concerns about privacy rights.

OVIC's stated purpose for its own resource is to provide a high-level understanding of AI and its uses in the public sector and to highlight some of the challenges and opportunities AI presents in relation to information privacy. Among the promises of AI, it lists increased efficiency and lower costs, huge improvements in healthcare and research, increased safety of vehicles, and general convenience. Worth knowing before citing it: OVIC states the resource is written for a non-technical audience and does not endeavour to solve the questions it poses, nor provide legal guidance.

Five risk categories worth keeping separate

1. Opaque collection

The University of Texas Libraries' AI ethics and privacy guide describes ongoing privacy concerns and uncertainties about how AI systems harvest personal data from users. That is the first category: not a specific breach or a specific misuse, but the collection process itself.

2. Sensitive data handed to AI assistants

Yale's Privacy Office, addressing its own community, states that AI assistants are designed to collect, process, and analyze vast amounts of data to provide personalized and efficient user experiences — and that this capability raises significant privacy concerns when the data collected includes sensitive personal information, as set out in that notice to the Yale community. Read it as institutional guidance to a defined population rather than as a measured property of every assistant on the market.

Of the five categories here, this is the one most directly shaped by individual behaviour, since the exposure begins with what a person chooses to type into an assistant — an interpretation of the category, not a claim any cited source makes.

3. Exposure of data already stored

Collection risk and storage risk are different failure modes. A post on the Qualys security blog — a security vendor's own publication — states that AI systems can collect and store large amounts of personal data, making them attractive targets for cybercriminals, and that many AI systems store data over long periods, sometimes without clear limits. It also states that these tools often require access to large datasets and that, if those are not properly secured, there is a risk of sensitive information being leaked, which could include proprietary business data or personal information shared inadvertently.

The University of Texas Libraries guide's privacy reading list includes an item headlined "Huge Trove of Nude Images Leaked by AI Image Generator Startup's Exposed Database" — in that guide's own curation, an illustration of data accumulated by an AI product being exposed.

4. What training datasets absorb

The Qualys post states that generative tools learn from large datasets, and that this poses a risk if those datasets are unsecured or include sensitive information — which it says could include personal details or classified business data, posing a significant risk to privacy and security. This category is distinct from the previous one: the concern is not only a database left open, but what gets absorbed into a model's training corpus in the first place.

5. Downstream misuse of data that is already out there

Stanford HAI published a piece dated March 18, 2024 covering a report that analyzes the risks of AI and offers potential solutions. It names a concrete downstream harm: the data that report discusses helps enable spear-phishing — the deliberate targeting of people for purposes of identity theft or fraud. The referent there is the data the covered report is discussing, not all personal data everywhere. The same piece pushes back on fatalism: a lot of data has been collected about all of us, it says, but that does not mean a much stronger regulatory system cannot still be created.

What the research literature does with all this

For the scholarly view rather than the commentary view, a 2026 review by Voloch and Hirschprung in PMC systematically analyzes 94 research papers in the field of AI and privacy.

Its abstract describes the intersection of AI and privacy as presenting both significant challenges and opportunities. To model the issue, the review states that it categorized privacy in AI through a multi-dimensional approach that includes technological domains' privacy actions, privacy-preserving strategies, and AI–privacy interaction directions. Those three are the review's own taxonomy labels, reproduced here as it states them.

The review also grounds privacy as a concept rather than assuming it: a multifaceted concept spanning personal, social, and technological dimensions (citing Mulligan et al., 2016), usually revolving around the right of individuals to control information about themselves and to decide how and to what extent that information is communicated to others (citing Solove, 2004). The CSULB article likewise lists Daniel J. Solove's Artificial Intelligence and Privacy (February 1, 2024) among its references.

That control-centred definition is not only academic framing. IBM's explainer on privacy issues in the age of AI describes control over personal data as including the ability to decide how organizations collect, store and use that data. That definition is the one used as a measuring stick later in this guide.

The legal picture: several layers, not one instrument

A caveat before the list: IBM's Think explainer — published 30 September 2024, updated 24 February 2026, by Alice Gomstyn and Alexandra Jonker — is vendor-published commentary from a company that sells AI products, not an independent legal explainer. It is the only narrator of the legal layer in this evidence set, so read what follows as its account.

  • The EU's General Data Protection Regulation (GDPR), which the explainer describes as setting several principles that controllers and processors must follow when handling personal data.
  • US state privacy statutes, with the California Consumer Privacy Act and the Texas Data Privacy and Security Act given as examples.
  • AI-specific state law. The explainer states that in March 2024 Utah enacted the Artificial Intelligence and Policy Act, which — as characterized by IBM in that explainer as of its 24 February 2026 update — is considered the first major state statute to specifically govern AI use.
  • Non-US AI-specific rules. It states that in 2023 China issued its Interim Measures for the Administration of Generative Artificial Intelligence Services.

With a picture layered like that, asking whether a given AI product is compliant has no single answer — it depends on which layers reach you, which is a function of where you are and what the data is. That conclusion is drawn from the list; the explainer states the layers, not that verdict.

The same explainer opens on the general premise that as technology advances, so do the risks of using it — a framing worth holding alongside OVIC's list of AI's promises rather than instead of it.

The consent default and where it came from

One of the more useful historical points in the Stanford HAI piece is that today's collection default has an origin. As of that March 2024 account, the current default is described as the result of industry convincing the Federal Trade Commission about 20 years ago that if data collection switched from opt-out to opt-in, there would never have been a commercial internet. Against that default, the piece raises two proposals: a regulatory system requiring users to opt in to their data being collected, and one that forces companies to delete data when it is being misused.

That same March 2024 piece points to a platform-level control that had already shipped: Apple's App Tracking Transparency (ATT), which Apple launched in 2021 to address concerns about how much user data was being collected by third-party apps. As described there, when iPhone users download a new app, Apple's iOS system asks whether they want to allow the app to track them across other apps and websites. That description is a 2024 snapshot of platform behaviour; check Apple's current documentation before relying on the specifics today.

Those two things sit at different layers — a statutory default and a device-level prompt — and can coexist; treating the platform control as a substitute for the regulatory one would read more into the source than it says. That is an interpretation of the two proposals, not a position the piece takes.

Reading a vendor's explainer against its binding policy

Before the checklist, one worked demonstration of why both documents matter — using the only vendor in this evidence set that publishes both a general explainer and its own corporate privacy statement.

IBM's Think explainer defines privacy control, for the reader, as including the ability to decide how organizations collect, store and use that data. IBM's own privacy statement, effective 25 August 2026, answers the same question — who decides how data is collected and used — differently for business-to-business delivery: where IBM provides products, services, or applications as a B2B provider to a client, the client is responsible for the collection and use of personal information while using them. The statement adds that IBM's agreement with the client may allow IBM to request and collect information about authorized users of those products for reasons of contract management.

Name the gap plainly: the explainer speaks to an individual's control over their own data, while the binding statement locates that responsibility with the client organization and, at the same time, reserves the provider's own collection of authorized-user information. Both documents are IBM's, both address who controls data in an AI product, and they answer at different levels of the delivery chain. That is exactly the pairing to look for at any vendor — the marketing-adjacent explainer that tells you what privacy control means, and the policy that tells you who actually holds it.

This is a single-vendor worked method, not a cross-vendor comparison: only one provider's binding privacy statement is in the cited material here. Run the same two checks yourself against whichever vendor you are evaluating.

A checklist for reading an AI vendor's privacy documentation

Each check below is derived from something a cited source states. IBM appears throughout as the worked example because both document types are available for it — not as a recommendation of any product.

  1. Work out who is responsible for your data — and who still collects some of it anyway. IBM's privacy statement says the client is responsible for the collection and use of personal information in B2B delivery, and that IBM's agreement with the client may allow IBM to request and collect information about authorized users for reasons of contract management. Responsibility shifts; collection does not necessarily stop. Reading that forward: if you reach an AI tool through your employer or another vendor, the consumer-facing statement may not be the document that governs your data — an inference, not the statement's own words.
  2. Look for a product-specific supplementary notice. The same statement says IBM may provide additional data privacy information by using a supplementary privacy notice, so a top-level statement may not be the whole answer for one specific AI product.
  3. Count what account creation alone collects. IBM's statement says an IBMid provides IBM with your name, email address, mailing address, and related information you may provide, and that an IBMid may be required for certain services such as IBM Applications, Cloud and Online Services. Registration data is collection, before the product has been used at all.
  4. Check the wider legal surface, not just the privacy page. IBM's legal hub states that it also includes Fair Use guidelines for use and reference of IBM trademarks — a reminder that privacy terms sit inside a larger set of documents.
  5. Identify which regime applies to you — GDPR's controller/processor principles, a US state statute such as the CCPA or the Texas act, or an AI-specific law such as Utah's, per IBM's explainer.
  6. Ask whether collection is opt-in or opt-out, using the framing from the Stanford HAI piece above, and whether any tracking-consent prompt comparable to Apple's ATT exists for the product.
  7. Ask specifically about sensitive categories, per Yale's point to its community that the concern sharpens when what an assistant ingests is sensitive personal information.
  8. Ask about retention and training data, given the Qualys post's points that many AI systems store data over long periods without clear limits and that generative tools learn from large datasets that may include sensitive information.
  9. Ask what happens on deletion and misuse — the remedy Stanford HAI raises as a proposal, not as an existing universal entitlement.

Adjacent concern: this is not only about data

The University of Texas Libraries guide places privacy alongside other AI ethics topics, including bias and environmental impact — noting that AI is typically associated with virtuality and the cloud, yet relies on vast physical infrastructures that span the globe and require tremendous amounts of natural resources, including energy, water, and rare earth minerals. That is a different axis from privacy, noted here so the two are not conflated.

How current is any of this

The cited material carries different timestamps, and they matter when you quote it. The CSULB article and the Stanford HAI piece are 2024. The Qualys post is dated February 2025. IBM's explainer was published 30 September 2024 and updated 24 February 2026. IBM's privacy statement is stated as effective 25 August 2026. The Voloch and Hirschprung review carries a 2026 copyright. The University of Texas Libraries guide's reading list references items dated through late 2025. Legal specifics in particular — state statutes, regulatory defaults, platform consent prompts — are the parts most likely to have moved since capture, so verify them at the source before relying on them.

Sources