
Generative AI tools are now part of the working day for millions of people. Across Indian banks, hospitals, IT companies, and government departments, employees are using AI assistants to draft emails, summarize documents, generate code, and answer customer queries. The tools are fast, capable, and getting cheaper by the month.
What is far less discussed is the privacy risk that comes bundled with the technology itself. Generative AI is not a neutral tool that processes information and forgets it. These systems are built on vast quantities of data often including personal information about real people and the way they learn from that data creates risks that do not go away once the model is trained and deployed.
Personal data has appeared in AI-generated outputs. Confidential business information has leaked through AI systems used by employees who did not realize the risk. Organizations deploying generative AI without understanding how the technology handles data are, in effect, operating a pipeline they cannot fully see or control.
This article explains how these privacy risks arise in plain terms, what India's Data Protection Board and global regulators are already doing about them, and what organizations need to put in place before enforcement catches up with adoption.
To understand the privacy risks, it helps to understand simply how these tools are built. A large language model (LLM) is trained by processing enormous quantities of text websites, books, news articles, code repositories, and user-generated content. The model reads this text, learns the patterns in it how words relate to each other, how ideas connect and uses those patterns to generate new text when given a prompt.
The model does not store text the way a database does. It compresses what it learned into billions of numerical values. But that compression is imperfect. Under certain conditions, a model can reproduce fragments of its training data including personal information about real individuals when prompted in specific ways.
This is not speculation. In a landmark 2021 paper published at the USENIX Security Symposium, researchers from Google, Stanford, UC Berkeley, and OpenAI demonstrated that they could extract hundreds of verbatim text sequences from GPT-2's training data, including names, phone numbers, and email addresses data that appeared in the training set just once. A follow-up study in 2023 showed that this extraction was possible even from production systems like ChatGPT, at a rate that researchers estimated could yield approximately a gigabyte of training data if pursued at scale.
This phenomenon called memorization is a structural feature of how these models work, not a bug that can simply be patched. The implication for privacy is direct: if personal data was in the training set, there is a real possibility it can surface in someone else's AI-generated output.
The first privacy risk in generative AI sits at the very beginning of the pipeline: the data used to train the model. Most publicly available LLMs were trained on data scraped from the internet at enormous scale. Datasets like Common Crawl one of the most widely used training datasets contain billions of web pages assembled without individual consent. This data includes personal information: names, email addresses, phone numbers, medical disclosures shared in online forums, financial details discussed in public threads, and much more.
In most cases, the people whose information was collected had no idea it was included in an AI training pipeline. They posted in a health support forum, or their email appeared in a public directory, without agreeing to have that information used to train a commercial AI system. The question of whether this data collection is lawful under data protection frameworks is one regulator are actively working through and the answers are not comfortable for the industry.
For organizations building or fine-tuning their own AI models using internal customer records, employee data, or proprietary business information as training inputs the problem is more direct. Those organizations must be able to identify their lawful basis for using personal data as training material, apply data minimization, and put in place technical safeguards to prevent that data from surfacing in model outputs. Most organizations deploying AI tools have not worked through these questions carefully, if at all.

The second category of risk is output leakage situations where a generative AI system produces content that reveals information it was never supposed to share.
As the USENIX 2021 research demonstrated, models can and do reproduce fragments of their training data. The risk is higher for data that appeared repeatedly across many sources in the training set, and it increases with model size. Google's research into differential privacy for AI is a direct acknowledgment by the industry itself that memorization is a live problem one serious enough to invest significant resources in addressing.
For most general-purpose consumer AI tools, the practical probability of encountering memorized personal data in a given interaction is low. For organizations that have fine-tuned models on their own proprietary data customer records, medical files, financial data the risk is considerably higher.
The more immediate risk for most organizations is simpler: employees using external AI tools in ways that send personal data outside the organization's control.
When an employee pastes a customer's details, a contract's terms, or a medical record into an AI tool, that information leaves the organization. According to OpenAI's own privacy policy for consumer services, content submitted through those services may be used to train or improve models. OpenAI's enterprise privacy page is explicit that this default does not apply to enterprise customers but most Indian organizations using AI tools are not doing so under formal enterprise agreements with negotiated data terms. Their employees are using consumer-grade tools, often informally, and the data they enter travels accordingly.
This is not a reason to prohibit generative AI. It is a reason to understand the data flows, and to ensure that personal data is only entered into tools that have been assessed and approved for that purpose.
Regulators have not been waiting for the law to catch up. The enforcement activity is already underway, and it is instructive for Indian organizations.
In March 2023, Italy's data protection authority, the Garante, became the first regulator in the world to temporarily ban ChatGPT, finding that OpenAI lacked a lawful basis for processing personal data to train its model, failed to verify users' ages, and had not transparently informed users about how their data was being used. The ban was lifted after OpenAI made changes, but the investigation continued. In December 2024, the Garante issued a €15 million fine against OpenAI the first GDPR fine specifically targeting a generative AI system citing the same failures and OpenAI's failure to notify the authority of a March 2023 data breach.
At the European level, the European Data Protection Board issued Opinion 28/2024 in December 2024, the first formal pan-European guidance on personal data in AI models. The opinion addressed when AI models can be considered truly anonymous the answer being rarely and clarified that using personal data to train AI models requires a lawful basis assessed on a case-by-case basis, with no blanket exemptions. The EDPB also confirmed that if training data was collected unlawfully, models built on that data carry the taint of that unlawfulness.
In India, MeitY released the India AI Governance Guidelines in November 2025 under the IndiaAI Mission, built around seven principles including privacy, transparency, and accountability. These guidelines are not yet legally binding, but they signal the direction of travel clearly. As the Data Protection Board of India becomes operational, AI-related personal data handling will come under the same scrutiny Indian regulators have already indicated is coming.

The Digital Personal Data Protection Act 2023 does not mention artificial intelligence by name. It does not need to. Its obligations apply to any processing of personal data, including processing done through or by AI systems.
Under Section 6 of the DPDP Act, personal data may only be processed for a specific, lawful purpose. An organization that allows employees to input customer personal data into an external AI tool without a clear lawful basis, without informing the Data Principal, and without a proper Data Processing Agreement in place is likely in breach of these requirements.
The Act's data minimization principle requires organizations to collect and use only as much personal data as is strictly necessary for the stated purpose. An employee who pastes an entire customer database into an AI tool to extract a single data point has not applied minimization.
Where an organization engages an AI tool provider as a Data Processor that is, where the tool processes personal data on the organization's behalf the organization, as Data Fiduciary, remains fully responsible for how that processing occurs. This mirrors the Article 28 GDPR requirement for a formal Data Processing Agreement, and it applies with equal force under the DPDP Act. Standard terms of service offered by AI tool providers are not a substitute for a proper data processing agreement.
For AI use cases involving large-scale personal data processing, sensitive categories of data, or automated decision-making that affects individuals, organizations should conduct a Data Protection Impact Assessment. The GDPR makes this a legal requirement under Article 35 for high-risk processing. Under the DPDP Act, it is a recommended best practice for Significant Data Fiduciaries and any organization whose AI use could cause significant harm to individuals if something goes wrong.

Responsible generative AI use is not a single policy document. It is an ongoing practice built from specific, actionable steps.
Most organizations do not know which AI tools their employees are using day-to-day. Consumer tools adopted informally tools employees signed up for personally, using organizational email addresses or for work tasks fall outside any formal data governance framework. The first step is knowing what is in use and what data is flowing through those tools.
Before formally deploying any generative AI tool for a purpose that involves personal data, the organization should document what data the tool will process, where it goes, how long it is retained, and whether it is used to improve or train the model. Consumer tools and enterprise tools often operate under very different terms and the difference matters enormously for compliance.
AI tool providers handling personal data on the organization's behalf should be treated as Data Processors. This means a Data Processing Agreement that prohibits using personal data for model training without explicit consent, requires breach notification within a defined window, and gives the organization audit rights. Most standard AI tool subscriptions do not include these terms unless negotiated.
Employees should not input personal data into AI tools unless the tool has been approved for that use and only the minimum necessary data is included. Where possible, personal information should be removed or anonymized before a prompt is submitted. This is not just good practice it is what the DPDP Act's minimization principle requires.
Most employees using AI tools do not think of themselves as making data protection decisions. They are trying to work faster. Training should be practical and specific: which tools are approved for which tasks, what data must not go in, and what to do if something goes wrong. Compliance obligations explained in legal language will not change employee behavior; plain guidance will.
The ndia AI Governance GuidelinesI released by MeitY in November 2025 place privacy as one of seven foundational principles for responsible AI. The Data Protection Board of India is moving toward full operationalization. The enforcement gap between AI adoption and regulatory scrutiny is closing. Organizations that wait for a formal enforcement action to prompt action will find themselves behind.
The practical steps are straightforward:
Generative AI is not going away, and the right response is not to avoid it. The right response is to understand it well enough to use it responsibly.
The privacy risks in generative AI are real: training data that included personal information without consent, models that can reproduce that data under certain conditions, and employees who send personal data to external systems they do not fully understand. None of these risks require worst-case scenarios to matter. They are present in everyday use.
The DPDP Act 2023 applies the same principles to AI that it applies to any other processing of personal data: lawfulness, minimization, transparency, accountability, and protection. AI is a new context. The principles are not. Organizations that build those principles into how they adopt and govern AI tools will be better placed with the regulator, and with the individuals whose data they are responsible for, than those that treat compliance as something to address after the fact.
We at Data Secure (Data Privacy Automation Solution) DATA SECURE - Data Privacy Automation Solution Solution can help you to understand Privacy and Trust while lawfully processing the personal data and provide Privacy Training and Awareness sessions in order to increase the privacy quotient of the organisation.
We can design and implement RoPA, DPIA and PIA assessments for meeting compliance and mitigating risks as per the requirement of legal and regulatory frameworks on privacy regulations across the globe especially conforming to GDPR, UK DPA 2018, CCPA, India Digital Personal Data Protection Act 2023. For more details, kindly visit DPO India – Your outsourced DPO Partner in 2025 (dpo-india.com).
For any demo/presentation of solutions on Data Privacy and Privacy Management as per EU GDPR, CCPA, CPRA or India DPDP Act 2023 and Secure Email transmission, kindly write to us at info@datasecure.ind.in or dpo@dpo-india.com.
For downloading the various Global Privacy Laws kindly visit the Resources page of DPO India - Your Outsourced DPO Partner in 2025
We serve as a comprehensive resource on the Digital Personal Data Protection Act, 2023 (Digital Personal Data Protection Act 2023 & Draft DPDP Rules 2025), India's landmark legislation on digital personal data protection. It provides access to the full text of the Act, the Draft DPDP Rules 2025, and detailed breakdowns of each chapter, covering topics such as data fiduciary obligations, rights of data principals, and the establishment of the Data Protection Board of India. For more details, kindly visit DPDP Act 2023 – Digital Personal Data Protection Act 2023 & Draft DPDP Rules 2025
We provide in-depth solutions and content on AI Risk Assessment and compliance, privacy regulations, and emerging industry trends. Our goal is to establish a credible platform that keeps businesses and professionals informed while also paving the way for future services in AI and privacy assessments. To Know More, Kindly Visit – AI Nexus Your Trusted Partner in AI Risk Assessment and Privacy Compliance|AI-Nexus