Privacy and data disclosure risk when you call an LLM API
What leaves your machine on an API call, what may come back, and the rules that follow.

A call to a hosted language model is a data transfer, not a local computation. Everything in the request — the prompt, the conversation so far, and any document or record you included alongside it — travels to the provider’s systems. The output is a second departure point: it can disclose material from the model’s training corpus, and it can repeat sensitive material that was in your own context. Two things make this hard to reason about. Whether a provider may retain your request or train on it is a matter of the terms in force for your account and product, not a property of the model; and a restriction you place in the system prompt is not an access control. OWASP states that such restrictions “may not always be honored and could be bypassed via prompt injection or other methods.”
The mental model: what talks to what, and where data can leave
your machine provider
──────────── ────────
prompt + conversation + pasted material ──request──▶ receives content
├─ may retain logs
system prompt restriction ├─ may or may not train on it
(guidance, not a gate) └─ model has a training corpus
│
output ◀────────────────response─────────────────────────┘
│ (may contain corpus data or your own context)
▼
wherever you put it: history, notes, ticket, another system
This sheet’s reading, not a sourced claim: read the diagram as two leaks. The request is one: it leaves your control the moment you send it. The output is the other: whatever it contains sits wherever you paste it, and may be logged, indexed or forwarded. The provider box is deliberately vague: what happens inside it is set by that provider’s terms, not by the diagram.
Terminology the reader needs
- Request / context. The payload of one call: the prompt, the conversation, and anything attached or pasted. That is what the provider receives.
- Output. The text returned to you, produced from the model’s parameters and your context.
- Retention. Whether the provider keeps the request or output after serving the call, and for how long.
- Training. Whether the request or output updates the model’s parameters, rather than being processed once to answer you.
- Training corpus. The text the model was trained on before you used it; material from it can be reproduced in an output.
- System prompt / system instruction. Guidance the application sets for the model. It shapes behaviour; it does not enforce anything.
What happens, in the right direction
- You assemble a request and send it. From that moment the content — prompt, conversation and attachments — is on the provider’s systems.
- The provider processes the request to produce an output. What it may do with the content afterwards — keep it, review it, train on it — depends on the agreement for your account and product, not on the model.
- As one documented example, OpenAI states for its business offerings, “We do not train our models on your data by default”, and for data accessed from connected apps, “By default, we do not train our models on any data accessed from apps”; its data guide adds that API data “is not used to train or improve OpenAI models” unless the customer opts in, and that abuse-monitoring logs — which “may contain certain customer content, such as prompts and responses” — are “retained for up to 30 days”. These are one vendor’s statements — not a guarantee for any other provider, plan or endpoint.
- The output can disclose data from the training corpus, or repeat your own context. OWASP describes LLMs that “risk exposing sensitive data, proprietary algorithms, or confidential details through their output”, naming PII leakage and proprietary-algorithm exposure and noting that training-data exposure enables model-inversion attacks; a training-data extraction attack against GPT-2 “trained on scrapes of the public Internet” recovered “hundreds of verbatim text sequences”. NIST’s Generative AI Profile treats data privacy as “leakage and unauthorized use, disclosure, or de-anonymization” of sensitive data, and states that “Models may leak, generate, or correctly infer sensitive information about individuals.”
- You act on the output, and the data leaves the boundary once more.
Where disclosure can happen
| Point in the flow | What can leave or be exposed | What governs it |
|---|---|---|
| The request | prompt, conversation, pasted documents and records | classification and minimisation before you send |
| The provider | retention and any training use | the terms for your account and product |
| The model’s parameters | memorised training data, reproducible in an output | training-time choices you do not control |
| The output | corpus data, or your own context restated | what you ask for, and what you do with it |
| After the output | onward logging, storage, re-sending | your own handling |
The common conceptual error
The error is treating the system prompt as a fence: “I told the model not to reveal personal data, so it will not.” That is not how it works — a restriction in the system prompt “may not always be honored and could be bypassed via prompt injection or other methods” (OWASP). A model asked to repeat, infer or continue text does not check a policy first; the instruction is context, not an enforcement point. This sheet names prompt injection only to mark the failure mode; its mechanics are not developed here. The same applies to any prompt-level promise about retention or training.
Practical rules that follow
- Classify before you send. Decide what category the request would carry — personal, health, financial, confidential business — and whether it may leave your boundary at all.
- Minimise the context. Send the smallest amount that answers the question; a redacted excerpt instead of a whole file.
- Redact or replace identifiers. Substitute names, addresses and account numbers before they enter a request.
- Keep it local when the requirement demands it. Where a handling requirement cannot be met by hosted terms, local inference keeps the request inside your boundary — it removes the provider-side transfer, not the output-side risk (a model can still reproduce corpus data) and not your own logging and storage. Setting up a local model is the subject of the sibling sheet on model, runtime and application.
- Read the terms. The documentation in force for your account and product — not a general impression — says whether content is retained or used for training.
Level and prerequisites
L1 — what leaves the machine on a call and the conceptual rules that follow; no configuration, no retention-window engineering, no attack technique. Prerequisites: none.
Where to go next
- Automation & AI — the area this sheet belongs to.
- Data-protection law and policy belong to Cybersecurity Governance.
References
- OWASP, OWASP Top 10 for LLM Applications 2025 — LLM02:2025 Sensitive Information Disclosure: the exposure of sensitive data through model output, the PII-leakage and proprietary-algorithm examples, model inversion, and the warning that system-prompt restrictions can be bypassed.
- NIST, NIST AI 600-1, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (July 2024) — the data-privacy risk definition and the statement that models may leak, generate or infer sensitive information about individuals.
- Extracting Training Data from Large Language Models, arXiv:2012.07805 — the training-data extraction attack and the recovered examples.
- OpenAI, Enterprise Privacy and Your data — the provider’s documented defaults for training and retention, cited as one vendor’s statements about its own products.