
“We don’t train on your data.”
Good. Now ask where yesterday’s prompts are stored, who can review them, and what happens when your AI assistant sends a document to another tool.
Those answers can change whether you approve a workflow, redesign it, or keep it out of production. A no-training promise answers one important question. It doesn’t finish the security review.
And it’s not a good strategy to just “take their word for it.”
Why this matters
Fortune’s October 5 report describes the tension between enterprise privacy requirements and AI providers’ efforts to detect misuse across multiple interactions. It also examines interest in open models and sovereign AI.
There is a legitimate procurement issue here without assuming vendors are stealing intellectual property. OpenAI says API data is not used for training unless customers explicitly opt in. Anthropic says the same for its commercial products by default, while identifying feedback and other permissions as exceptions.
A company can honor that commitment and still retain data for safety monitoring or application functionality. Retention, model training, and human access are different questions.
My recommendation: approve a specific workload and its data path, not a vendor logo or a privacy slogan.
What changed, and what is actually available
OpenAI’s August 19 announcement introduced Private Safety Processing, or PSP. Its September 22 update says PSP is rolling out to API customers in phases.
OpenAI describes an arrangement in which content remains in customer-controlled storage and automated safety review produces limited signals without exposing protected content to OpenAI personnel. OpenAI’s current API documentation also describes Private Retention with PSP: customer content is retained in encrypted logs on OpenAI-managed infrastructure and excluded from human review unless applicable law requires it. This is distinct from ZDR with PSP.
Anthropic’s September 1 announcement describes Enterprise Frontier Safeguards, or EFS, with a phased rollout planned for later this fall. It says eligible customers receive ZDR on Fable 5 and Fable 5.1 during the transition. Customer-owned storage, customer-managed keys, and fully automated review are each opt-in controls, not an automatic bundle.
These are vendor descriptions, not proof that your account has those protections enabled. Ask for the applicable contract, model eligibility, and configuration evidence. A rollout announcement cannot approve your production architecture for you.
Follow one request all the way through
I’d use a “request-to-deletion” test before approving a sensitive AI workflow. Pick one representative request and trace every place its content can go, from the user’s screen through the model, tools, logs, and backups.
Ask four questions at every stop:
- What data arrives here, including attachments, retrieved documents, and outputs?
- What is retained, for how long, and under which exceptions?
- Who or what can access it, and who controls the keys and permissions?
- What evidence shows deletion, and who handles an incident or policy change?
This exposes a common gap: the model endpoint may have an acceptable retention policy while the application saves complete conversations in a debugging service. Or the assistant may send sensitive text to a connector governed by a different company’s terms.
The AI vendor’s promise does not erase those copies.

ZDR has a scope. Read it before you build around it.
Zero data retention (ZDR) can be useful. Treat it as a defined commitment for particular data, models, and features rather than a claim that nothing anywhere gets stored.
OpenAI’s API data-controls documentation distinguishes abuse-monitoring logs from application state. Standard abuse logs may retain customer content for up to 30 days, with stated exceptions. ZDR requires approval. Eligible configurations exclude customer content from those logs, but ineligible endpoints can still retain application state.
For example, its table marks files, conversations, and vector stores as not ZDR-eligible. Features such as background processing and prompt caching have additional retention details. Data sent to remote MCP servers, the services an agent can call as tools, follows those services’ policies.
Human access needs its own question. OpenAI documents a narrow but important exception: images flagged as potential child sexual abuse material are retained for manual review even with ZDR. It also documents circumstances in which particular models can become ineligible for a customer’s retention controls, with advance written notice. “Nobody can ever see anything” is too broad.
Anthropic’s API documentation likewise requires organization-level ZDR enablement and distinguishes eligible Messages API features from stateful features such as files, batches, and code execution. Using a non-eligible feature can put that data outside the arrangement. Covered Models require retention unless Anthropic expressly authorizes an exception; flagged content and legal holds have separate retention rules.
Check the exact model, endpoint, feature, and account configuration together. An approved organization is not permission to use every new feature with the same sensitive data.
Sovereignty is a control decision, not a server location
Keeping records in your own cloud account can improve custody. It also gives your team responsibility for storage permissions, lifecycle rules, monitoring, and cost.
OpenAI’s PSP developer guide says OpenAI keeps an index containing operational metadata and a storage reference while protected content sits in customer-controlled storage. The guide requires customers to retain encrypted PSP records for at least 30 days. Automated review still retrieves and processes that content. Customer-controlled storage therefore does not mean the provider has disappeared from the system.
Separate where data is stored from where inference runs, who administers the service, which jurisdiction applies, and which external tools receive data. Even data residency can have scope limits: OpenAI distinguishes customer content from system data such as account and usage information.
Self-hosting an open-weight model can remove a hosted inference provider from that path. It does not remove vulnerable software, excessive permissions, careless logging, outbound telemetry, or an agent allowed to email a confidential file. Sovereignty gives you more control over certain dependencies. You still have to exercise it.
Choose the deployment by workload
Use hosted AI when its contractual and technical controls meet the workload’s requirements and its quality and operating simplicity justify the choice. A public-document research assistant should not automatically inherit the infrastructure requirements of a sensitive research repository.
Consider private-cloud or customer-controlled deployment when you need tighter network, identity, key, or regional controls. Verify what “private” means: a private network connection to a provider’s service is different from running inference inside your own environment.
Consider self-hosting when a concrete requirement, such as disconnected operation or a prohibition on external inference, warrants it. Then test whether the model is good enough and whether your team can patch, monitor, recover, and support it. Include staffing and reliability in the economics, not just GPUs and tokens.
A mixed architecture may be the sensible answer. Keep restricted workloads inside an approved boundary and use hosted services where they fit. Don’t buy hardware to avoid reading a contract, and don’t accept a contract as a substitute for understanding the system.

A healthcare example
Consider an illustrative hospital workflow that drafts staff guidance from approved policy documents. Start with synthetic questions and documents cleared for that environment. Restrict retrieval to the approved collection and require a person to check the answer against its source.
Adding patient records changes the approval decision. Before introducing protected health information, confirm the applicable business associate agreement, covered services, enabled configuration, access controls, and every downstream tool. Decide what audit evidence to retain without casually logging full patient conversations.
A BAA is not an automatic HIPAA-compliance guarantee. Neither is ZDR. The organization still has to assess and operate the complete workflow appropriately.
Your AI Pathfinder action plan
Bring one proposed workflow to your next vendor meeting. Ask for written answers to these procurement questions:
- Which exact models, endpoints, features, and deployment channels are covered today?
- Which protections require approval or separate enablement, and can you demonstrate our settings?
- What survives the request: content, cached representations, application state, safety signals, or metadata?
- Under what circumstances can a person review content, retention change, or deletion be delayed?
- Which tools, subprocessors, support systems, and customer-managed stores fall outside this commitment?
- What notice, audit evidence, and exit options do we get when the policy or model changes?
Have security, the application owner, and procurement reconcile those answers against the actual data path. If a critical answer is missing, keep that workload on synthetic or approved nonsensitive data until it is resolved.
Frequently asked questions (FAQs)
Does “no training” mean no storage?
No. A provider can exclude data from model training while retaining it for application features, safety, or legal obligations.
Do we need ZDR for every AI use case?
No. Some workflows need retained history or audit records. Set a justified retention policy and confirm where those records live and who can access them.
Is a local model automatically safer?
No. It changes the trust boundary and operating responsibilities. Compare complete systems, including tools and logging, rather than model locations alone.
The bottom line
Before approving sensitive AI work, require a request-to-deletion map with named owners and documented exceptions. Then choose the deployment that meets those requirements. You may need a different configuration, a narrower workflow, or your own infrastructure. Make that decision from the evidence.
About Jason Fleagle
Jason Fleagle is Head of AI at Netsync and writes AI Pathfinder about practical enterprise AI decisions. Find more of his work at thejasonfleagle.com and netsync.com.
Originally published on LinkedIn.



