Security 9 min read
What a zero-data-retention agreement actually is — and what it isn’t
There’s an argument moving through the industry right now, and it deserves a real answer rather than a brochure. It goes like this: model providers promise not to train on your data, but the promises are written in defined terms — Input, Output — and a reasoning model generates something in between: intermediate work product, produced from your data, that you’re billed for and never shown. If the defined terms don’t cover it, the argument runs, the provider keeps a notebook full of your thinking. And the strongest counter the providers offer — zero data retention — is granted only on separate request, which tells you what the default is.
We think the argument is worth taking seriously, because most of it is simply true as a description of the industry. Retention is the default. The defined terms in most agreements were written before reasoning models existed. And “we don’t train on your data” is a weaker sentence than most people hear it as. This article is about what a zero-data-retention agreement actually does — and about the specific ways Scout’s architecture was built so that the sharpest version of this worry has nothing to attach to.
A ZDR is a contract, not a setting
Start with what the term means, because it’s usually cited and rarely explained. Zero data retention is not a checkbox in a dashboard or a line in a privacy policy. It’s a negotiated contractual commitment — typically an addendum to the provider’s commercial terms, granted on separate approval — under which the provider agrees that when a request completes, it holds nothing. The pages traveled inside the request, were read in the moment, and were never written to storage on the far side. Not retained for thirty days, not retained for abuse review, not retained at all.
That distinction — contract versus policy — is the whole point. A policy is a page the provider can edit. A contract is an instrument with parties, defined obligations, and remedies for breach, enforceable in a specific jurisdiction’s courts. When your examiner asks what happens to member data at the model provider, “their website says” and “our signed agreement says” are different categories of answer. VisionFI’s model relationships are the second kind, and the shipped promise — zero retention, US inference, no training — is the summary of what those agreements put in writing.
The skeptic’s observation that ZDR “takes a special request” is correct, and it cuts the other way: the special request is the diligence. Retention is the industry default precisely because most customers never ask. We asked, and the answer is signed.
The boundary also carries a cost we accept without much ceremony: not every model is offered under zero-retention terms. Some of the most capable models in the industry — including some of the newest — are available only with retention, and Scout simply doesn’t run them. The roster is chosen inside the boundary, not the other way around. That’s the same posture the review article describes elsewhere in the product: feature decisions lose to the boundary, and so do model decisions.
Enforceable is not the same as verifiable
Here’s the limit, and we’d rather state it than have you discover it: no customer can technically verify non-retention from the outside. You cannot audit the absence of data. A ZDR is enforceable — breach carries contractual liability, and for a provider serving regulated industries, litigation discovery and regulatory consequence besides — but you cannot point a scanner at it.
So what does the agreement actually buy? The same thing every critical-vendor contract buys, assessed the same way. Your institution cannot technically verify that its core processor or its item-processing outsourcer never misuses the data flowing through them either. What diligence actually evaluates is the counterparty: who they are, what law they operate under, what they’d stand to lose. A large US frontier lab with enterprise customers in examined industries has profoundly asymmetric incentives — whatever marginal value one customer’s traffic might add to a training corpus is trivial next to the cost of being caught breaching a signed retention agreement. That’s not faith in anyone’s goodness. It’s the ordinary logic of third-party risk, applied to a vendor category that the economics article argues deserves exactly that treatment: findable, contractable, auditable, and enforceable where our customers are examined.
And the ZDR changes what’s at stake even in the failure case. Under a retention regime, the questions compound quietly — who accesses the stored copies, how long they persist, whose subpoena reaches them, what a breach exposes. Under zero retention, those questions have no object. Nobody can lose what nobody possesses. A contract can be breached; a store that was never created cannot be.
One signature between your pages and the model
There’s a second structural choice hiding inside “we have a ZDR,” and it matters more than it sounds: with whom?
Much of the industry reaches frontier models through intermediaries — cloud marketplaces, aggregators, resale layers. There are legitimate reasons to do that, and some regulated institutions deliberately prefer it, because inference then lives inside a cloud agreement they already govern. But every intermediary is another party in the chain of custody: another contract with its own retention terms, another monitoring regime, another set of infrastructure your pages transit under someone else’s paper. The diligence question stops being “what does the agreement say” and becomes “what do all the agreements say, and do they compose?”
Scout’s relationships are first-party and direct: the entity performing the inference is the entity that signed the retention agreement. That’s why the control-plane article can make its claim in the singular — exactly one processor of document content, transiently, under ZDR — and why the examiner’s questions have one place to land. Not because intermediated channels are unsafe, but because every additional party is another contract to read, and we chose the chain with the fewest signatures on it. One is the fewest.
The notebook question, answered for our traffic
Now the sharpest part of the industry argument: reasoning tokens. When a reasoning model works through a hard problem, it generates intermediate tokens before the final answer — real derived data, billed for, generally not shown. Are those tokens Output, owned and protected like the answer? The industry’s standard agreements were mostly drafted before the question existed, and the skeptics are right that few providers have answered it in a sentence.
Here is our answer, in the two parts it divides into.
First: for the work that touches your documents, the notebook mostly doesn’t exist. Extraction — the reading of actual borrower pages, the overwhelming share of Scout’s model traffic — runs with extended reasoning off. The model is asked to read pages and return specific facts; what comes back is the answer, and no hidden intermediate reasoning is generated, billed, or left behind on anyone’s server. This isn’t a privacy posture we adopted for this article — it’s the economics of the harness doing its other job. A system that never asks the model to reason out loud doesn’t produce a scratchpad. Where Scout does enable reasoning, it is selective and deliberate, a per-call decision made in code — never a default running over every page of every file.
Second: the agreement, read as written. Like most retention agreements in the industry, the defined commitments speak of inputs and outputs; intermediate reasoning is not separately named. We’d rather describe that precisely than pretend a definitions section says something it doesn’t — the skeptics found a real gap in the industry’s paper, and it isn’t closed by wishing. What we can say is what zero retention means in operation: when a request completes, the provider stores nothing — and intermediate computation that was never stored cannot be retained, trained on, or subpoenaed, whatever ownership category it might someday be assigned. We’ll also say plainly what we think the industry should do: providers should state, in a sentence, that intermediate reasoning carries the same non-retention and non-training commitments as input and output. The ambiguity is resolvable, and customers in examined industries are entitled to the sentence.
The alpha never enters the prompt
But the deepest version of the industry worry isn’t about tax returns. It’s about alpha — the idea that what you send a model teaches it your proprietary judgment: how your institution decides, what it tolerates, where its thresholds sit. The fear is that the provider keeps a notebook about how you think.
This is where Scout’s architecture gives an answer that no agreement has to carry alone, because it’s the same rule this site keeps returning to: the model reads; code decides. Your institution’s judgment — the rules, the tolerances, the exception logic, the policy encoded in a recipe — is deterministic code, executed on the workstation. It is never sent to the model. The model is shown pages and asked for facts: the APR on a disclosure, the guarantors on a note. It is never shown what you do with them. There is no prompt containing your credit policy, so there is no reasoning about your credit policy to retain, train on, or leak — under any definition of any term in anyone’s agreement.
What transits, transiently, is borrower documents — genuinely sensitive, and covered by everything above: one direct counterparty, zero retention, no training, US inference. What never transits at all is the thing the alpha argument is actually about. The pages are protected by contract. The judgment is protected by architecture — and architecture doesn’t need a definitions section.
If the question behind the question is why we rent frontier models at all when so much of the system is deterministic, the economics article is the full accounting. The short version: the model’s one job is reading real documents, and reading quality is the one line item where cheap is expensive. The alpha lives in code precisely so that the model can be chosen for how well it reads — not for how much of your judgment you’re willing to send it.
What to ask any vendor, including us
If you’re evaluating an AI product for an examined institution — this one or any other — the argument circulating in the industry compresses into five questions worth asking as written:
- Is the no-training promise a policy or a signed contract, and in whose courts is it enforceable?
- Is retention actually zero, or a deletion window with exceptions?
- How many parties stand between your data and the model, and does each one’s paper say the same thing?
- Does the system generate hidden reasoning over your data — and if so, what covers it?
- Does your proprietary decision logic enter the prompt at all — or does the vendor’s architecture keep it out of the model entirely?
Scout’s answers: contract, zero, one, mostly-none-and-covered, and never. The last one is the one to press every vendor on, because it’s the one no agreement can substitute for. A promise about what happens to the notebook is good. A design in which the notebook about your judgment is never written is better.