Skip to content
Open to board advisory and board seats: 2H 2026, then CY 2027-2028.
See details →
AI

Country of Origin Is Not a Control

Risk committees ask whether to allow Chinese models as if nationality were a control. Your regulator defines foreign by where the work happens, not the flag.

By Michael YorkAugust 3, 2026 10 min read 2,220 words All AITable of contents

Nationality is not a control. It is a proxy, and the third-party guidance your examiner already works from says in writing that a proxy is not due diligence.

I have sat through this meeting enough times to predict the agenda. Someone brings an open-weight model to a risk committee, the capability case is good, and within four minutes the discussion has collapsed into one question: do we allow Chinese models? The question has somewhere to live, because most third-party questionnaires already carry a country-of-origin field, and a field with an answer in it looks like a finding.

It is not a finding. It is a proxy standing in for determinations nobody made, and it fails in both directions at once: blocking a deployment that would never send a byte outside your VPC, clearing one that quietly ships prompts into a jurisdiction you did not choose because the vendor's letterhead was domestic. I own security and DevOps for a fintech platform serving more than 1,500 financial institutions, so I both write the determination down and operate what it authorizes. Those two jobs disagree with the country field constantly.

Your regulator already defines "foreign" by where the work happens, not by whose flag flies over it

The agencies settled this three years ago, for every third party, in language nobody in that room has read. The June 2023 Interagency Guidance on Third-Party Relationships (opens in new tab) from the Federal Reserve, FDIC, and OCC does tell you to weigh the "use of foreign-based third parties" when you plan a relationship. Then it hangs a footnote off that phrase to define the term, and the footnote is the whole argument: third parties "whose servicing operations are located in a foreign country and subject to the law and jurisdiction of that country."

The footnote spends its two remaining sentences closing both doors. The term "does not include a U.S.-based subsidiary of a foreign firm because its servicing operations are subject to U.S. laws." And it "does include U.S. third parties to the extent that their actual servicing operations are located in or subcontracted to entities domiciled in a foreign country." The regulators' test is where the operations sit and whose law binds them, and the definition exists mainly to rule corporate nationality out of that test in both directions.

The guidance makes the point a second way. The due-diligence section enumerates thirteen lettered considerations (information security, operational resilience, incident reporting, reliance on subcontractors, insurance coverage, and nine more) and country of origin is not one of them. Ownership surfaces once, inside legal and regulatory compliance, as the first of five items: identify beneficial ownership, "whether public or private, foreign, or domestic," then determine whether the third party or its owners are sanctioned by the Office of Foreign Assets Control. That is an instruction to look a name up on a list, not to render a verdict on a flag. Sanctions screening is a real control with a real list, and not the same instrument as a mood.

The guidance also says the quiet part outright: "Relying solely on experience with or prior knowledge of a third party is not an adequate proxy for performing appropriate due diligence." Substitute nationality for prior knowledge and the sentence still holds. What it asks for instead is proportionality, that "the scope and degree of due diligence should be commensurate with the level of risk and complexity of the third-party relationship." Proportionality requires knowing what is in the payload, and the country field tells you nothing about it.

Two Chinese labs, two opposite answers to the only question that matters

Here is where the proxy breaks in public. Read the published policies of two providers a country field treats as identical. DeepSeek's consumer privacy policy (opens in new tab) states that it collects, processes, and stores personal data "in People's Republic of China," lists User Input among the categories it processes, and names among its purposes "to improve and develop the Services and to train and improve our technology, such as our machine learning models and algorithms." It also grants a right to opt out of that training use. That is a self-declared answer on residency and training, and for most regulated payloads, a decline.

Now Alibaba. Alibaba Cloud's own Model Studio documentation, the managed platform through which Qwen is commercially served, answers the same FAQ with "Alibaba Cloud protects data privacy and will never use your data for model training." The platform runs in six regions, among them Singapore, Frankfurt, Tokyo, and US (Virginia), and the docs say the region "determines the access point and data storage location." Its deployment-scope selector goes further than plenty of Western vendors bother to, letting you pick global nodes that explicitly exclude the Chinese mainland. Those are data-flow findings, published by the vendor, that a country field cannot represent.

So the word "Chinese" does not tell you which company receives your prompt, what it does with it, or where it rests. One row merges them and returns one verdict. The mirror image is the failure I have actually had to chase. A domestic vendor with a clean no-training clause can still move your bytes along a path nobody asked about, which is why "we don't train on your data" answers one leg of three. The country field clears that vendor. It clears every vendor whose incorporation documents are in English.

The country field does not predict cost any better than it predicts risk

The origin question smuggles in an economic assumption too, that the Chinese option is the cheap option and the committee is trading risk for a discount. That framing is also wrong, and the evidence is the government's own. In NIST's Center for AI Standards and Innovation evaluation of DeepSeek V4 (opens in new tab), published in May 2026, CAISI filtered out any U.S. model that performed significantly worse or cost significantly more per token, leaving GPT-5.4 mini as the only surviving comparison. On the seven benchmarks it priced, DeepSeek V4 cost less on five, and the full range ran from 53 percent less expensive to 41 percent more expensive.

Sit with the shape of that number rather than its direction. One model, one evaluator, one baseline, and the cost verdict inverts depending on which job you point it at. There is no per-model discount to approve or decline, because the economics belong to the workload rather than the checkpoint, and certainly not to the country. Which is the same reason model selection is a routing and capacity decision rather than a brand decision. The unit you evaluate has to be the lane, because that is the only unit where cost and quality both become measurable.

Replace one field with determinations recorded per lane

So retire the row and put findings in its place. Not policy statements, determinations, each with a date, an owner, and an expiry, made per lane rather than per vendor. This is the form I hand a committee.

  • Job class and data tier. What work does this lane do, and what is the highest classification that can appear in the payload. If the answer is public or synthetic, most of the argument evaporates. If it is member data, it never touches a path you cannot audit, whoever built the model.
  • Resolved model endpoint, not family name. "Qwen" is a family, not a model. Record the served checkpoint and the endpoint serving it, because behavior, license, and hosting terms all vary inside one family and none are attributes of the name.
  • Deployment path, with its own control set. First-party lab API, third-party host serving open weights, or weights you run yourself. Three counterparty sets, three data flows, three bodies of evidence.
  • Where the payload travels, where it rests, and who can read it there. The same three questions I would put to a domestic SaaS vendor, answered with retention windows and subprocessor lists, not adjectives.
  • Weight provenance and license. Who trained this checkpoint, on what, under which license, and what happens to you if the middle answer is contested two years from now.

This has to be per lane because a single lab's consumer app, its enterprise platform, and its downloadable weights are three counterparties wearing one logo, with different answers in every row. One verdict for all three will be wrong about at least two.

Weight provenance is the line item no vendor can satisfy

That last row is the one almost nobody carries, and it deserves the anxiety the country field has been absorbing. On February 23, 2026, Anthropic published its own investigation (opens in new tab) alleging, as one interested company's reading of its own logs and not a court finding, that DeepSeek, Moonshot, and MiniMax ran industrial-scale distillation campaigns through roughly 24,000 fraudulent accounts and more than 16 million exchanges with Claude in breach of its terms, attributed by IP correlation, request metadata, and infrastructure indicators. None of it has been adjudicated, so file it as an allegation and weight it accordingly.

That does not mean ignoring it. It means noticing that your file now holds a question no vendor on either side of the Pacific can close. Distillation itself is ordinary and disclosed: DeepSeek's R1 repository ships six dense distilled models on Qwen and Llama bases, from 1.5B to 70B, fine-tuned, the card says, with 800,000 samples curated with R1. That is authorized self-distillation. The category that matters for a contract is the one where a checkpoint's capabilities may derive from another provider's outputs against that provider's terms, because that is an IP-contamination question with a downstream user on the hook.

You cannot resolve it with an artifact, because the artifact does not exist. I can produce a software bill of materials for any container in our pipeline and diff it against the last release. There is no equivalent for a checkpoint. The nearest thing, a CycloneDX ML-BOM, documents declared datasets, training methodology, and framework configuration, which are assertions by the publisher rather than attestations you can verify against the weights. That is the gap I have complained about before, where an SBOM nobody wires into a gate is compliance cosplay. A bill of materials you cannot check is a claim with better formatting.

Which makes this a contract problem, not a questionnaire problem. Ask for the indemnity. Ask who bears the cost if a checkpoint you built a lane on becomes subject to an injunction, a takedown, or a revoked license eighteen months from now. Ask what the fallback is on that day, then rehearse it. Those are ordinary vendor-risk questions, and a better use of the hour than relitigating geography.

Self-hosting changes who can read the prompt, and nothing else

The last thing the country field does is push people toward a conclusion they have not costed. If the worry is that prompts leave the building, run the weights yourself. But price it honestly, because open weights are a legal review and a capacity plan, not a download.

The GLM-5.2 card on Hugging Face lists 753 billion parameters under an MIT license in BF16 and F32, which at two bytes each is roughly 1.5 terabytes of weights before you serve a token. The Kimi K3 card, whose weights are posted rather than promised, lists 2.8 trillion parameters with MXFP4 weights, and it is not MIT or Apache 2.0. It ships under a bespoke Kimi K3 License whose commercial-use clause requires a separate agreement with Moonshot once the aggregate revenue of the licensee and its affiliates passes 20 million U.S. dollars over any consecutive twelve months. Counsel reads that threshold before your platform team downloads anything, and it has nothing to do with nationality.

Then there is everything self-hosting hands back. It changes who can read the prompt, and nothing about the telemetry your inference server emits, the tool calls your agents make to third parties, the logs you now retain, the authentication you now owe, the patch cadence, or the rollback you execute at two in the morning. You did not remove a vendor from the diagram. You became one, with its own concentration questions once you trace the column down to the accelerators.

Delete the row, keep the determinations

The practical ask is small and I would run it this quarter. Delete the country-of-origin field from the AI section of your third-party questionnaire, because it produces a finding your policy cannot act on while giving a committee false confidence that a determination was made. Replace it with the five rows above, per lane, each with a named owner and a review date. Keep OFAC screening and any applicable export-control review, which are real controls with real lists, and stop letting them stand in for a data-flow analysis.

Then hold every candidate to the identical standard, incumbents included. The discipline only works if it is symmetric. The moment your form asks harder questions of a Beijing lab than a San Francisco one, you are running the proxy again with extra steps, and the real reason the domestic vendor passed is that nobody made it answer.

I would genuinely like to hear how other regulated shops are writing this determination down, and specifically whether anyone has gotten weight-provenance indemnity language into a model contract that survived the vendor's redlines. That is the row I still cannot close, and I would rather learn it from someone else's negotiation than my own.

AIAI GovernanceVendor RiskModel RiskFintech