What is a Compute Threshold? Why do regulators prefer thresholds over evaluating each model's risk individually?
A compute threshold is the most common practical tool under the broader concept of Compute Governance — using a concrete, quantifiable number (a chip's processing performance, or the total floating-point operations consumed training a given model) as a dividing line. Cross that number, and specific regulatory obligations kick in automatically, without needing a separate risk assessment for that particular model or chip.
Regulators favor threshold-based approaches for reasons tied to compute's own properties: a chip's processing performance and the compute scale a training run consumes are concrete, measurable numbers that are relatively hard to fake — a far cheaper thing to enforce and audit than the subjective, often contentious judgment of "how dangerous is this model," which requires complex capability evaluation to even attempt answering. That's exactly why, even though compute governance spans many different tools, "set a numeric threshold" remains the form most consistently reached for first.
How are compute thresholds actually used in the real world? Can you give examples across different policy layers?
Compute thresholds show up across at least three different policy layers. The first is export control thresholds, used to decide whether a chip can be shipped to a given country or region — the Bureau of Industry and Security's new rule from January 2026 classifies chips with a Total Processing Performance (TPP) at or above 21,000, or memory bandwidth at or above 6,500 GB/s, as fully restricted, while chips below that threshold shift to case-by-case review instead of a default presumption of denial.
The second is training-scale reporting thresholds, used to determine whether a model's developer has to disclose information about that training run to the government — the U.S.'s 2023 executive order requires developers of any model trained above 10^26 floating-point operations to report it; the EU's AI Act sets a similar compute-reporting threshold for general-purpose AI models. The third is performance thresholds used to sort risk tiers rather than simply trigger reporting — OpenAI's own Preparedness Framework, for instance, maps a model's performance on a specific capability dimension (cyber offense, say) reaching a given level directly onto its highest risk category, "Critical," which triggers deployment pauses and extra review procedures the moment it's crossed, not just a disclosure requirement.
Once a Compute Threshold is set, does that solve the problem for good? Are there limitations it can't escape?
No. Compute thresholds face at least two structural limitations. The first is threshold drift caused by efficiency gains: algorithmic efficiency and chip manufacturing both keep improving, meaning the actual compute required to reach a given level of capability keeps falling over time — a threshold set today may, within a few years, end up covering more and more general-purpose chips it was never meant to restrict, or fail to catch a next-generation model that reaches equivalent capability with far less compute. That's exactly why a technical threshold like TPP has to be continuously revised alongside hardware generations — the U.S.'s chip performance thresholds have been revised multiple times since first set in October 2022.
The second limitation is more fundamental: a compute threshold governs whether that much compute was used, not how dangerous the resulting model actually is — which means a model trained below the threshold, but that reaches high-risk capability through some algorithmic innovation, could in principle trigger no reporting or control obligation at all. That's why analysts broadly treat compute thresholds as just one piece of the AI governance toolkit, generally needing to be paired with direct capability evaluation (like the Preparedness Framework risk tiers mentioned above) to close the structural gap where thresholds catch scale but miss innovation.
Why should an ordinary reader care about a technical number like a Compute Threshold — what does it have to do with everyday life or investment decisions?
The specific numbers behind a compute threshold often signal whether a company's future room to grow will be constrained earlier than that company's own model capability scores do. If a company relies heavily on a particular tier of chip, and that tier happens to sit close to a newly set control threshold, any future adjustment to that threshold — tightening or loosening — could directly affect whether the company can keep obtaining the compute it needs to train its next-generation models, which in turn ripples into its business positioning and market valuation.
For a general reader, practical things worth watching include: whether major chip suppliers specifically mention in earnings reports or calls that shipments of a given performance tier are constrained by a threshold in a particular market; whether regulators publicly signal an upcoming threshold adjustment (shifts in how tight or loose a threshold is tend to signal a structural bottleneck in the industry's future compute supply earlier than any individual model release does); and whether a company proactively discloses that a given model or training program had to be reported or paused because it crossed a specific performance threshold. This kind of concrete, verifiable threshold-related information tends to be more useful for gauging a company's actual future position than a bare capability benchmark score.
The Bureau of Industry and Security's rule, published January 15, 2026, set a chip Total Processing Performance (TPP) of 21,000 and memory bandwidth of 6,500 GB/s as the full-restriction threshold, replacing the Biden-era AI Diffusion Rule's country-tier framework. The day before the new rule took effect, the White House separately announced a 25% tariff on advanced computing chips meeting the same performance thresholds — producing an unusual policy combination of loosened export licensing paired with a tightened import tariff, running side by side.