The Price of Intelligence: Open-Source AI, Compute Controls, and the State Balance Sheet in the AGI Era

From Power, Scaling, and Compilers to Self-Improvement: Possible Endgames for U.S.–China AI Competition

This article discusses developments up to July 24, 2026. For convenience, this text uses the term “open-source AI,” although most so-called open models are more accurately open weights. Training data, full training code, data-cleaning pipelines, and training recipes are usually not fully public, so they are not equivalent to traditional open-source software.(hai.stanford.edu)

1. AI Is Both a Genuine Technological Shift and a Genuine Asset Bubble

There are now two seemingly contradictory narratives about AI.

One view is that humanity is entering an era of machine intelligence and that existing investment is still not enough. The other is that tech companies are using the AGI narrative to capitalize unrealized future profits early, creating a new technology bubble.

Both can be true at the same time.

The railway was real, and the railway bubble was real. The internet was real, and the internet bubble was real. Whether AI has technical value and whether current valuations, capex, and financing structures are sustainable are separate questions.

Using an approach we can call “balance-sheet geopolitics,” it is easier to understand what is happening if we set aside corporate vision statements, national slogans, and ideology, and treat AI as a balance sheet. In discussing national competition, the focus is not a single industry but who can establish new collateral, sustain cash flow, and prevent legacy liabilities from exploding first. AI is becoming new collateral that both the U.S. and China are trying to build, though what each country wants to collateralize differs.(yuzhes)

On AI’s asset side are models, algorithms, data, talent, chips, memory, networking, data centers, power grids, generation capacity, industrial use cases, and global users.

On the liability side are large capital expenditures, chip depreciation, power contracts, supply-chain dependence, technology iteration risk, unemployment, income distribution issues, and a harder-to-value liability: if a rival obtains a system with self-improvement capability first, the laggard may lose more than market share.

Cash flow has three layers.

The first layer is whether model companies can make money through subscriptions, APIs, and enterprise services.

The second is whether society can raise productivity, company profits, and tax revenue through AI.

The third is whether a country can turn AI into scientific, industrial, intelligence, defense, and international rule-making power.

The most reliable cash flow remains concentrated today in chips, cloud services, and capital markets. As for the massive future profits AI could bring in an AGI era, most of that is still expectation.

The Bank for International Settlements estimated that the top five hyperscale cloud providers would spend more than $1 trillion on AI-related capex across 2025 to 2026. Those expenditures have already begun to outpace growth in company profits and free cash flow, and some firms are increasing debt financing. BIS also notes that compute purchasing, model-company investment, and cloud-service contracts are entering more complex interdependence and circular financing.(Bank for International Settlements)

That is the most concrete part of an AI bubble.

Companies are not necessarily overinvesting because management has lost rationality; they are facing a contest-style game. If, in the end, only a few models and cloud platforms win the main market, any firm that cuts investment too early may forfeit its entire future. A single company can make a rational decision while the industry as a whole can still overinvest.

The logic is even starker for states.

A data center might not produce a high standalone financial return, but if AGI could alter military, scientific, and governance capability, building it is no longer merely a commercial investment—it becomes a national security option. Even if that option is never fully exercised, governments and firms are unlikely to voluntarily step back during competition.

So AI bubbles and the “AI national destiny” race are not contradictory. Precisely because this is seen as a national destiny contest, the bubble can persist longer than in ordinary industries.

2. Objective Reality: The Gap Is Narrowing, but AGI Has Not Appeared

As of 2026, the first key fact on model capability is that the public gap between U.S. and Chinese models has shifted from a generational gap to a dynamic gap.

The Stanford AI Index says that by March 2026, the gap between the top U.S. and Chinese models on its aggregated indicator was about 2.7%, and since 2025 leadership has shifted multiple times. The U.S. still produces more frontier models and has more high-impact patents; China performs strongly on paper volume, citation count, patent counts, and industrial robot installations.(hai.stanford.edu)

The second fact is that open-weight models are no longer far from the closed frontier. A July 2026 cyber capability benchmark by the UK AI Security Institute found that leading open-weight models lag closed frontier models by about 4 to 7 months in this area. While this finding cannot be projected to all capabilities, it does indicate the monopolistic window for closed-source firms is shrinking.(AI Security Institute)

The third fact is that capability remains jagged. Models can be very strong in math, code, and scientific benchmarks, yet still fail in simple perception, long-running execution, and real computer operation. Stanford’s 2026 report shows AI agents still fail on about one third of certain structured computing tasks.(hai.stanford.edu)

This means current AI is already enough to reshape a lot of software, knowledge work, and enterprise workflows, but it is not yet a reliable AGI that can independently run entire organizations and execute long-horizon research programs without supervision.

The fourth fact is that China’s open models are moving from “catching up” toward ecosystem diffusion. Hugging Face statistics from 2026 show Chinese models accounted for about 41% of model downloads on its platform in 2025, exceeding U.S. models in both monthly and cumulative downloads. Download counts cannot be mapped directly to production deployment, but they do indicate shifts in developer attention, derivative model activity, and technology trajectories.(Hugging Face)

So the current picture is not one where the U.S. has already achieved AGI, or where China has fully closed the hardware gap. A more accurate description is:

Model capabilities are converging rapidly, while compute and infrastructure still show clear divergence, and the spread of an open ecosystem is starting to challenge the pricing power of the closed ecosystem.

3. What Matters Is Not Chip Count, but Realizable AI Capability

When discussing AI competition, people often treat chip numbers, PFLOPS, or model leaderboards as proxies for national power. They matter, but none can explain the outcome alone.

A country’s realizable AI capability can be approximated as:

Realizable AI capability = Effective compute × Algorithm efficiency × Iteration speed × Deployment density × Social absorption capacity

And effective compute is not the same as the chip’s stated peak:

Effective compute = Peak compute × Memory efficiency × Interconnect efficiency × Software utilization × Cluster stability × Power availability

A chip specification sheet tells us the theoretical limit, not how many training runs a cluster can truly complete in a year, how many reliable tokens it generates, or how many valuable tasks it solves.

That is why “how many PFLOPS a country has” is often a misleading metric.

As of the end of March 2026, official Chinese figures put total intelligent-compute scale at roughly 1.88 million PFLOPS, with about 1.37 million PFLOPS connected to the national monitoring and dispatch platform, about 72% of the total. Over 80% is concentrated in eight national hub nodes. The number shows rapid expansion, but PFLOPS from different chips, precisions, memory designs, interconnects, and software environments do not add linearly, and they cannot be compared one-to-one against foreign GPU clusters.(nda.gov.cn)

The real issue is not whether China has compute, but whether it has enough frontier-effective compute.

Frontier-effective compute means stable support for large-scale synchronous training, sufficient high-speed memory and interconnect bandwidth, mature software stacks, controlled failure rates, and the ability for researchers to spend less time on drivers, communication, and compatibility problems.

That is what “talent is sufficient, compute is constrained” means more precisely.

It does not mean China has no bottlenecks at every talent layer. Advanced manufacturing equipment, EDA, HBM, packaging, compilers, hyper-large cluster operations, and hardware-software co-design still require long-term accumulation. It does mean China already has enough model researchers and engineering teams to produce near-frontier results under constrained resources, but not enough compute headroom to run all valuable research tracks simultaneously.

What computelimit cuts first is not the size of a single final training run, but the width of experimentation.

A lab with a huge cluster can run different architectures, data mixes, reinforcement-learning methods, and agent frameworks in parallel, and keep many failed lines as part of exploration. Teams with tight resources must drop candidate options earlier, rely more on judgment, and do fewer large-scale ablation studies.

What reaches the public may just be two models with similar benchmarks. What is invisible is how many experiments each side ran to get there, how many failures they tolerated, and how many fallback routes remain for next rounds.

In frontier R&D, the core metric is not one-off training cost, but how many high-quality learning cycles can be completed each calendar year.

4. Why China Needs Open-Source AI

China’s support for open models should not be explained simply as techno-idealism, nor entirely as government command.

It first emerged from developer culture, competitive pressures among firms, and practical choices under compute constraints, then gradually became part of state industrial strategy. Stanford’s study of China’s open-weight ecosystem also finds state support concentrated in talent, infrastructure, policy environment, and open-ecosystem building, and notes that DeepSeek’s early success cannot be explained simply by direct state subsidies.(hai.stanford.edu)

But China has now clearly elevated openness into explicit policy. The 2025 “AI+” document proposed building a globally oriented open-source technology stack, including open models, tools, and datasets. The 2026 “AI+Manufacturing” policy further requires building a globally influential open ecosystem and promoting open models in industrial applications.(cac.gov.cn)

There are several structural reasons behind this choice.

1. China Cannot Replicate the U.S. Advantage in Closed-Model Rents

U.S. AI firms’ asset base includes global cloud platforms, USD capital markets, enterprise software channels, chip design capabilities, and international customers.

A U.S. firm can turn model capability into API revenue, convert expectations into valuation, and then use valuation and cash flow to build next-generation data centers. Closed models, cloud platforms, and capital markets can form positive feedback loops.

Even if a Chinese firm trains a model of comparable quality, it may not secure the same global subscription revenue, valuation multiples, and enterprise software control.

If Chinese firms package models as closed in the same way, they are effectively competing on terrain where the U.S. has the strongest advantage.

Open weights alter the metric. The competition becomes not only who can charge the highest price per million tokens, but also whose model is adopted by more developers, governments, devices, and industries.

2. Openness Is a Commoditization Strategy at the Model Layer

Valuations for frontier U.S. AI imply a key assumption: advanced intelligence remains scarce for a long time, allowing leading models to charge high rents.

Every time a Chinese firm releases an open model close to a closed-source flagship—cheaper and locally deployable—that assumption weakens.

It may not earn the most directly from the model itself, but it can generate broader strategic returns:

It can lower global model prices, shorten the excess-profit window for closed firms, reduce AI costs for domestic enterprises, and shift value from the model layer into servers, energy, robotics, automobiles, factories, and industrial software.

This is a grounded industrial strategy. China’s strongest complementary assets are not global enterprise software subscriptions; they are manufacturing, engineering capacity, power equipment, telecom equipment, robotics, cars, and consumer electronics.

For this kind of economy, the cheaper the model itself, the more other sectors can benefit.

From this perspective, China does not need model firms to capture the highest global profits. If open models lower intelligence costs across the industrial system, a lower model margin can be exchanged for a larger total surplus in manufacturing.

3. Open Models Can Turn Global Developers into External R&D Resources

In a closed-model system, adaptation, quantization, inference optimization, and industry-specific fine-tuning are mostly handled internally by the model company.

With open weights, global developers contribute hardware adaptation, low-bit quantization, inference frameworks, edge deployment, and industry fine-tuning. The model owner exchanges weights for distributed global engineering work.

This is especially important for the side with tighter compute. China does not need to complete every application and hardware adaptation alone, and can shift part of experimentation and deployment cost to the open ecosystem.

More importantly, open models generate many derivative models and tools. They can gradually form de facto standards so API formats, model architectures, inference frameworks, and fine-tuning methods evolve around Chinese models.

4. Openness Is Also a Sovereignty Product

Many countries want advanced AI but do not want to permanently hand governance, healthcare, finance, and industrial data to foreign cloud providers.

Open weights let them run models locally, stay within domestic law, fine-tune in local languages, and continue operating if international relations deteriorate.

So the U.S. and China are offering different dependency models.

The U.S. government has explicitly called for exporting the full American AI stack to allies and partners, including hardware, data systems, models, applications, security systems, and standards. It is a packaged, controlled, continuously updatable full-stack service.(The White House)

China’s open models are closer to downloadable, modifiable, and locally deployable base capabilities. They may be less complete than a full U.S. stack, but more compelling on price, editability, and data sovereignty.

Many medium-sized countries will not choose AI systems by ideology. They will weigh who is cheaper, who allows data to stay local, who can provide financing and energy infrastructure, and who is less likely to abruptly cut API access due to diplomatic conflict.

5. Open Models Could End Up Reinforcing the U.S. Stack

There is an often-missed point.

Even if a Chinese open model is widely adopted globally, if it mainly runs on NVIDIA GPUs, the CUDA ecosystem, and U.S. cloud platforms, it can compress U.S. model-company profits while still expanding U.S. chip and cloud-service revenues.

So download volume is not equivalent to tech sovereignty.

To become a real substitute order, open models must pair with domestic hardware, compilers, inference stacks, cloud services, data-center financing, energy plans, and overseas maintenance capabilities.

Otherwise, China may provide free models while the U.S. still captures compute rent.

That is why TileLang, CANN, domestic interconnect protocols, inference engines, and distributed file systems appear less headline-grabbing than model releases but may have longer strategic significance.

The model decides what users want to run; compilers and runtimes decide which hardware stack the model actually runs on.

If China can move an open model across domestic chips with low cost while keeping utilization high, then open model adoption can actually expand the hardware ecosystem. If every domestic chip needs heavy manual tuning, developers will eventually return to the mature CUDA environment.

So a true AI open strategy is not publishing a weight file; it is lowering migration cost across an entire replacement stack.

6. When Will China Close the Compute Gap?

There is no single answer, because “closing the gap” includes at least four distinct things:

Catching up in model capability, achieving domestic inference self-sufficiency, having multiple frontier training clusters, and approaching the U.S. in total compute, energy efficiency, software, and cost.

They will not arrive together.

Model capability: dynamic convergence is already happening

From public model results, China is no longer stably a full generation behind; rather, within a few months, it catches up, leads locally, and is then pulled back by a newer generation.

That suggests the compute gap does not map one-to-one to capability gaps. Algorithmic efficiency, systems optimization, data quality, and open ecosystems are offsetting some hardware disadvantage.(hai.stanford.edu)

But that does not mean compute stops mattering.

When efficient algorithms are public, U.S. labs can also adopt them. For the side with more compute, efficiency gains do not lower the incentive to invest; they allow larger models and more reinforcement-learning environments under the same budget.

Progress in efficiency can help followers, but it can also expand the leader’s absolute exploration space.

Domestic inference self-sufficiency: likely to arrive first

Inference workloads are easier to quantify, distill, shard, and deploy heterogeneously. Many requests are independent and can be scheduled across different machines and regions.

So stacking more lower-performance chips is often more effective for inference than for training.

Frontier training requires thousands, even tens of thousands, of chips synchronized in the same step. Any interconnect congestion, memory bottleneck, or node failure can slow the whole cluster. Inference can split user requests across clusters; one weak node does not necessarily block all others.

So China’s compute catch-up is likely to be asymmetric:

First, it solves most domestic inference needs; then a handful of frontier training clusters emerge; and only later can it approach U.S. total compute stock and full software ecosystem.

Given current domestic chip roadmaps, low-precision inference, and a national compute network, I believe a strategic level of domestic inference self-sufficiency in the late 2020s is a plausible scenario.

Here, “self-sufficiency” means runnable, scalable, and resilient to external supply cuts—not necessarily leading in per-watt efficiency, maintenance cost, or developer experience across the board.

Frontier training clusters: interconnect matters, not just chips

China is trying to compensate for single-chip limits through massive interconnect scaling. Huawei’s roadmap states that under advanced-process constraints, it can connect more chips through supernodes and clustering, and plans to launch new product generations through 2026–2028. Because these numbers include corporate roadmaps and promised future performance, actual on-time mass production and stable utilization still require deployment validation.(huawei.com)

Scale is not fake, but it is not free.

More chips means more traffic, higher fault probability, and more complex needs for optics, switches, power, and cooling. With MoE training, large amounts of data must move between experts and nodes. DeepSeek-V3’s technical report says cross-node expert parallelism saw near-one-to-one compute-to-communication overhead, so it used DualPipe, computation-communication overlap, and custom communication kernels to reduce losses.(arXiv)

So 10,000 weaker chips are not naturally equivalent to several thousand stronger chips. That substitution holds only when interconnect, scheduling, low-precision compute, and software utilization are all sufficiently high.

My scenario view is this:

By the late 2020s, China has a reasonable chance of having several domestic clusters capable of training frontier models. As for completely catching up to the U.S. on total compute, per-watt performance, software maturity, and cost per token, it could take longer—or remain a dynamic near-pass.

More important, China may not need absolute parity.

If China can keep model capability within a few months and rapidly diffuse results through open weights, U.S. hardware advantages may persist but may not convert into a multi-year model monopoly.

Export controls will still matter, but their role is likely to shift from “blocking China from obtaining capability” to “raising the cost for China to obtain capability.”

U.S. chip policy is also not pure blockade. In January 2026, the U.S. began allowing case-by-case approvals for H200, MI325X, and similar products under conditions, while still tightening licensing requirements for entities linked to China to obtain advanced compute products globally.(bis.gov)

This suggests chip control is becoming a tunable strategic tariff rather than a simple zero-sum on/off switch.

People tend to talk about compute and only mention GPUs. A real AI system, however, includes at least chips, HBM, networking, storage, power, cooling, data, and software.

In the short term, the scarcest inputs remain advanced chips, HBM, advanced packaging, and high-speed interconnect.

In the medium term, power and the grid become increasingly critical.

The IEA expects data-center electricity consumption to rise from about 485 TWh in 2025 to about 950 TWh by 2030, with AI-dedicated centers growing faster. Data centers are highly concentrated in few regions; even if a country has abundant total generation, local substations, transmission lines, gas pipelines, cooling systems, and grid-connection capacity can still be bottlenecks.(IEA)

China’s case is particularly notable.

In 2025, China’s compute-facility electricity use was about 170 billion kWh, or 170 TWh, up about 30% year-on-year, about 1.6% of total social electricity consumption. The National Energy Administration expects that share could rise to 6% by 2030. It also judges that during the 14th Five-Year Plan, inference load would exceed training load and become the main compute electricity demand.(nea.gov.cn)

That changes the logic of China’s compute layout.

Training can queue, can move west, can use cheaper electricity, and is less sensitive to user latency.

Inference is different. Massive agents, industrial control, autonomous driving, financial services, and real-time applications need low latency and have larger demand swings. Inference demand will stay closer to population, enterprise, and industrial centers.

So “East-West Computing” will not disappear, but it may shift from one-way migration to a layered structure:

Western regions handle large-scale training, batch processing, and non-real-time inference; eastern regions handle real-time inference, industry data, and edge intelligence, while the national network dynamically balances cost, latency, and energy.

The National Data Administration has already included compute-power co-planning, nationwide compute monitoring, and cloud-edge-terminal coordination in follow-up construction.(nda.gov.cn)

Power advantage is more than generation volume

China can rapidly build generation, grid, transformer infrastructure, and data centers, and can better coordinate western energy with eastern demand.

But AI needs not just annual total generation; it needs predictable, schedulable, data-center-proximate power with stable 24/7 supply.

Solar and wind can lower average prices but do not alone ensure uninterrupted training clusters. Storage, fossil generation, nuclear, hydro, long-distance transmission, and load dispatch remain essential.

Water is another constraint. Some energy-rich regions do not necessarily have sufficient cooling water. Liquid cooling can reduce some limits but raises equipment, maintenance, and supply-chain requirements.

The U.S. may face the opposite problem. It has advanced chips, cloud platforms, gas, and capital, but many data centers are still clustered near existing hubs, making them vulnerable to local grid, transformer, permitting, and transmission-cycle limits. The IEA notes that about half of U.S. data centers under construction remain in existing large clusters, potentially deepening local grid bottlenecks.(IEA)

So in the future, AI competition may shift from measuring raw FLOPS to:

How many reliable, valuable, user-acceptable intelligence tasks a kilowatt-hour can produce.

Efficiency gains can increase total electricity use

Cheaper models reduce per-token electricity use, but they also raise token demand.

As inference prices drop, firms embed AI into search, support, coding, manufacturing, transport, and personal devices. Agents can plan repeatedly, call tools, and verify outcomes to complete one task, so a user may consume many more tokens per day.

This is a classic Jevons effect: improving efficiency can increase total resource use.

So DeepSeek-style efficiency innovation does not necessarily mean the AI industry will need fewer chips and less electricity. It may instead mean AI becomes cheap enough to enter every sector, generating much larger aggregate demand.

8. Bypass Will Not Come from One Magic Invention, but from Multi-Layer Co-Optimization

Under advanced chip constraints, a common misunderstanding is to wait for one technology to replace top GPUs directly.

More likely, bypass happens across multiple layers.

Layer 1: Algorithmic bypass

MoE activates only part of parameters per token, reducing compute. MLA compresses KV cache and lowers inference memory. Low-precision training and inference reduce bandwidth and storage requirements. Sparse attention, distillation, quantization, caching, and speculative decoding further cut per-task cost.

DeepSeek-V3 uses MoE, MLA, FP8 training, computation-communication overlap, and multi-token prediction. Its report says formal training used about 2,788,000 H800 GPU-hours, roughly 5.576millionat5.576 million at 2/hour, but that figure explicitly excludes earlier architecture research, ablation experiments, and data-development cost.(arXiv)

So what matters is not that “it was trained for only $5.6 million”; rather, it demonstrates that algorithm-system-hardware co-design can materially raise utilization of existing chips.

Layer 2: Systems and software bypass

TileLang tries to move high-performance operator development from hand-written low-level code toward a more expressive, portable block-level programming model. It separates data flow and scheduling spaces so compilers handle more thread binding, layout, pipelining, and tensorization work.(arXiv)

DeepSeek then open-sourced TileKernels based on TileLang, plus DeepGEMM, FlashMLA, and 3FS, which handle high-performance matrix math, attention kernels, and distributed storage respectively.(GitHub)

TileLang does not replace advanced process, nor does it turn low-performance chips into high-performance chips by itself.

What it does change is another key point: how many engineers, how much time, and how much proprietary knowledge are required to make a non-mainstream chip performably usable.

If compilers and DSLs can reduce adaptation cost, the entry barrier for hardware ecosystems falls. Theoretical peak does not change, but more theoretical compute can become effective compute.

At this layer, compilers can become strategic infrastructure as important as chips.

Layer 3: Cluster-level bypass

When single-chip performance lags, larger memory, higher-bandwidth interconnects, network accelerators, and supernodes can assemble more chips into a logical machine.

The core challenge here is not merely card placement in racks; it is preventing communication, memory, and fault-recovery costs from consuming scale gains.

Future designs may include more AI-native interconnect protocols, optical links, network computing, compute-memory co-design, and near-memory computing. DeepSeek’s technical report already proposes offloading some communication to network accelerators and performing partial precision conversion near HBM to reduce data movement.(arXiv)

AI hardware evolution may therefore shift from “each generation gets a faster chip” to “the data center itself is a machine.”

Layer 4: Heterogeneous bypass

Not every workload needs the same expensive GPU.

Training, prefill, decoding, recommendation, vision, robotics control, and small-model inference have different requirements for compute, memory, and latency.

One plausible architecture is:

Most advanced chips handle a small set of critical training and difficult inference tasks, domestic general accelerators handle large-scale ordinary inference, specialized ASICs handle stable workloads, edge chips run local models, and CPUs/storage handle retrieval, preprocessing, and some sparse compute.

This is more realistic than expecting one chip to cover all workloads.

Huawei already treats prefill and decode as different hardware requirements and designs different memory and bandwidth allocations in its roadmap. Its impact still needs production validation, but such a specialization strategy is likely to become an industry direction.(huawei.com)

Layer 5: AI helps design AI infrastructure

There are already early cases of AI optimizing algorithms, code, data-center scheduling, and chip design. Google DeepMind’s AlphaEvolve has been used to improve data-center efficiency, chip design, and AI training itself, including training a model to support AlphaEvolve.(Google DeepMind)

This is not recursive AGI yet, but it highlights a key route:

AI first optimizes kernels and compilers, lowering training costs; lower training cost supports stronger models; stronger models then optimize systems software and chip design.

If each improvement can be reliably validated, this loop can compound.

9. Iteration Speed May Decide the Winner More Than Any One Generation

Model capability is stock. Iteration speed is throughput.

A lab leading today does not guarantee leadership a year later. What matters is how quickly it can complete this cycle:

Propose a hypothesis, write code, train a model, evaluate results, deploy to users, collect feedback, then revise algorithms and systems.

Compute determines how many experiments can run at once; software tools determine engineering time per experiment; open ecosystems determine how many external developers participate; real users determine how rich feedback data are.

U.S. frontier labs have scale advantages in clusters, platforms, global users, and vertical integration. China’s open ecosystem has faster diffusion, broad participation, rich industrial usage, and rapid replication plus improvement of public methods by multiple firms.

Open weights effectively expand a single company’s iteration loop into parallel search across the ecosystem.

This is the most difficult-to-measure part of China’s open strategy. It gives up some short-term exclusivity in exchange for broader engineering feedback and faster diffusion.

Open models also diffuse China’s innovation to U.S. labs. Public techniques on low-precision training, MoE scheduling, and inference optimization can be adopted by better-resourced competitors as well.

So openness does not automatically shrink the gap. It raises industry-wide iteration speed and forces China to innovate continuously in subsequent rounds.

Once open ecosystems stagnate at replication rather than sustained original breakthroughs, that advantage disappears quickly.

10. How Self-Improving AGI Would Change This Competition

“Self-improving” AGI actually spans very different levels.

The lowest level modifies prompts, tool configuration, search strategies, and agent frameworks.

Next is automated code writing, GPU-kernel optimization, novel algorithm discovery, and training-data generation.

Higher still is modifying model weights, designing reinforcement pipelines, and selecting the next training target.

At the top are tasks like chip design, building factories, expanding power supply, and changing organizational systems.

There are already early signs of the first several levels. A 2026 SIA study tries to have one system modify both agent frameworks and model weights, reporting improvements in legal classification, GPU-kernel optimization, and biological data processing. The study still states openly that discovering how to continuously improve one’s own system remains an open problem.(arXiv)

True recursive acceleration requires several conditions simultaneously:

The system must propose useful improvements; improvements must be reliably validated; training and deployment costs must be below returns; and the next improvement cycle must not be interrupted by model errors, objective drift, or infrastructure shortages.

The most easily overlooked condition is validation.

Code, mathematics, chip layout, and some experimental sciences provide clear feedback—whether code runs, performance improves, and proofs hold can be judged relatively quickly.

Social governance, strategic decision-making, education, and organizational management do not have equally clear evaluation functions. A system may propose highly confident complex policies, yet it can take years to observe outcomes.

So even if AGI appears, the first domains to undergo rapid self-improvement are likely code, algorithms, cybersecurity, chip design, materials search, and experimental planning—not whole-society synchronized acceleration.

The physical world limits the speed of “intelligence explosion”

An AGI can modify code in minutes, but cannot build a fab, power plant, or transmission network in minutes.

It may design better chips, but lithography, equipment installation, and ramping to volume production still take time.

It can optimize data-center scheduling, but cannot conjure transformers, HBM, and cooling equipment out of nowhere.

Hence, AGI self-improvement may unfold at two speeds:

Fast acceleration in software layers and slow follow-through in physical layers.

That gives early investment in power, networks, chips, and data centers special value. They become real-world interfaces for future AGI.

From this angle, today’s AI capex—even if partly excessive—is not purely irrational. Firms and states are buying an option: if algorithmic breakthroughs arrive, they will have enough physical resources to scale quickly.

But this also means being first to AGI does not automatically mean winning in the long run.

If a leader has only a model, with insufficient power, chips, and deployment systems, that lead may not scale.

Conversely, if it holds large idle compute, mature automated R&D systems, and reliable safety boundaries, a small algorithmic edge can compound through sustained iteration.

So the genuinely dangerous phase of AGI competition may not be the day a company claims “AGI is achieved,” but the moment a system begins reliably accelerating development of the next generation.

11. Is China’s Open-Source AI the Biggest Threat to the Closed AI Order?

Possibly, but the first threat is not the existence of closed models itself, but the profit model of closed firms.

Closed frontier models can still keep advantages in reliability, complex agents, high-value enterprise use cases, safety accountability, and full-stack user experience.

The real risk is if a stable market expectation forms that:

Capabilities that are expensive and closed today will be replicated by cheap open models within months.

Then companies will avoid building long-lived systems tightly tied to one API, and investors will reassess how long high model profits can persist.

In that case, closed firms may still own the most advanced models but fail to secure long-duration monopoly rents sufficient to cover massive capex.

That directly affects the valuation foundation of U.S. AI assets.

The strategic effect of open models is not necessarily to take all users; it is to shorten the duration of high-priced access. If the window contracts from years to months, many commercial models need to be reworked.

For China’s open-source AI to become a challenge to the broader U.S. tech order, three conversions are still required:

First, move from open models to open and usable domestic software stacks.

Second, move from those stacks to stable, low-cost domestic chips and cloud services.

Third, move from domestic technical stacks to full-stack solutions that can be financed, built, maintained, and upgraded abroad.

Only then would Chinese open models become more than inexpensive software on U.S. GPUs and become an alternative path of international dependency.

The U.S. already recognizes the competition is not only model-performance competition. Its official AI export plan explicitly emphasizes coordinated export of hardware, models, software, applications, security, and standards.(The White House)

China’s real response is not to publish a few more models, but to build a loop among “open models, domestic compute, low-cost power, industrial deployment, and overseas infrastructure.”

12. What Whole-of-State Mobilization Can Solve—and What Problems It Can Create

It is inaccurate to describe U.S.–China competition simply as “China’s mobilized-state model versus U.S. free markets.”

The U.S. is also building a mobilization framework, only organized differently.

It channels investment through large technology firms and capital markets, uses export controls to constrain rivals, speeds data-center and power deployment through policy, and then uses diplomacy to export a full stack to allies.

That is a financialized, corporatized mobilization model.

China, by contrast, more effectively builds infrastructure through planning, state banking, local governments, grid enterprises, telecom operators, and manufacturing supply chains.

That is an administrative, industrial mobilization model.

Both have strengths and typical failure modes.

The U.S. model can quickly turn large future expectations into actual capex, but it can also create overinvestment through market competition, valuation dynamics, and debt buildup. BIS’s concerns about AI capex are precisely a risk for this mode.(Bank for International Settlements)

China’s model can build grids, compute hubs, and domestic supply chains even when returns are unclear, but it may produce duplication, high nominal compute, and low utilization.

The National Data Administration’s push for national compute monitoring, unified dispatch, and a “single ledger” indicates that pure capacity building is no longer enough; the next stage is matching demand and improving utilization.(nda.gov.cn)

The mobilized-state model is best at clearly defined, engineerable goals—generation, transmission, data centers, chip replacement, interconnect standards, and procurement.

It is not necessarily best at selecting which model architecture will produce the next breakthrough. Frontier research needs tolerated failures, preserved heterodoxy, and avoiding all teams optimizing to the same narrow KPI set.

The ideal is not government direct selection of every model; instead, it is public support for power, compute, talent, basic research, and open ecosystems, while multiple teams explore in parallel.

In other words, infrastructure can be centrally coordinated while research paths should remain somewhat decentralized.

13. What Remains After an AI Bubble Bursts

Even if an AI bubble bursts, AI will not disappear.

But the railway and internet analogies do not map mechanically.

Power grids, generation assets, land, fiber, substations, and cooling systems can last decades.

GPU, servers, and specialized accelerators depreciate faster. Some data centers designed for specific chips and interconnects may not migrate cheaply to the next hardware generation.

So AI overinvestment leaves asset quality that differs by type.

The most durable assets are power, networks, industrial capability, and talent.

The easiest assets to impair are high-cost purchases with low utilization and fast obsolescence.

If a bubble bursts before AGI, the U.S. may face shocks through tech-stock declines, corporate debt, private credit, and household wealth effects.

China may show a different pattern: underperforming regional compute projects, low utilization, loan rollovers, and continued support through state systems.

Both countries could socialize private or local losses because AI has been defined as strategic capability.

So a burst need not end competition; it may make competition more state-driven.

14. Should AGI Be Open Source?

My view is:

AGI should be widely used, but unvalidated frontier AGI should not be unrestrictedly released as weights.

“Widely used” is not the same as “copyable, modifiable, and anonymously deployable by anyone.”

Permanent AGI Control by a Few Firms Is Not Acceptable

If AGI can replace large amounts of cognitive labor, accelerate research, shape information flows, and participate in national decision-making, the firm controlling it gains more than commercial advantage.

It can shape the price of cognitive work, decide which companies gain advanced capabilities, influence what the public sees, and create structural dependence in government, military, and research institutions.

Closed-model firms can easily justify a permanent capability monopoly with safety claims.

But when safety governance is defined unilaterally by model owners, it gradually becomes private sovereignty.

A society cannot hand over its most important general-purpose productive capability permanently to a handful of boards and expect corporate ethics to serve public allocation.

Immediate Release of Frontier Weights Is Also Not Acceptable

Once model weights are public, recall is nearly impossible.

Users can remove safety controls, perform dangerous fine-tuning, copy the system to unregulated jurisdictions, and connect it to networks, labs, and automation tools.

The UK AI Security Institute found open-weight models have narrowed the cyber-capability lag to closed frontiers to a few months. That means capabilities still API-gated today could quickly move into non-recallable models.(AI Security Institute)

The U.S. NTIA assessment of open-weight models also does not offer a simple pro- or con-for-open decision. Open weights can improve competition, research, auditing, and local deployment, while also increasing irreversible diffusion and dual-use risk.(NTIA)

So the real split is not “open versus closed,” but five distinct forms of openness:

Whether usage rights are open, whether audit rights are open, whether methods are open, whether model weights are open, and whether governance is open.

A plausible institutional design is:

Frontier capability is broadly available as regulated compute services with price and access rules under public constraints; model safety assessments, incidents, and governance procedures undergo independent auditing; lower-risk or lower-uncertainty older models release weights with delay; frontier weights are held jointly by multiple independent actors so no single company or government can alone decide release, shutdown, or high-risk use.

This can be understood as a form of public AI infrastructure.

The public gets wide access, while raw control over high-risk systems is bounded by multiple authorities.

It still faces serious problems: regulators can be captured by industry, states may refuse mutual trust, and capability evaluation can fail. But it avoids two extremes:

One extreme is permanent private monopoly over humanity’s most powerful intelligence, the other is uploading network-capable, biological, and self-replicating systems directly to a network anyone can download.

A reasonable AGI equilibrium is likely neither full private ownership nor fully anonymous diffusion, but broad access, bounded possession, independent audit, and phased openness.

15. Possible Endgames

What follows are not certain forecasts, but condition branches inferred from the structure above.

Outcome 1: Closed Frontier, Delayed Opening

This is the scenario I think is most likely in the medium term.

The most advanced models stay closed, preserving commercial and security advantages for several months to around a year. The previous generation then opens its weights to expand ecosystems, promote national standards, and erode competitors’ profits.

The U.S. builds a system of advanced chips, closed frontier models, cloud platforms, and tiered access for allies.

China builds near-frontier open models, domestic compute, industrial deployment, and lower-cost infrastructure.

Neither side becomes fully closed or fully open. The actual competition shifts from model leaderboards to power, deployment cost, robotics, industrial data, and global stacks.

Many middle-power countries likely use both systems to avoid single-point dependence on either side.

Outcome 2: China Successfully Commoditizes the Model Layer

If China’s open models remain near frontier for a long period, inference costs keep falling, and domestic chips and compilers become usable, model-level profit margins may fall quickly.

U.S. hyperscalers can still profit; chips, power, and enterprise integration remain valuable, but pure model-company valuations come under pressure.

AI value shifts further into energy, data centers, robotics, manufacturing equipment, proprietary data, and industry channels.

China could benefit from this structure because it has a relatively complete industrial system.

Yet this outcome carries a domestic risk for China as well.

If AI mainly raises productive capacity without simultaneously lifting incomes, consumption, and social security, China may produce more low-cost goods but still lack sufficient domestic demand.

AI can strengthen manufacturing capacity while amplifying excess capacity and export pressure.

Then China may win at model-layer commoditization but still fail to fix its own balance sheet. Higher productivity can intensify, not reduce, global trade friction.

Outcome 3: One Side First Achieves Closed, Self-Improving AGI

If a firm first achieves a system that meaningfully accelerates AI R&D and preserves a one-to-two-year effective secrecy advantage, competition would shift abruptly from market rivalry to a national security event.

Governments would restrict talent flows, control chips and data centers, harden cyber protection, and fold models into military and intelligence systems. A firm might remain nominally private while operating as a quasi-state institution.

If the U.S. gets there first, it may integrate AGI into its existing dollar, cloud, chip, and alliance stack, building tiered intelligence-access regimes.

If China gets there first, it may first show advantage in industry, research, cybersecurity, and governance—but whether that quickly translates into global financial and tech order depends on international trust, capital mobility, and overseas infrastructure.

The most dangerous part of this outcome is that other states cannot be sure of the lead size.

If followers believe the window is closing, they may intensify espionage, chip controls, cyberattacks, and physical sabotage. Larger a lead is may not mean lower short-term strategic instability.

Outcome 4: AGI Weights Leak or Are Deliberately Released

Once true-AGI weights enter the public network, model-layer scarcity could disappear quickly.

Then the scarcest resources become power, chips, robotics, lab equipment, proprietary data, and institutions capable of safely running systems.

Open AGI would reduce monopoly by single firms and states, but significantly raise risks from cyberattacks, bio misuse, automated fraud, and strategic capability diffusion to non-state actors.

In that world, national strength would no longer depend mainly on owning models, but on owning physical execution capacity and institutional resilience.

A small organization may have very strong digital intelligence but no chip factories, robots, or energy. An industrial country can retain advantage through mass physical deployment even if it no longer monopolizes models.

Intelligence becomes cheap; safety and real-world execution become expensive.

Outcome 5: AGI Does Not Arrive Soon; Capex Gets Repriced First

Models keep improving and agents become more useful, yet reliability, long-horizon planning, and data or real-world feedback become barriers that do not easily break.

AI continues changing software, customer support, R&D, and manufacturing, but no recursive intelligence explosion emerges.

In that case, part of current capex may fail to generate expected profits. Model companies consolidate, GPU and data-center assets are impaired, and financial markets reprice.

Open models may become infrastructure for most ordinary tasks, while closed models remain for high-reliability, high-responsibility, and high-value environments.

This is not AI failure so much as a shift from religious narrative to ordinary utility AI: important, widespread, and widely distributed in profit, but no single model company captures the economy’s total rent.

Conclusion: The Final Bottleneck May Not Be Intelligence, But Whether Society Can Absorb It

In the coming years, the easiest thing to overestimate is a single model release; the hardest part to appreciate is the slower-moving layers: plants, transformers, transmission lines, cooling systems, compilers, cluster stability, enterprise processes, and income distribution.

Early AGI release is of course important, but it is not the full answer.

Real leadership requires translating algorithms into stable compute, stable compute into cheap services, cheap services into industrial productivity, and industrial productivity into wages, tax revenue, consumption, and social legitimacy.

If AI raises productive capacity while destroying most people’s income sources, it can increase technical capability while weakening demand foundations and political stability.

If China turns intelligence into cheap means of production without improving household balance sheets, it may gain stronger factories but face deeper demand shortfalls.

If the U.S. keeps intelligence as a high-price asset held by a few firms, it may preserve financial and technical rents, and may also produce unprecedented wealth concentration and private power.

So the ultimate outcome of the AI national contest will not be determined only by who owns the “smartest” model.

A more practical question is:

When blockades come, who can still build and maintain equipment; when power is tight, who can keep clusters running; when model prices fall, whose commercial system still has cash flow; and after productivity rises, whether ordinary people have enough income to buy what machine production creates.

AGI should be used widely across society, but should not be permanently owned by a few firms, nor copied infinitely when no governance capacity exists.

Intelligence can become a public productive force or a new private sovereignty. The dividing line is not model parameters, but who ultimately controls governance, infrastructure, and distribution of returns.