The Geopolitics of Open Weights

Ever since Kimi K3 was released, it really captured the attention of the broader AI and investor community. As open weight closes the gap with closed, frontier models, many are understandably worried about the implications across the value chain. As far as I can tell, while you can legitimately argue about the potential margin erosion in the model layer if open models gain broad adoption, the long-term profitability of other parts of the AI value chain should not be affected even if open models become popular. Gavin Baker eloquently made this argument on X after Kimi’s release. Some excerpt from his post:

“Kimi K3 may be an important inflection point for AI. Potentially negative for Anthropic and OpenAI while being net positive for essentially every other company in the world. I mean that very literally.

A world where there are only 2-3 dominant frontier labs with 90% inference margins is net negative for every other layer while being awesome for those 2-3 labs. Those labs would become monopsonies for power, data centers, semiconductors and hyperscalers and would obviously vertically integrate over time into all those layers while also completely subsuming the application/software layers.

Anything that lowers margins and increases competition at the model layer is good for every other AI layer: power, semiconductors, hyperscalers, neoclouds and yes even software.

An open-source model requires the *exact* same amount of compute to run as a closed frontier model of similar size and architecture. Kimi K3 is roughly the same price as GPT 5.6 Terra on a per token basis, which actually suggests that it is less computationally efficient as I am sure that GPT 5.6 is priced to a higher margin than K3. And given that K3 is a token wastrel, i.e. token inefficient, it is significantly more expensive per task than GPT 5.6 and Grok 4.5, which are much more token efficient. Cost per token and token efficiency (i.e. intelligence density per token) are the drivers of intelligence per unit of cost. The winning AI companies will be those that offer the most intelligence per $ over time.”

The reason I said the “long term” profitability should not be affected even if model layer margin proves to be thin is that the transition from OpenAI and Anthropic’s aggregate value from ~$2 Trillion to “just” a few hundreds billion can potentially be rather challenging for other parts of the value chain in the “short term”. After all, hyperscalers such as Microsoft, Amazon, and Alphabet all have massive backlogs from OpenAI and Anthropic. If frontier models cannot follow through their commitments made to hyperscalers, we may see a temporary moment where demand-supply gap evaporates. It might be only temporary given a more broad adoption of AI seems pretty secular and demand will eventually broaden out even if frontier models’ economics falter. Of course, even a short-term overcapacity can create some pressure on the economics in other parts of the AI value chain and their respective stock prices. To be clear, I’m not quite ready to announce “game over” for model companies; it may still prove to be the case that as Baker pointed out “Claude and ChatGPT products and harnesses may be more important than their models today”. It may be unsatisfying to not be able to infer anything definitively, but it is perhaps more dangerous to conclude more than we currently can. Observing closely but not being able to infer any long-term outcome confidently will likely be the default state for much of the AI value chain for quite some time.

One reason open models can be a real concern for frontier models is not only the overall capability gap between closed and open models is diminishing, they may simply be more performant due to their lack of guardrails compared to closed models, especially in certain high value work such as cybersecurity. I would highlight the following post from Guillermo Rauch, CEO of Vercel (emphasis mine):

“Based on internal evals:

Kimi K3 is top-tier at cybersecurity
There is chatter on X that Moonshot benchmark-overfit. These are stealth evals. Model has raw IQ.

Sol is a leap ahead in cyber capability
At a significantly higher cost, but quite remarkable still.

Fable refuses everything
We couldn’t get it to complete the run at all.
What’s interesting is that Sol in comparison was much more open to helping with defensive cyber hardening

TL;DR: frontier, open-weight cybersecurity capability is here.

Incidentally, I’m very bullish on cybersecurity as one of the best benchmarks for superintelligence. The “IQ test” of software engineering.

The best engineers I’ve worked with in my career have usually had a deep background or interest in security.

It’s actually easy for a model to “one-shot an XYZ clone” and impress people on X. But that’s not a good test.

Finding, patching, reversing, and exploiting require a cognitive skill that transcends programming languages, runtimes, frameworks… It demands true reasoning power from the model and “corner thinking”. Very, very few humans excel at this, let alone in ordinary day-to-day software writing.

Seeing Kimi K3 do so well here bodes well for open models.”

China is clearly emboldened by the success of their open models. And in case if there was any doubt at all, Chinese President Xi Jinping in his speech at World AI Conference (WAIC) made it clear that they intend to offer a counter position to closed, frontier models by American companies. Some key excerpts from Xi’s speech:

“we should adhere to the principle of openness and win-win and boost innovation-driven development. As a new engine of world economic growth and an accelerator for the shift of growth drivers, AI is moving from the digital world into the physical world. We should seize this rare, historic opportunity to encourage open source, openness, collaboration and sharing. We should facilitate technological innovation, industrial development and scenario-based application of AI. We should make coordinated advances in the transformation and upgrade of traditional industries, the cultivation and growth of emerging industries and forward-looking planning for future industries, so that all sectors and businesses can benefit from AI.

We often say in China, "A single string cannot make music, and a single tree does not make a forest." AI development should not be a solo performance by a single country, but a symphony of international cooperation.

China is ready to be more open, take more practical actions, and assume a more visionary perspective. We are ready to work with all parties to seize the opportunities of AI development and meet the challenges, and join hands to create a brighter future for humanity.”

Interestingly, even though China has built a reputation of releasing open weight models, not all Chinese companies were actually following the same approach. For example, Alibaba’s Qwen models are closed models. But following Xi’s ardent defense of the “openness” at WAIC, it appears every single Chinese company will pivot away from closed models. It is a bit amusing that Alibaba’s verified twitter page actually retweeted a post that insinuated that the strategic direction came directly from Xi. Both in the US and in China, governments are clearly becoming integral players in the AI puzzle.

China’s government is perhaps much more comfortable in calling the shots about what the strategic direction for their AI companies should be, especially in light of their own national interest. I wonder if running similar pre-training run by four different Chinese companies is of the best interest given advanced chips remains their primary constraint. If three/four Chinese labs each possess a cluster too small or fragmented to conduct the best possible training run, aggregating those chips into a nationally scheduled cluster could permit a substantially larger and more reliable training run. Centralized purchasing, networking, utilization and power allocation could also remove genuine waste. A shared open-weight base model would then turn frontier pretraining into national infrastructure.

A big assumption I am implying here is US AI companies running different, expensive training runs as pure duplication. However, OpenAI, Anthropic, Google, xAI and Meta do not merely take an identical recipe and repeat the same run. They make different bets on architecture, data curation, native multimodality, mixture-of-experts design, long-context attention, reinforcement learning, synthetic data, safety and inference-time reasoning. Nobody knows beforehand which combination will work best. Four independent training programs produce four shots at discovering a new capability or efficiency improvement. Nonetheless, it wouldn’t surprise me if China takes a more concerted approach in aggregating their limited resources in the pre-training stage. Alibaba, Tencent, Moonshot and thousands of startups could begin from the same strong checkpoint and spend their resources on continual training, reinforcement learning, inference optimization, tool use, memory, retrieval, agents and applications. The fixed cost of pretraining would be amortized across the entire Chinese economy. China could treat the base model as a subsidized public good and intentionally drive the market price of comparable intelligence toward inference cost.

This speculation, of course, rests on the assumption that advanced chips remains the long-term bottleneck for Chinese AI companies. There are indication that we may want to hold even that opinion a bit loosely. See this excerpt Bloomberg piece:

“Z.AI, the Chinese artificial intelligence company formerly known as Zhipu and focused on developing its GLM model platform, has completed construction of a massive 1-gigawatt data center powered entirely by Chinese-made chips. The facility has started partial operations and is designed to provide the computing capacity needed to develop Z.AI's most advanced GLM systems.

Investors may view the facility as a major test of whether China's domestic chip industry can support increasingly advanced AI models over the longer term. Huawei Technologies, China's leading designer of AI accelerators, is competing with Cambricon Technologies, a Chinese chip company, and Alibaba Group Holding, a major Chinese technology and cloud-computing company, as local suppliers work to narrow the performance gap with NVIDIA. The scale of Z.AI's new facility would place it among the largest data centers developed by a Chinese AI laboratory, although Alibaba and China Telecom, a major Chinese telecommunications operator, remain among the country's largest builders of computing infrastructure.”

I suspect Zhipu is not the only Chinese company to build a gigawatt scale data centers. As the training cluster for the next models become bigger and bigger, you can bet that other Chinese AI companies will also want to undertake similar projects. And if the Chinese government wants to pursue a more centralized training run in some future date, perhaps China may even go for the largest training run in the world! As you can see, the AI race is not only far from over among the companies involved, it may also be very much alive on the geopolitical front. The dominance of US AI companies may be far from certain even if they appear to be better positioned today.


Subscribers get the daily journal and five+ years of Deep Dives, i.e. full-length analyses with financial models on 65+ companies. The daily is just how I think out loud between the Deep Dives!


Current Portfolio

Please note that these are NOT my recommendation to buy/sell these securities, but just disclosure from my end so that you can assess potential biases that I may have because of my own personal portfolio holdings. Always consider my write-up my personal investing journal and never forget my objectives, risk tolerance, and constraints may have no resemblance to yours.

My current portfolio is disclosed below:

This post is for paying subscribers only

Already have an account? Sign in.

Subscribe to MBI Deep Dives

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
jamie@example.com
Subscribe