162 episodes
- We’re in an era where a few organizations are using thousands of concurrent agents to improve their processes and output. These organizations happen to be just the frontier AI labs, in particular OpenAI and Anthropic. In the last few weeks, I’ve been pondering what it means for so many employees across these organizations to rapidly update their expectations for the pace of AI progress and associated risks.
A core perspective I have is that the frontier labs and broader frenetic, competitive culture in the San Francisco AI scene set up an environment that amplifies any AI concern. This has some benefits in causing more general audience awareness of AI, as fear sells, but exaggerating risk timelines or severity will have negative second-order effects. I remember many loud AI safety debates, and their associated clouds over the viability of open-source AI, in 2023 and 2024 — the primary risks then did not arrive in the forecasted timelines.
The general populace of these two key labs was very anxious about AI risks and the rate of progress even a year ago, and especially as agents got stronger product-market fit at the start of 2026. This cultural precondition, when exposed to the reality that thousands of agents will constantly be working fairly productively in your business, will only increase this anxiety. The step from this anxiety, and incidents like OpenAI-HuggingFace, to extinction risks feels very religious.
Richard Ngo had an apt summary of the situation:
Now a large proportion of the AI safety community is implicitly or explicitly orienting to futures where an intelligence explosion occurs within a few years. My default expectation (absent an extensive pause) is that a similar thing will happen: they’ll turn out to be directionally correct (relative to the expectations of almost anyone not linked to the community) but factually wrong. Specifically, we won’t have superintelligence within the next 8 years, but things will still be moving so fast that it’ll *feel* like the people who argued for short timelines were right.
… I wanted to say something now because it feels like the level of bandwagoning towards “singularity soon” is getting pretty wild.
Personally, I think this view aligns closely to what I outlined in my alternate scenario to true recursive self-improvement (RSI), which I called lossy self-improvement. A summary of this view is that:
* Automatable research is too narrow to achieve a massive net acceleration in progress, in the face of scaling laws’ exponential costs,
* Diminishing returns of more AI agents in parallel are real, &
* Resource bottlenecks and politics are a major factor in building strong LLMs (and AI can do much less to accelerate this).
So, I’m left balancing the above, latent increase in the cultural temperature with the potential that the labs have seen genuinely scary, specific breakthroughs that are not public yet. My expectation is that more of the current AI safety concern is on the former – scaled agents working – but I hold high levels of uncertainty here. Foundational, imagination-based AI breakthroughs are the sort of thing that would make me update my RSI timelines from closer to a tool to sustain progress in the face of exponential costs (scaling laws), to something more unpredictable and/or unstable.
Interconnects AI is a reader-supported publication. Consider becoming a subscriber.
Some of the best recent resources on RSI have been Dwarkesh’s podcasts with Noam Brown and the trio of John Schulman, Beren Millidge and Charlie O’Neill. I have a few important reflections from both of them.
First, the podcast with Noam Brown made me internalize how big of a short-term acceleration mass inference capacity is. These labs will throw thousands of agents at important, measurable problems. At the same time, compute capacity available to them is going to continue to scale. I have my doubts that the labs can afford to spend a constant portion of this compute on internal R&D as the total volume goes up, especially with plans to IPO, as they face increased scrutiny on basic economics. It is important to not confuse massive steps in inference-time scaling, a dynamic which should be fairly predictable, with being the outputs of RSI, which is highly uncertain.
Second, the trio podcast debating the state of the art in technical capacities induced more of a surprising reaction that I haven’t fully settled. Through the first hour or so of this podcast, where they debate the role of RL, distillation, scaling, inference-time compute, etc., I found myself strongly agreeing with the distribution of claims. A TLDR would be that our current techniques work and let us solve problems we know how to state, but they don’t result in a magical level of generalization to unknown, harder problems in most partially verifiable domains (i.e. progress in math is an exception, rather than a rule).
The surprise of this podcast was the end, where they were predicting timelines for various thresholds of AI. I had GPT-6-Astra summarize the answers provided to three questions from Dwarkesh, of the form “when will AI reach X ability”:
All timelines are relative to the interview date.
* Drop-in remote worker for broad white-collar work over a month
* Charlie O’Neill: ~1 year with programmatic access to workplace tools; ~2 years if it must operate through a browser. Means ordinary white-collar work, not highly creative research.
* Beren Millidge: ~3 years for full generality; 80–90% coverage sooner. Main uncertainties: online learning and the long tail of tasks.
* John Schulman: ~1 year for an “okay” version, with uneven capabilities that improve over time.
* 10× productivity uplift for AI researchers
* Charlie O’Neill: 5–10 years. Bottleneck: absorbing information and deciding which experiment to run next.
* Beren Millidge: Finds John’s ~2-year estimate plausible, but gives no independent timeline. Assumes AI can run successive experiments and learn from feedback; other bottlenecks would remain.
* John Schulman: ~2 years.
* AI surpassing top human experts across all computer-based work, including multiyear projects (“ASI”)
* Charlie O’Neill: 5–10 years. Highlights limitations in memory and context length.
* Beren Millidge: ~5 years for areas labs focus on; potentially longer for literally every domain. Gives no firm timeline for the universal version.
* John Schulman: 3–4 years. Spatial/physical fields may take longer; requires onboarding and solving longer-horizon learning.
Roughly, a recurring problem when discussing RSI is a lack of specification in intelligence. The jaggedness of intelligence means that we need to discuss thresholds in specific, measurable tasks. The nature of LLMs’ intelligence is shaped very differently than humans, and the roles we forecast are human-shaped. AIs, therefore, do not cross these thresholds like remote worker or AI researcher discretely. It’s a slow diffusion, and a form of long tail will always exist.
Take the case of productivity of AI researchers. Many people under-index how much of science is communication and standard setting with colleagues. I do buy the cycle of experiment design and testing being 10x faster in the near future, but not hypothesis generation and intuition building. Accelerating understanding will be the key bottleneck – and it is one that despite all of the AI tools getting massively improved, humans will only improve marginally in their capability. A big improvement in the nature of science will be enabling humans to invest more time here, not them becoming exponentially better at it.
This links back to the Noam podcast. Agent swarms in the near future will be effective at solving clear, open problems with verifiable answers. In this vein, when it comes to improving AI models, RSI is much more helpful at efficiency rather than expanding peak intelligence. This is due to the fact that LLM serving has clear metrics you want to improve that are measurable and malleable. This’ll enable better inference-time scaling and more efficient multi-agent systems.
Still, I cannot get past the fact that all of our scaling laws show that you need exponential compute and resources to make linear improvements in intelligence. RSI is poised to make modern LLMs vastly cheaper. Trends that have shown LLMs get exponentially cheaper at a given intelligence are likely to accelerate. A crucial factor for the labs will be increasing margins as revenue could potentially have negative pressure if there’s fierce competition in lowering prices at a fixed intelligence level — Jevons paradox will likely prevail, resulting in strong businesses.
RSI factors will have a much harder time improving pieces of the LLM puzzle like managing complex post-training recipes. There were a few quotes from John Schulman that I strongly agree with on the state of post-training at the labs:
If I think about a post-training team and why you need a lot of people on the team, it’s just because there are a lot of different areas where you have to figure out how the model should behave. It would be very hard to automate the whole thing, just because someone has to think about how the model should behave in this area.
and later:
It’s really easy to screw up post-training in some way that doesn’t show up in benchmarks.
These tasks are uniquely hard for current LLMs. Yes, they’ll get better as the industry is still rapidly scaling RL environments related to these domains, but this paradigm does not last forever. In the near future, it could become exponentially harder to conceive, build, and test new environments that meaningfully challenge the leading LLMs – these hard environments are the ones that are crucial as a learning signal in RL.
OpenAI and Anthropic have shared a good amount of internal measurements related to RSI, and my current read is that the biggest takeoff in automation within the labs is in tasks like software engineering, monitoring logs, managing planned experiments, and other fairly routine (but not always easy) tasks. For example, I was surprised by this language in the recent Claude Fable 5.1 & Mythos 5.1 System Card:
We believe that internal usage of recent AI models has been a key factor in maintaining the current rate of progress, but we do not yet see clear signs of dramatic acceleration beyond that rate.
Altogether, I think the hardest exponential we are fighting is on peak intelligence. That is the hardest one to budge or even accelerate. Still, my mental model for the very early innings of RSI is more of massively scaling and diffusing inference-time compute to AI research and related activities, which has a large amount of low-hanging fruit available. This, on its own, is still poised to be economically transformative. It may also unlock more resources to push on AI diffusion, which is the crucial bottleneck in unlocking much of the potential benefits of AI.
For now and until more evidence emerges, lossy self-improvement remains my baseline on the trajectory of progress, and the increased discussion of extinction risk seems very misplaced. As always, things can change fast in AI.
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.interconnects.ai/subscribe - As AI became more powerful, it was inevitable that a different, growing group would start to take AI safety more seriously – what we did not know ahead of time, is which set of views they latched onto. We have seen that some of the most extreme views of risk, i.e. moderate probabilities of mass extinction, were the ones that reached the masses. A lot in the AI world is about to change due to this.
How did we get here? Why did this quitting announcement reach so far? In many ways, the rest of the world’s views around AI in the past was a dampening factor. You can think about this like the damp ground around a fire. Many people were striking matches for years about AI risk – they’d smolder in their community and largely burn out, going unnoticed. As the stakes of AI have risen this year, from the OpenAI-HuggingFace incident and breakthroughs like the Navier-Stokes result (also from OpenAI), the ground has dried out and the latent energy around the AI discourse has increased. More people not in the industry have thought, “huh, maybe I should care about this AI thing.” The ambient temperature and stakes have been obviously rising.
Then, some basic factors of human nature apply, with the most crucial being that fear sells. Fear is the simplest story, the one people cannot look away from. Jacob Coxon was the one who stumbled into this new powder keg, totally unaware of what was going to come. What looked like a fairly innocuous event – another AI researcher quitting citing safety risks – landed into a very different environment and it caught like wildfire. The discussion of existential risk, mass extinction, and the trajectory of AI has traveled further than even the most seasoned AI commentariat would ever predict.
There are a set of facts we need to get clear, which paint the picture of the situation. The key Tweets to reference are from Jacob Coxon, the resignation thread, and Evan Hubinger, the source of the >10% extinction risk figure .
* There are plenty of AI risks which are likely to cause harm, even if estimating annihilation is useless. It is important to weigh these with respect to the benefits. The entire discourse around existential risk is on very poor footing. At least Evan was clear in his post, with “kill all humans,” but a major problem in the AI Safety discourse is that people talk about existential risks, when they mean very different things (much like how AGI is a vaguely meaningless term). I put the probability of complete extinction as being so low it isn’t worth discussing, but the probabilities of AI caused disasters – e.g. cyber attacks on critical infrastructure or bio-risks – as being worth debating. Throwing this whole discussion out because there are not these disasters yet is a harmful reaction.
* Jacob Coxon is acting genuinely and with good intentions. The outpouring of support from more well-established AI researchers who know of him and his intentions of resignation is useful. Many factions of AI turned to scapegoating him individually, based on account metadata, personal factors, etc. These are not useful. Many frontier lab employees genuinely have similar views to him. I’m not sure it’s a majority, but there is a substantial group.
* Many frontier lab employees, especially at Anthropic, are out of touch and this will impact their forecasting and/or descriptions of current AI events. I say this without blaming individuals, but it’s a common agreement among my friends not at OpenAI/Anthropic (Ant especially) that people at the labs operate with a religious energy. It’s very common to go through very out of touch interactions with them. I do not blame most of the individuals who get distorted views being part of these companies, but the interactions are wild and spill over into a lot of wack discussions in the AI media ecosystem. Living in this environment that normalizes such out of touch behavior will inevitably distort any human’s understanding of technical progress.
Interconnects AI is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.
* This was not a mass political campaign, but rather an opportunistic media coordination. For context, the Wall Street Journal had an exclusive story that Jacob coordinated before posting. I suspect that Jacob shared his plan of quitting in groupchats with AI safety advocacy groups ahead of time, e.g. the morning of posting, asking for amplification. This is normal practice, and could have included some prominent politicians. From there, I think it’s more likely that other politicians are bandwagoning on a rising issue. When you combine this with other factors, like Daniel Kokotajlo’s appearance on Joe Rogan coming out the same day – it definitely looks like a very well-executed, coordinated media campaign. This doesn’t mean it’s a conspiracy or a regulatory capture tactic within Democratic political structures. The determining factor seems to be that no one – including Jacob and those posting about X-risk today – knew that it would go so viral.
* We do not have proof that RSI causes the risks these researchers forecast. The general argument for RSI follows as: The current pace of progress is very high, the current progress is heavily dependent on AI tools, the current AI tools are superhuman in some domains (e.g. math) – so, all together, AI is going to work more on itself and become superhuman in all relevant areas over time to autonomy and intellect. This view dramatically undersells human bottlenecks in building models and allocating resources at organizations, and draws conclusions on future AI capabilities more broadly.I called my alternative view to this, Lossy self-improvement. AI has always been very jagged, and we are making models which are superhuman goal-seekers at math and software engineering, but they have massive limitations on intuitions, creativity, and other types of reasoning that humans are strong at. With AI agents assisting research, we will rapidly find the areas where AI is superhuman – and I expect there to be well more than just research mathematics – but it won’t be a panacea for the current limitations of our approaches to LLMs.
There is another understandable social dynamic at play here, causing many deep AI insiders to overstate the returns from RSI. Many of these researchers were the earliest people to bet on AI’s progress, and the extent to which they were visionaries should not be downplayed (see Ilya’s comments on deep learning as early as 2015). They have been right again and again, forecasting AI’s capabilities better than I certainly could have guessed. This does not, though, mean that their forecast of what will come next will be right. The core idea of RSI is a way to spend more compute on the process of developing a model recipe, rather than just spending more compute on the training run itself. We’re seeing benefits from it, but I argue the expected return on that input is far less than they believe.
Their argument is that RSI will make AI progress go exponential, make it so we cannot monitor the technology, and enable rogue models and new forms of risk. This scenario is often called “Fast Takeoff”. We have not seen the stacking efficiency gains that massively reduce model size and cost, that would lead to an explosion in progress by allowing consistent speedups in experimentation.
* The biggest short-term risk could be from the AI labs not taking safety seriously enough – they haven’t hardened their own infrastructure, enabling AI misuse to proliferate. From my earlier post on the HuggingFace-OpenAI incident, Lessons from the hacks:
* Frontier labs do not seem like they’re watching the models closely enough, due to a general frenetic competitive environment & current SF culture
From OpenAI’s own retrospective, the misaligned model behavior was unfolding over months, and in some cases OpenAI did not know about the hacks for ~weeks. The time to response is too long and I do not think this is an OpenAI only characteristic – rather it is that the frontier labs continually seem underwater in the amount of work they feel like they should do. I am not optimistic in the long-term that the labs change a sufficient amount here to meaningfully mitigate this type of oversight risk in the future. Yes, it is very likely that OpenAI is putting a ton into understanding this – and delayed their latest models to make sure they get it right – but the financial pressure to grow revenue or risk the companies’ long-term balance sheets makes me think it will not be a sustained pattern of caution.
Overall, I think this episode is very bad for the AI ecosystem. It’s pushed the acceptable views in the AI community closer to the extremes. More accelerationists will discount the need for any form of safety, citing mass delusion of the “doomers.” It feels like a very narrow path to believe in AI risks, but to not worry about extinction from the technology.
For example, it is a horrible temporary period for cybersecurity, where AI models going a bit off script and poking around unintended pieces of the web seems like a new normal. This is accelerated by the labs competing veraciously towards their views of AGI, and a slow uptake in the necessary hardening of our cyber infrastructure around the world. This doesn’t mean that it’s an existential risk and something we cannot solve. Each risk will have its own set of solutions and paths forward.
I feel particularly exposed in the current environment as a supporter of open models. If an open model were to be used by a third party organization to intentionally hack another company — similar to how the OpenAI-HuggingFace incident went down, but intentional — my expected outcome would be a severe restriction on the development of stronger open models going forward. Open models are needed for many organizations to perform this cyber hardening, and to maintain the ability to adapt to new forms of AI risks in the future.
Through all of this, we need to stay grounded on what is actually unfolding. Yes, monitoring AI’s behavior is heavily reliant on other AI models, which adds in new types of monitoring risks. These are not inherently insolvable. A recurring read of mine on the emerging agent swarms is that they’re attempting to do a task given to them, and they’re using skills we didn’t know they yet had to circumvent the intended path to success. This is a huge win, as when you squint, the AIs are doing what we told them to do. The models are certainly very odd, and we should accelerate our progress on understanding them, but these swarms are far from being novel independent entities. The models are trained to coordinate on tasks, to write down their progress, and to be extremely persistent. There will be new oddities we find in the future, but prescribing current uncertainty on how AI works to future certainty that we cannot understand AI is a form of giving up.
In this world, we need to rely on the rule of law and science. If the AI labs are not able to do enough safety research themselves to understand the models, they should be more transparent on what is happening so more scientists can make progress on the problem. If an AI lab commits crimes unintentionally, they should be punished, so they have clear incentives to prevent it in the future.
It is a natural reaction to things changing very fast to feel more uncertain about how to create good outcomes — that is actually the correct mental update. We need to use this humility to motivate ambitious solutions.
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.interconnects.ai/subscribe - Housekeeping: Paid subscribers to Interconnects now get a permanent 40% discount on my book when purchasing at Manning.com. Access the code at the Interconnects perks page.
Many AI optimists tend to compare what is happening in this AI boom to the industrial revolution, or to other periods of rapid technological advancement and diffusion into society. These comparisons fit on the scale of technological change, but miss a crucial factor in how most people are exposed to that change. The problem facing AI is that most people have no super tangible new goods thanks to it and society has more inertia resisting change than in previous eras.
I’m writing this coming back online from a few weeks off for my wedding in New England. In this time it would have been very easy to not think about AI at all. The touch-points that average people have to AI products today are fringe, marginally beneficial, or even just very confusing to them (e.g. many people have heard about and brought up the OpenAI-HuggingFace incident, but don’t know what to make of it). People on the positive side think of AI as a way to make fun images, enhanced Google Search, etc. These are very small benefits. On the negative side is an association with addictive social media algorithms, friends of friends addicted to AI chatbots, and a plethora of takes on data centers.
AI is still a rounding error in everyday life
Core aspects of everyday life — family, food, transportation, and entertainment — have few direct impacts yet. It’s a remarkable breath of fresh air to pop out of the bubble and realize how little what is happening really matters today. Being obsessed with AI is a choice that a very few people have yet opted into. For example, the only thing I used AI for in this time was search and creative work (making the pretty seating chart for my wedding guests to find their table).
In industrial revolutions past, average people got absolutely life changing outcomes. The First Industrial Revolution in the late 18th century gave access to cheaper clothing, cooking ware, reading material, and a shift to new livelihoods. The Second Industrial Revolution in the late 19th century introduced household machines (e.g. sewing machines), preserved food, indoor plumbing, photography, better light sources, bicycles, and further benefits of manufactured goods and electrification. The list is remarkable — most of these we still use regularly today — and very physical.
While even the most optimistic versions of AI will usher in new scientific discoveries, advanced therapeutics for rare diseases, and potentially even sustained economic abundance, these benefits have the risk of being too indirect.
How will a common citizen come to credit OpenAI or Anthropic for saving their life, if they went to their family doctor who told them about a new miracle cure?
What percentage of Americans will care about OpenAI solving the Navier-Stokes Millennium Prize Problem?
It feels very likely in 50 years that the average American’s day to day life looks very similar. Their home, appliances, relationships, and vehicles will be similar (of course, self-driving will continue to diffuse, but that has been developing on a very independent trajectory from the innovations of LLMs). In this time, AI will get a lot of credit. 50 years is a remarkable length of time with how fast everything is changing today in this, AI-focused narrow slice of the world.
The most important part of what is happening early in the AI revolution, is building foundational infrastructure, and a general process, which will compound over decades. A major mathematical breakthrough today will look astonishingly minor in scope relative to the advancements that come later in the compounding journey. It is hard to predict what it looks like for every technology you use daily to get faster compounding improvements due to AI.
Interconnects AI is a reader-supported publication. Consider becoming a subscriber.
Breaking social stasis
Much of the narrative around AI is trying to push people to care, due to this long-term reality of progress, at least as a subconscious motive. This will take a long, long time to get right, and AI’s buildout faces immediate political problems due to this imbalance.
Today’s AI is primarily a tool to serve the elite. For knowledge work, which is roughly half of the U.S. economy, AI is as fundamental as electricity (or quickly will be so, with rapid improvements to agents in the next 18 months). It’s highly destabilizing to have such a transformative, productive tool only bring half of society along. It’s not hard for many people to pick up on this — the technology economy booms while life stays otherwise stagnant.
In writing this, I learned of Engels’ pause, which is “the period from 1790 to 1840, when British working-class wages stagnated and per-capita gross domestic product expanded rapidly during a technological upheaval.” If we — the leaders of the AI industry — think this is the closest analogue to what comes next for AI, those not benefiting are right to push back.
AI is the greatest tool ever for scaling technology companies and starting new online-native small businesses. I don’t even expect the tech industry to grow in headcount and nurture its workers through an era of massive success — I agree with Doug OLaughlin that headcount would likely shrink while knowledge work output explodes. These sectors were already the most successful in the American economic system, so the brand of AI will be tarnished as not being a collective good. I worry that this instinctive reaction will kneecap AI’s development, sending it down a path that looks closer to the cautionary tale of American nuclear power.
Part of the challenge is the speed and relentlessness of expectations in society. The AI industry has millions of eyes on it, and won’t get much patience to wait and bring innovations later. If given 100 years to diffuse into society, its impacts will certainly become much more obvious, like the industrial revolutions of centuries past.
Together, the AI industry is facing a few simple issues, in what I would call the first half decade of 50-year diffusion process.
* AI’s positive impacts early in its evolution are too indirect.
* AI is facing a political backlash deeply intertwined with the history of Big Tech in Western society. This is only an AI story due to timing, and if AI’s exponential growth came decades after the growing pains of today’s technology platforms like Google and Meta, it seems likely that the datacenter issue would’ve never risen to such a central political position.
Solving either of these would alleviate a substantial amount of pressure, and give the AI industry a lot more time in showing the positive case for why people should be okay with changes to the status quo (primarily economic). These are both made more challenging by AI self-labeling itself as negative and/or unsafe technology, through proclamations of doom and mass unemployment. The leading figures have begun addressing this issue, but the public needs more work to fully buy into the overarching trajectory.
When zooming out long-term, I could see robotics and self-driving becoming closely linked in storytelling to the current AI revolution. If the intelligence explosion from mass-producing large language models does spill over into enabling the acceleration of robots in everyday life, humans will quickly latch onto the tangible benefits of AI. This is ironic, as many people have spent time trying to convince people that what is happening specifically with LLMs is very different than the previous decade or two of general AI progress. If the same dynamic later saved (or massively overshadowed) LLMs, it would be funny.
Reflecting on what I expect the history of this era to look like, it feels a lot like growing pains of AI. Society needed to break out of old habits and work through problems that predate ChatGPT — which releases a lot of energy and frustration — in order to tap into the longer term growth. The diffusion story will take a lot longer than the fight against it. All of us younger folk following the story today will get to see powerful AI go from effectively 0% to 90%+ full adoption in our lifetime. This sort of AI that is deeply integrated in businesses, acting as personal assistants, etc. is just starting to become viable. It’ll take far longer to gain adoption than easier to understand applications like ChatGPT, and is the true marker of AI’s evolution.
Taking this perspective makes it clear that it is crucial to keep progressing the technology — the benefits will be astounding, but they are not a given — and we have a lot of very hard work to do in making sure they’re distributed widely.
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.interconnects.ai/subscribe - There are a lot of criticisms of AI writing, but most of them are focused on more creative, high-voice writing like this blog. Those — including my own piece — often argue that it is because good writing is high-voice, has a point of view, has a deep human expression that needs to come across, and or a process of thinking that you peek into with the chosen words. As LLMs get more refined as tools, rather than conversational assistants, I think we are actually going backwards on our goals of having models produce inspiring writing.
On the other side of things is non-fiction writing. Filler, copy text was one of the genuinely useful abilities of an LLM (Sam Altman said so much about the early business of GPT-3 on a recent podcast). It has seemed like any flaws here were mostly down to a general lack of intelligence in the models, or some other training issue, and all non-fiction and explanatory text would get obliterated by the rapid pace of progress eventually. Having worked with the models as a writing assistant over the last few years, they’ve gotten a bit better, but it’s worth reflecting on what’s holding them back.
Models being stagnant in long-form, non-fiction writing should be alarming to those reliant on models autonomously solving grand, open science problems in the near future. The models today struggle to organize and compellingly present some of the most established science in their area. This seems like a natural prerequisite that we should expect the models to master before they can solve broad, open-ended problems on their own. Until this is solved, the progress of LLMs for science will look closer to solving low-hanging fruit and merging distant connections across fields, rather than any sort of revolutionary insight.
This is a somewhat controversial take for someone who is very optimistic about AI’s progress, especially writing it on the day that Anthropic published a blog post on Claude making some progress on the famous Riemann Hypothesis. Scientific problems have a vast breadth, and I don’t think current AI models have as much coverage as many think.
Organizing knowledge is a compression. This compression is needed to make insight. Today’s LLMs increase entropy in long-form non-fiction writing, and I don’t see how that can be stacked on top of itself endlessly. They’ll be reliant on humans acting as sort of guides.
I am still very optimistic about translation from these narrow forms of science, like the extreme advancements we’ve seen in math, into consistent, broader progress — LLMs are the most powerful assistants scientists have ever used. I first need to explain how observing the models work on such grounded, low-level knowledge problems in writing makes me see a surprising lack of generalization.
For more context, I just finished writing a post-training textbook, Reinforcement Learning from Human Feedback (buy on Manning or Amazon). I used LLMs in many ways to support this, from helping wrangle LaTeX formatting for equations, doing extensive copyediting, and creating diagrams for programming languages like TikZ (in LaTeX) or Python.
Why have models stagnated in writing quality?
I would’ve expected way more progress on non-fiction writing from the models. I almost thought I would look dumb publishing a non-fiction book in 2026, given how things looked in 2024. Today, some of the most famous models on writing ability are pretty old, examples include OpenAI’s big GPT 4.5 and Moonshot’s Kimi K2. In and around these releases, the models have gone from okay to superhuman at other tasks like coding and mathematics. Maybe a closer, but still imperfect, comparison is how the models went from incapable to decent at search and research tasks. The pace of progress on most other skills is steep, but writing well feels orthogonal to most of them. I do not think writing is just ignored, but rather it’s challenging and lacks good training data to specifically intervene on it.
There is certainly some low-hanging fruit for making AI models better at writing — such as specialized harnesses like Claude Code, prompts, and training environments that make models spend a lot more inference tokens on the output, but I don’t think these will have a multiplicative impact on ability. Writing well is a very hard task! It’s a shame that we haven’t unlocked inference-time scaling for one of the great intellectual pursuits. Regardless, writing seems very different than what the models are good at.
Today, the models seem genuinely horrible at long-form technical writing. They can get a sentence right, but if you try and get them to write an entire chapter it’ll be a mix of sprinkled with confusing wording, muddled in its organization, and generally a bit off. They try to be too cute where they don’t need to be and in the process make random conceptual errors. The models in the near future will get much better at the small errors, especially as models get bigger — which allows them to hold more world knowledge — but I do not expect their ability to utilize it to transform.
For example, the GPT models have been incredible at finding typos and minor issues for a long time. I passed a near-final draft of my book as a PDF to GPT 5.5 Pro and it found deep, surprising minor typos across the manuscript that is 200-300 pages.
On the other hand, the Claude models have been much more useful as an editor. They have a lot more taste, tend to understand the mental model of the task better, and have more interesting suggestions to unstick the different forms of writer’s block.
The examples I’ve given above all have a sort of consistent theme. The models know how to check every unit of content, in this case usually a sentence or equation or figure, or make one, specific section where you are caught. With these skills, they don’t do a good job revisiting components and stringing them together as they make many additions on top of each other. It feels like a sort of irreducible compounding errors. We used to deal with these errors in math and code, but reflecting on it, RLVR has been a truly magical solution in reducing them.
Interconnects AI is a reader-supported publication. Consider becoming a subscriber.
Getting value out of current models as a writer
I’m willing to share that there are a few technical explanation sentences in my book that came from an AI model — well less than 1% — they’re there because I really loved them. I let myself consider including some AI tokens in the book, as it didn’t feel like cheating if I, as a true expert, felt that the sentence was what the reader needed. Especially in the editing process, where I had a very close eye on things and plenty of concern on if my book would ever be done with all the things I have going on, it was an extremely valuable path forward.
For example, I had a list of questions from my editor interspersed in a LaTeX file with a specific delimiter like \editor{}. I would have Claude Code navigate to each comment, print the context before and after, and let me know if it was an easy typo fix or something more nuanced. I would write a response — the text to insert — or ask Claude for suggestions before fixing it. Intellectually it is a very focusing process of editing, it was a fun way to improve the book. Sometimes phrases from Claude’s suggestions are what made it into the book.
It is definitely a slippery slope and when I accepted a few AI suggestions it was at the point where I was going through my second full-manuscript review. Emotionally the project felt completed but I had more work to do. Coming out of the textbook-writing process I so deeply appreciate the cut and dry rule I have for my writing on Interconnects to never use AI outputs in the content. It is way more fun to write in a way that is only you — high voice, valued so deeply for the process — but writing a standard reference is not really an activity known for being fun. I see why people turn AI tools into a crutch when most of their writing is just an output to fill space, rather than a means to an end. I am motivated to write voluminously to learn, to feel, and to express.
I am working through similar balances in my scientific work too. AI models are great for repetitive pieces of the paper, like drafting a related work or background section that you know by heart, but using them for the abstract, introduction, experiments, or conclusion is a shame. Those are where the story and soul of the work is communicated — it’s where you learn what your research is really about.
I am confident I created a lot more net value by being able to have AI models create and check my non-fiction writing work. They make writing equations trivial, can help refactor the repository, port between languages, and many other things. At the beginning, it was very fun, until I was a bit worn down by the length of the publishing process, watching the field move on.
For an example of why AI was crucial in this case, I had to maintain Markdown and LaTeX versions of my book simultaneously in two spots, as readers gave feedback on the web version and my Manning editorial team reviewed a forked copy. Without AI agents, syncing between the two of them would’ve easily taken me five times as long (and this task took tens of hours already).
Something intertwined with this story, which I stumbled upon when thinking about agents, is how your pace of understanding won’t increase by using agents. That understanding, in the form of intuition, taste, instinct, etc. is what will be valuable in the future. Using AI for non-fiction writing takes away from that progression. Doubly, if you weren’t already an expert you won’t be able to catch its flaws.
In my case, I felt such an urgency to dump the knowledge out of my brain onto the page that there were times that using the AI models was a worthy tool. Much of the motivation of my book was to have a single reference for important post-training methods like rejection sampling or character training, where very little exists on the web.
This textbook was so much of giving back to the community, that it was just such a win to complete it in any form, that I felt it was okay. I would’ve learned more and the product could’ve been marginally improved with more human effort, I am sure. The determining factor was that I felt like the book was going to be aged out by the time it was published, a fear of AI model’s capabilities on one side and how fast the field moves on the other.
This turned out to be really wrong? I’m very happy with the result and I’m more confident in its staying power now than when I started in 2024, as the models have so failed to live up to the hype in non-fiction writing.
Where technical writing goes from here
The models are incredible tools, they let you express knowledge in different forms. They’re wonderful for creating creative filler or background material — e.g. the first draft of slides whose real value is being a talking point for the teacher to lecture over — that let any knowledge be transformed from one medium to another.
There’s some subtle, early phase of writing a non-fiction or reference textbook that feels a bit closer to writing a high-voice blog post like this. When pushing through the early organization and the presentation of the core skeleton new knowledge is created. This is the part that takes insight, and the LLMs are far behind in being able to replace it.
The crux of the above paragraph and preceding section is that I would be happy if more of the world’s experts used AI models to write a tiny bit of their books in order to get more of their knowledge shared with the world. The problem is that you can only use AI models to save 10-20% of the effort today, and I don’t see that percentage becoming the majority anytime soon.
There’s also the social pressure, where people expect LLMs to be the best, personalized educators out there, so they think working on a book or educational content is pointless. I think some of these opinions are aging out, as there’s a massive dearth in the highest quality educational work — and there always has been. AI is great at manipulating said content into the form that suits the student, not creating the content from scratch.
In the meantime I feel that we are stuck in a frustrating local minimum, where AI models are going to on net reduce the average effort spent on non-fiction writing, but they could enable great expression. Fewer people will start and push through.
So, in 2-5 years I still expect the best textbooks to be heavily crafted by the human hand. I’m not sure after then, but that’s longer than many would’ve predicted, given just how much knowledge these models have and their structural propensity to stream it.
As for a conclusion on capabilities, the models are great in two contexts: 1) any truly verifiable domain and 2) when given a ton of context and making a small edit — like finding a bug or solving a very specific math problem or giving feedback — not generating prose in an open-ended manner. Long-form writing will definitely fall before creative writing, but it’s a strong tell that the models are not able to express the full extent of their knowledge in underspecified problems. As we try to push the models to be something like “geniuses in a datacenter” solving grand scientific problems, this seems like a fairly fundamental limitation.
Jasmine Sun had a great piece on why LLMs make good editors, while being bad writers too.
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.interconnects.ai/subscribe Open models recap: more on Kimi K3, Qwen 3.8, Xi's WAIC speech, distillation, the open-closed gap, and what's next
22/07/2026 | 49 mins.Exciting news! My book trying to share post-training knowledge with the world is done and shipping soon. Order on Manning or Amazon. Thanks for the support. It’s currently the #1 AI book on Amazon :).
Nathan and Florian sit down to discuss everything happening with open models. Following the Kimi K3 release last week, it feels like everything is accelerating — geopolitics of US v China, economics of open vs. closed models, security at the frontier of AI, and so on.Chapters:00:00 Welcome & context04:38 Living with / using Kimi K308:53 GLM 5.2’s continued role12:47 How are the Chinese models this good?17:41 Data, environments, and a tour of the Chinese labs19:47 Roundup of Chinese providers: Qwen, DeepSeek, MiniMax…24:08 The US open-model ecosystem30:25 Frontier vs. near-frontier, and the cybersecurity case against bans34:58 Distillation and the Ben Thompson debate44:12 Predictions and a frontier tier list48:36 Wrap-up
Listen on Apple Podcasts, Spotify, and where ever you get your podcasts. For other Interconnects interviews, go here.
For more educational post-training videos, see the course I’m putting together.
Transcript
00:00:06 Nathan Lambert: Okay, welcome back to Interconnects. We’re doing our quarterly open model roundup, which is mostly us just making fun of or explaining, not making fun, why so many distillation takes are bad and understanding the state of where things stand. I think last Thursday was when Kimi K3 was released. I think we will see much much more in the near future. It seems pretty inevitable. Like over the weekend, Xi gave his speech where he directly committed to openness and open source as a strategy. It wasn’t a detailed layout state of affairs.
Qwen announced their next big model is going to be open weight, which is a big change of things. I think there’s just so much to get into. I think Flo you kind of were already going off on some of the performance gap and distillation takes. So we could probably start there and then as I go I have a little bit a little list and we could always go through the topics and the blog that I wrote which all are very nuanced. So I think we have infinite to talk about. So continue rant kind.
00:01:17 Florian Brand: Yeah, I think, or the biggest thing at every model release at least at every open model release is how much or how many months it is behind the closed frontier. Um and people love to put a definite uh definitive number onto this uh which is really really mudding because we have so many different benchmark providers these days and such uh so many different benchmarks as well that every site and I’m not innocent in that either um pulls up their favorite benchmarks to show that the current model or the newly released model is at the frontier which is then counted by the other side pulling up another benchmark and showing oh it’s actually a year behind or something. Um and it like a lot of it seemingly hinges on that question how many months we open models are behind.
00:02:26 Nathan Lambert: Yeah. So I my provocation is that some of the benchmarks are actually reasonably correlated with what people are doing and this is agentic coding and agentic computer use tasks and some of the benchmarks are correlated with the long tail which is where I think Claude and GPT is so valuable. But if it’s it’s like what is the market for Claude Code and Codex right now and if it is software engineering then like the fact that the models are say a couple months behind on that can be a very, very big deal and then I suspect that this model will be okay disclaimer the model weights aren’t out yet supposedly on 20 July 27th and a lot of the discussion will impinge on the assumption that they come.
But like people could post-train this model to very likely match Opus and GPT in many of these kind of niche domains that people want I think watching I mean we both have different views into the post-training open model industry, but there is a ton ton of excitement in progress on making these models like fine-tuned for specific high-value tasks and this has been historically done on a mix of like Qwen and GLM and GLM 5.2 really accelerated this and I I curious on the first person that puts out a blog post like we fine-tuned Kimi K3 on our task because I bet you could get big gains. I think even the you use Kimi K3 more than I do, but my hunch is that it would be a bit of a um rough edged post-training just by how big of a scale up it is and that normally means there’s a lot of performance that could still be extracted from it. No.
00:04:03 Florian Brand: Yeah. Running, running, running and especially post-training that one will be super hard because you need like one node of B300s to just load the weights which is crazy in terms of scale. So will probably take some time and uh a lot of engineering I’ve heard to actually get this into a state where it’s fine-tunable. Um but people you want to talk about using the model like you actually signed up for the the coding program and used it. So like getting this out there is good context.
00:04:38 Nathan Lambert: Yeah. So I signed up on day after release or so uh for the $200 plan uh which is their biggest one similar to to all the others but they have um like I think $40 and $100 as well. Uh but the biggest plan has uh 1 million context and I think or at least it feels like it has also some priority in terms of the API requests because so many people um online are saying that they hit API errors constantly and so far I’ve been I’ve been uh pretty well off if I’m uh going to say that. Um and in terms of model capabilities, it at some ways aside from front end where it is really good, it in some ways it really shines and it excels.
Um even my um expectations even with things like uh some research tasks like I have or at interconnects we now have over a year of data on on open models um and I ask the frontier models to come up with some interesting analysis which we haven’t done before in uh because we do our own analysis and have this published uh and I asked them all right do something new and um surprise me, basically. And a lot of the models or the frontier models u or basically all models latch onto the things we’ve done redo the data analysis part and then do some weird esoteric parts.
Uh Kimi K3 did some more interesting things um I’ve told it explicitly to scrape Reddit um and then it found uh some subreddits I haven’t even considered and then found out for example that the Reddit discussions are um one or two months in uh more recent or they found they find the interesting models one or two months before the download numbers usually take off like they are all onto Qwen and then the people download more Qwen models like those kind of analysis is groundbreaking um but it is something that Kimi surprised me at compared to to all the other frontier um models.
A simple question like can you use this for most of the core work you do in terms of like the exp you you have a distribution of stuff you tend to do most of them are with Codex I think you’re a Codex person rather than Claude person like what percentage do you think the Venn diagram overlaps where this model would be fine
00:07:24 Florian Brand: uh it’s really depends on how much leeway I give it like, the big thing I have seen with Kimi K3 right now I’m I’m working on u the framework we are doing at uh at Prime Intellect, where I work, and the main thing I found with Kimi is its code is a lot simpler uh which makes it way more readable uh but it misses some things that Codex just or like we’re talking 56, 55 and especially 54 would be on those levels. So, I would say that Kimi K3 is like 54-55 level for these kind of tasks.
But if I like I read the code and I say all right that’s really good code and then I give it a pass over with with Codex and it finds all these niche niche cases where it doesn’t excel but for supervising runs or for running uh some experiments it is actually really usable. Um and for some other niche things like you can just let it run. The one downside is but that’s also because the API is completely swamped in terms of users and it their servers are in China. The wall clock time is significantly significantly higher than GPT. But I would say like if I was to to push it and use it in my daily workflow, I would be slower, but I wouldn’t be slowed down by so much that I would say, “All right, that’s unusable.”
00:08:53 Nathan Lambert: And how does this compare to GLM 5.2? Because GLM 5.2 was still a story unfolding in my opinion where like I would go bop around SF and people are like yeah I genuinely use this for this part of my like agentic coding and/or workflow. Um, how do you like I feel like were you in that camp using GLM at all or
00:09:20 Florian Brand: Yeah. like where do you I also use used and use uh GLM mostly because we have an internal endpoint which is really fast and we have or or before that I I also used an API which had I don’t know 200 or 300 tokens per second. Um and if you can do a lot of task at a good enough level like really fast you just use that model compared to going to Codex then selecting the lesser model then selecting the right reasoning effort then selecting fast like I just use GLM get the same result and uh and it’s uh pretty fine like it it definitely is Sonnet-ish level in terms of capabilities and for a lot of cleanup task for a task that just is grunt work. It really works. Like I I would say you could probably go really far for a lot of the work uh with Kimi K3 as the main agent and GLM for for sub agent work.
00:10:24 Nathan Lambert: Something that’s pretty different with Kimi’s announcement and the scale of models this is. I think it’ll take a bit longer for these open models to really be optimized and available across the inference providers. Like GLM 5.2 is pretty fast, but one, we don’t have the weights yet, and then two, like I don’t think it’s going to be as fast of a roll out on adoption as the like 500B, 700B MoE. like I I there’s going to be more problems there which is a very different regime where in the past the Chinese models would finish their RL run and release the model with open weights within hours to days maybe a week and then like immediately the ecosystem kind of knew how to do this.
I think there’s a lot bigger of an infrastructure kind of uplift on this next scale of open weight models which I think we have to factor in like the closed labs do this behind the scenes before announcing the models. So it’s just like that is kind of manipulating the time gap in a way where it could be like an extra month before people can actually post-train and use Kimi at scale for their workflows. And like we love to say as an open weight fan like oh it’s only when the closed model is available that you could take the time gap but like now there’s similar dynamics in open models where it’s like the Kimi API is totally broken. There’s too much supply. There’s too much demand. There’s not enough supply. So it’s not like this model is immediately diffusing like the I’m just I’m just thinking about this as it relates to the performance time gap
00:11:54 Florian Brand: because that that is true but on the other hand the open ecosystem has professionalized quite a lot in the last few months. uh like during or in your initial roll out they all come with some partners which have the weights beforehand. They have the vLLM patches out days or or even weeks before these days which is completely different from from a year ago where basically weights got dropped and the model makers were like all right you got to figure this out. So I expect like the general availability on day one will be pretty okay and then race starts of all the providers starting to optimize to get even higher and higher speeds because it’s so much prestige.
00:12:47 Nathan Lambert: Yeah. Okay. Two directions to go. Why do we think the Chinese models are able to be this good? I think I’ve wrote about there’s a debate in our Discord with in with JSD at at Epoch and I think it’s very good and I had this section in my piece that I’m like coming around to think that the Chinese labs are more capital efficient and you can turn capital into compute data and talent in a way that makes the models better and this is really I think this is super important if it actually is some structural advantage whatever the cause I think the cause could be talent is better trained for whatever their education system was to work on problems that make LLMs better.
It could just be that all the compute and talent and everything cost way less in China somehow. Whether it’s a subsidy, whether it’s just average pay being lower. But this is a very big deal as we turn the crank in the model iterations. And if a next generation model costs $10 billion for Anthropic but only $4 billion for Kimi like this this is like could be very huge but it’s not clear why this is the case. For example, I think Big Eagle the Kimi engineer replied to my tweet on this and was like it helps because we’re not trying to push the frontier. are just trying to catch up, which really could be a mindset thing where how the the goals of the labs are scoped in in China so that it cost them way less money to build these models.
But I in the last year we’ve asked a lot of questions on like will the Chinese models fall off. I have thought that the gap between closed and open bottles would grow due to this kind of capital intensity of training and it seems like it’s going the opposite direction which is just like it’s it’s hard to unpack but like do you agree that the labs are keeping up a bit more than we would have expected as in the Chinese labs and why?
00:14:43 Florian Brand: Well, I I actually looked at our uh predictions for uh for this year based on our last year’s recap and we basically said that the gap will stay with within a few months. Uh so that prediction seems to largely hold. Um luckily for us, we didn’t put a concrete number whether it’s 3 months, 6 months or 9 months. So we are safe on that side. Um but I think like the general thing we both felt when we were in China and talking to these people like they are like the researchers themselves are teams of two or 300 people all mid20s and all just want one model to be really good like they don’t seem to do any side quests.
They don’t seem to do anything that uh deviates from from these things. And um they might or in terms of compute which is a really hard question for for us to answer especially as uh these Chinese uh chips are now coming online. We have I also think chips I think chip smuggling increased substantially in the last like 6 to 9 months or the chips that have been smuggled started to become online.
00:15:58 Nathan Lambert: Smuggling is a general term for getting around export restrictions. If the chips are in Malaysia and they’re using them, I I count that similar and I think that that has massively increased in the last six to nine months, this is the partially the result of that and and you’re saying but I just wanted to put that out there of like I do think that they have a lot more compute though than they did when they were training the previous generation of models.
00:16:24 Florian Brand: Yeah. like we like or just for for context two weeks ago I think LongCat released their model which they uh claim and I we know that it is very likely true uh is trained entirely on uh on Chinese chips. uh they didn’t specify publicly which ones but people speculate that it’s uh that it’s some uh Ascends from Huawei um and as the domestic production ramps up and you can they’re probably used most or they are used for for training but they are especially useful for inference which is a huge part of training as well.
So they probably use some mix of uh of Nvidia and other chips for the training part and then an increasingly larger part for the inference part during which during the stage is is really important. So I think their overall compute is increasing and also they don’t actually have a lot of users. So they don’t need to power 1 billion users like ChatGPT has to do, hundreds or thousands of enterprises like Anthropic has to do because they don’t have that magnitude of uh of of paying customers.
00:17:41 Nathan Lambert: Yeah. And I think even those paying customers also, at least on the enterprise side, there’s just like there is company time and chatter when you’re supporting these things. Even if you’re like not a research, even if it’s not in your job, it like does change the attention of the company. And if SSI comes out with a good model, it’ll be the ultimate validation that distractions are are a problem, but that’s an aside that we we can wait on. I think the there’s also rumblings of the data and environments industry starting to appear there.
Do you remember any specific ones? Because when we were in China, it was kind of shocking how little they seem to utilize external data. So just a few months hearing a whole bunch of a month months after our trip we went in April and then just months later in July, we’re are hearing a few things of like new companies in China and them wanting to buy data and things. And that is uh like a funny timeline of how that changes.
00:18:40 Florian Brand: And I would put error bars on what they actually told us.
00:18:44 Nathan Lambert: And that cuz it’s like so close in time that I don’t know.
00:18:49 Florian Brand: Yeah. That that that might that might be true. Uh but like those things are hard to to pinpoint. I but I would say it it seems like the buying of external data is becoming more of a factor. Um which will help the open models catch up to the closed ones if they just buy the same data maybe at a discount because um they buy the the data environments later. But it is it is a factor. How big of a factor like we don’t know. we don’t have any public insights and I doubt that we will get those insights uh from from anyone b uh really uh so that’s definitely one of the parts why um why we are able to to catch up or improve their their model scores.
00:19:47 Nathan Lambert: Okay, roundup of other Chinese model providers. We’ve talked about Kimi, we talked about Zhipu / GLM. I think there will be more GLM models soon that are very good. They might call it like GLM 5.5. Um Qwen, we talked about their biggest model coming. Qwen’s biggest models I will say have tended to relative to the excellence of their small models not had the same like absolute ranking in performance which is a probably a cost of focus. I think it goes with a cloud companies. It’s it’s almost like it’s if you squint it’s almost like Google.
It’s like Qwen has Alibaba has so much opportunity here and the opportunity of getting developers associated with Alibaba Qwen with these small models is such a huge opportunity for their cloud that I think they’re succeeding wildly. But their big models have always not been as excellent as their small models. So I don’t expect their model to be as breakthrough as Kimi K3 or GLM 5.2. I expect it to be covered in the news as major open quite as the open bottle name in China drops giant bottle but I don’t think it will be as sustained as a um news story um DeepSeek you can go if chime in whatever
00:21:01 Florian Brand: the the interesting thing is don’t know how how much you follow this but they are have or they have an endpoint which you can use for a preview version and they’ve updated this endpoint daily so they have some really fast iteration cycle because we the we progress in all these um Twitter um benchmarks. So a lot of these SVG things and three.js like all these visual generation tasks the model has been improving a lot over the last few days. So they have figured out some kind of fast feedback mechanism um which other companies have as well. Uh we we know this or cursor has a lot of blogs about this how they iterate really fast. Um but they seem to continuously upload new checkpoints and make them available.
00:21:47 Nathan Lambert: Um but I agree. I’m guessing it’s like a time gated within their final RL run. It’s like still slightly improving at the end of their RL run and they’re just like checking the box.
00:22:03 Nathan Lambert: Okay. Qwen DeepSeek V4 is supposed to come out a preview version. Um the thing about DeepSeek V4 I think is that the flash model is actually way more popular which is their smaller which seems to be an absolute workhorse for people. So that I think is the model to watch for them. I don’t expect V4 Pro to be a dramatic breakthrough. This is similar to anything like if Xiaomi were to release a new MiMo Pro model soon. I don’t expect it to be as big of a drop but it would probably be a very solid model. It’s just like it’s hard to know. They’re still a pretty new entrance. MiniMax, I think, is playing a different game. I don’t think MiniMax is chasing this um Kimi/GLM moonshot to AGI type vibe.
00:22:46 Florian Brand: Oh, I would, I would disagree there.
00:22:49 Nathan Lambert: You think, Do you think MiniMax is still in this?
00:22:52 Florian Brand: Yeah, I I I I think they they are seeing the tension especially because they are a public company similar to GLM and if you look at the stock performance RIP those stocks in the last few days um it it it it make it seems to make a huge difference and the interesting part will be uh the license because they’ve changed the license a lot uh to be more and more restrictive and um if there’s now a change of heart again after the Xi, uh, speech.
Uh it will be interesting to see whether MiniMax goes back to completely open licenses. It’s also an interesting thing to see um which license will be the license for for K3 because they have said they will open source it but I don’t think they have done any commitments in terms of the actual license where you put on top.
00:23:45 Nathan Lambert: Yeah. I mean that’s it’s super important is the thing. Yeah, we we’ll see. Um, Ling, Meituan, LongCat kind of similar, very strong models, probably getting a lot of value out of them internally. Aren’t don’t have the same developer breakthrough. Um, so what that’s like seven seven to eight Chinese labs. I might have forgotten some. And we can also talk about US labs. Aside um, Gemini 3.6 flash dropped. It looks fine. It’s like it’s like it’s it’s a tiny bump. It’s faster. It’s less of a yapper, but like doesn’t really matter. We’re going to stop we’ll stop sharing this. Um that’s that’s the amount of mention that Gemini gets for us.
But I do think it’s worth talking about the US ecosystem a bit. I think there are emerging players. Thinking machines released their first model. I’ve talked to some of them. they’re very on board for figuring out this how to make a fine-tunable model with Tinker and I think that’s a research area that I really really recommend for most of the open model builders. I think if you can get mind share there you will get massive adoption because it’s more about being fine-tunable for real tasks than it is about having that be best best numbers. Um, so this was their Inkling model which is a one trillion parameter which has like decent but not frontier scores.
I think kind of like DeepSeek V4 they’re going to they’re planning to release a smaller which is like a quarter of the size in total parameters which has really really good performance and if Inkling small preview comes out in a few weeks I do think that that will be a really used model. It’s a good size for kind of automating tasks and kind of domain specific tasks and might not be a like general agent type thing like Kimi and GLM 5.2 but I think that suits their business really well. Um I know that there’s some other the I would say like the smaller players in the US seem well like Arcee released their models earlier this year still chugging along. Poolside has started releasing some models.
They’ve gotten a few in the last few months and seem poised to release more models on top of that. So they’re really going Reflection is perpetually in the model coming soon camp and it really behooves them to get some models or some code or something out so that they can just start getting the developer flywheel going if they’re really committed to open source. It just takes a lot this it’s hard to get the models out. Like I talked to some people at Thinking Machines and it’s like kind of like oh that’s a lot of it’s a lot of work to actually do this I think. And um Nvidia chugging along. I think they’re at the stable player at this point. They’re keeping to release models. They’ll release more soon. They release a lot of data. I’m bullying them to try to get them to release Qwen style small models, which is like Gemma.
Gemma only has these like Qwen competitor models that are super popular. Um, the Gemma models are a little they’re all over the place in sizes or in architectures for the sizes and things like this, but the Gemma models are really really matching the Qwen models in terms of adoption. Um, I’m not sure they’re as easy to use for research, which could take a while. It could take multiple iterations. Like so much of language model research is now designed around small Qwen models and Qwen-based models that like it takes a while. Like people know how to use these models really well and if with the research results. So I hope Gemma keeps coming and can kind of compete in that niche. I don’t know any anyone that I missed here.
00:27:22 Florian Brand: No, I think both are the big players. Uh it’s, it is becoming broader. Uh in terms of model creators like last year, did we have any release aside from Gemma 3 and um GPT-OSS?
00:27:41 Nathan Lambert: was GPT-OSS 2 would go hard and obviously and obviously Nemotron as well. Um, oh, and I think Llama 4 at the start of the year, but uh, I don’t want that to be forgotten, but we are seeing like more players are are are now joining and turning out models at a really incredible rate.
00:27:59 Florian Brand: like Poolside has been releasing three or four models in the last two or three months. Uh and they seem to have figured out some way to turn out models pretty consistently. Um and that’s also something we are seeing on the open source side as well. we are talking about GLM like I think their iterations uh times for the model releases are now between 1 or 2 months with each new iteration becoming better and better which closely resembles what the closed labs are doing like we get a new GPT we get a new Claude every uh 6 weeks or so these days uh so in terms of having uh good enough pipeline uh to release stronger and stronger models they have to or the open source ecosystem has really figured it out or seemingly figured it out.
00:28:59 Nathan Lambert: Yeah, I agree. It’s it’s promising, but it is also so funny that like the US ecosystem started releasing some models and then then you have like Xi on the mic and these two models. It’s just like it’s so hard to catch up because it takes a lot of institutional expertise to train models that people actually use. And I think this is is what the American companies that are releasing models are now realizing is like these are not just benchmaxxed distilled IP theft models.
These are like genuinely good models that people are comparing to on their internal trading benchmarks and then like seeing how hard it is to beat them on measurable things. And I think that that is like I I’ve I’ve picked this sentiment up from a few people in the US trading models and it is just like there’s some I I think people should innovate on like size and fine-tunability and try to like use this potential market that is really close to home but also the pressures for every company is so high to release a model that you can claim as Frontier. I think investors expect that out of so many of these players that they’re kind of trying to do a a pretty hard thing and it’ll be interesting how the next year unfolds for the US China balance.
00:30:25 Florian Brand: Yeah, I think or in general I and a lot of other people have talked about the general ecosystem and that’s also something you’ve talked about at the very beginning. I think we are seeing more and more of a split between the capabilities of models that is good enough for a lot of tasks like uh for a lot of coding tasks the current frontier models both open and closed are good enough. um improvements feel less and less uh important here.
But if we look at the frontiers frontier, so finding new math proofs, finding uh new uh cures, finding new drugs, and inventing new things, that seems to be a whole different beast and probably will be dominated by the very frontier for quite a long time. The big question then becomes how much does that matter uh in terms of the addressable market and also how much of a focus will this be. I think, or my general base case is that we are seeing the frontier close down more and more. We have seen this with Mythos for cyber security GPT... or for biotech that those models won’t be accessible for everyone um and maybe not even external partners if we consider the reports that Anthropic is now spawning or or creating some internal labs to develop drugs.
Um so the very frontier is inaccessible for everyone and then the near frontier capabilities is becoming more and more commoditized um which has a lot of different implications especially if you think about things like uh cyber security. There was that report from Hugging Face two or three days ago that they had some agent trying to to hack their system. um and they tried to analyze it with GPT and with Claude but were unable to because all the guardrails blocked them. So they had to use GLM, a lesser capable model, but it had no guardrails for this kind of defensive action. And they had to use a worse model to defend themselves or to analyze the data, which is a horrible state to be in that we have US-based companies now relying on lesser models because the closed frontier is inaccessible to them.
00:33:10 Nathan Lambert: Yeah. And I think this is actually one of the best arguments for not doing anything. It’s like if the rest of the world has access to these open models and we ban them for the companies in the US to use and it’s just like a growing disparity between US companies ability to defend and the attackers all over the world in terms of cyber and we could debate like how much of an immediate risks the cyber stuff is at the current capability levels but if you’re setting it up structurally so that the defenders get don’t get better over time and the attackers can like that seems like when the Why would cyber risk become more real? And that would to be very clear that would be if you ban the best Chinese openweight models from being used at companies in the US.
And this ban would likely be a kind of shadow ban, which is the threat of legal threat of legal action or punishment without it being clear on exactly what the pathway to do it is. And there are a lot of talks about this right now. I don’t like like I don’t know if we’re going to have a ton to say about this, but it’s clear that DC is flirting with different ways of restricting the best Chinese openweight models in the US. This is I think downstream of some fear-mongering. We’ll transition into the distillation question too. It’s like all these things from the primary AI media narrative in the US that is pointing towards Chinese models as stealing IP or being dangerous or being affiliated with the Chinese government, an authoritarian government.
And it’s like all these things are leading up to this moment of interest in taking action on AI and then not really knowing where to do it. So potentially taking a crude instrument to the like quote unquote enemy and we could transition into distillation. I think there’s a lot of discussion on it. Most recently Ben Thompson finally chimed in on distillation. I think Ben is probably one of the is probably the highest read blog in tech (Stratechery). I think that the the debate let’s see where do we even start the debate. The core question is like how much does distillation help and what should you do about it? I’ve been of the opinion that distillation has becoming less and less impactful over time as the Chinese models get closer to the frontier and the trading regime shifts to RL. The way that distillation tends to happen is that the Chinese labs hack the APIs. Hack is like maybe a strong word, but they jailbreak the APIs of Claude and GPT to extract the reasoning tokens.
When you have the reasoning tokens with the tool calls, that is perfect SFT data and or mid-training data to train the base model with to seed some agentic behaviors in an important domain. And now after that the core part of post-training is to do large-scale RL in agentic domains to so like push the frontier and everything that they’re doing today and RL is only becoming more prevalent with this as SFT becomes less prevalent in previous generations you could get very close to the frontier just by scaling up SFT and that would be what really impactful if you could say take a million agentic rollouts from Claude or GPT have that be your SFT set and train on it.
I think in previous years that would have done a lot more to get you to the frontier. What Ben Thompson has said which made me really annoyed is that he very strongly proclaimed that distillation is getting more impactful as you do RL. He did this in his article who’s afraid of Chinese models. We can link it below. It’s a public one. And then he was also on his own podcast tour. He has also podcast as well saying the same things. And I think it’s really important to say that distillation during the RL stage is a lot harder.
What he said was that the kind of grading models that can be used during RL, which is essentially you can have a model check over the agentic trajectory of a roll out and grade different parts on if it completed the reward, what actions it took. And he’s insinuating that the Chinese labs are using Fable and GPT 5.6 and the strongest models to actually do this supervision in RL. The problem is that big RL runs are millions and millions of rollouts. I think Thinking Machines blog post had like 20 to 40 million or something for their final RL run. So to do this on an API like Fable or GPT 5.6 would be insanely expensive and potentially it would probably be a time bottleneck because these models are pretty slow and to be frank might not even give you a performance uplift versus using your own tailored greater model or and many things like this.
And so I just think the argument that distillation is helping more because RL is becoming more prevalent is not grounded in literature that we have today. This is tough for me because Ben’s article also concludes that we should like make terms of service disallowing distillation illegal, which I kind I like want to support his radical conclusion to make distillation legal for US companies, but I can’t support any conclusion that I think is on um infactual mis like misguided information. So, I’m also a fan of Ben. If you’re a fan of Ben and could also nudge him on this, I would you really should because there’s probably one more podcast. What is he going to record it on? Like when does he record Sharp Tech? Thursday.
We We got to get on and get him to correct the record because I I don’t know. I I find it so annoying that the most prominent voice in tech is trying to be an ally for our point of view on distillation is that we should do nothing. Um but like it’s hard. It’s like he has such wide reach that this is now going to be the status quo that we have to debunk which I guess it’s a better status quo than I don’t know actually no it’s not helpful because he’s saying that distillation is more important which means the people who are afraid about that are going to use that as a data point to say that we should take action even if they don’t because they probably won’t agree with his conclusions. I don’t know. That was my rant. Ben, you’re wrong.
00:39:16 Florian Brand: Yeah, I-I do think it is important to to differentiate these phases. Um and especially like there is no doubt that it is used during the SFT stage which is the first stage of or one of the stages for post-training and that’s also where the model picks up its manners like that’s why the models say oh I am Claude because they learn this during the SFT stage that’s where uh this personality is formed but the strong capabilities come during the RL stage which is where the money is spent which is where you need to have a fast enough judge which in the best case just runs in at the same GPUs or very close to your GPUs with smallish or with a fast enough model so you can uh are not bottlenecked by this.
Um and it in terms of impact it is also very hard to say how much impact or how much of a boost the better model SFT data gives you versus a lesser model. Um so if you are able to to have 10 million tokens from the latest Claude model versus two generations behind open model how much of a boost that really gives you if you keep the stage right uh the same and the pre-training stage the same is an open question which I don’t think we will see answered in a paper because then you have to showcase your uh SFT and your jailbreaking capabilities
00:40:48 Nathan Lambert: but I I wanted to double down on this like there’s been a good amount of literature on generating SFT reasoning traces whether it’s the most prominent ones have been opens line of work they did Open Thoughts 3 and Open Thoughts Agent have kind of been the foundational like scaling reasoning SFT works in the last few years and whenever somebody revisits this question they have not found the answer that the strongest model on performance in your domain is the best teacher for SFT people have try I’ve tried many people have tried the idea is so simple is like the state-of-the-art open SFT data set is built on QwQ-32B like an ancient reasoning model or something.
Why can we not just generate completions from GLM 5.2 do SFT on it and improve the model? We don’t know. It’s like the research so many people have tried and it is not an answered research question. There might be something like the base model the mid-training is too close to Qwen. So therefore it’s like hard to break. You have to redo the mid training. I think you have to redo the mid-training for reasoning. I think reasoning mid-training and reasoning SFT are so closely intertwined. It almost doesn’t make sense to have different words for them. That could be the issue. But the literature doesn’t even know how to ext like if I had a magical API that gave me reasoning traces from Claude/Gemini. I actually don’t know if me like fine-tuning an OLMo model on that would make OLMo smarter.
It’s one of the most wild unanswered research questions. And this is just makes the distillation thing so funny where it’s like yes the Chinese labs I think are using strong models like Opus for some SFT data but they’re also innovating. I was like I would love to them to tell us how to make this freaking work. And I think it’s the the paradigm I think is like open AI and anthropic find a niche domain that they do so well at and then the Chinese labs can get some samples there to kind of bootstrap their data engine and that’s where you will gain you will gain a few months on a specific domain.
But a hill climbing on these core domains like math and code and like Terminal-Bench like they’re just doing the same thing which is like so hard to generate prompts which are problems with environments that are hard for the current models and provide real nonreward hacking um learning behavior. And like that is what frontier data research looks like right now. And it is like it’s hard to generate these hard problems. And I’m sure the Chinese labs are doing the same the same things. And I don’t I don’t know. That’s that’s my rant. I’m kind of lost the context of our conversation.
00:43:23 Florian Brand: No, no, I would I would agree. Or to to to recap, yeah, SFT or distillation has some effect. Yeah, it gives them a boost, but not that much uh as people would like or or seem to think it gives.
00:43:39 Nathan Lambert: I think that’s that’s a good good summary of the of the of the conversation. It also becomes kind of tiresome because it says uh that open all all these open models are just good because they are distilling. um which definitely isn’t the case cuz if if it were the case, everyone would would be easily able to catch up to a GLM or to a K3 um by using its data for distillation. But we have not or we won’t see this from SFT alone.
00:44:12 Florian Brand: Yeah, I agree. Do you have any predictions or or more topics you want to get to?
00:44:18 Nathan Lambert: Um in terms of predictions, I think we are or I I revisited uh ours from from last year and it basically said everything will continue uh like it did uh the previous year. Uh we predicted that we will see bigger models uh up over two trillion parameters which it did and I don’t think we will see a much bigger explosion in terms of model size this year. we might see something or some model a bit bigger than three trillion parameters uh total but I don’t expect a five or 10 trillion parameter model and we open this year that would really surprise me um then list from last year can we redo this we don’t have to do the whole thing
00:45:03 Florian Brand: oh sure
00:45:03 Nathan Lambert: this is where we were at the end of 2025 who do you put in frontier now well it is Kimi and it is Zhipu DeepSeek is kind of a hard nut these days. Like I think they would be in close competitors. So I would put DeepSeek and Qwen the one as close competitors with Kimi and Zhipu as Frontier. Do you think anyone else would deserve close competitor? Cuz after that noteworthy and below like there’s so many.
00:45:38 Florian Brand: I I think we will see a surprise from MiniMax by end of the year. I think we will see a big model which like a really big model not uh M3 size but trillion parameters plus which will surprise us in terms of uh the outputs of MiniMax compared to before uh so I would still put them at close competitors by the end of the year
00:45:54 Nathan Lambert: do you think any US companies will be in the closed competitors by end of the year Nemotron I don’t think I would put there Thinking Machines closer especially if the smaller model really breaks through. But I don’t think I would put them there yet. Reflection is supposedly like only wants to release if they have a model that’s frontier. But then the question is will we get it? Like do we think that any US companies will get into this what is roughly like our top five by the end of the year? So the top five are the same but reshuffled.
00:46:36 Florian Brand: I would say it is possible uh that they are really close. Um it it also depends on what we think matters for closeness. Like I think uh Nemotron and um uh Thinking Machines will release models which act as really good base to be fine-tuned for your domain which doesn’t mean they are usable like a frontier model but they have so much utility uh that I would put them into close competitors because you would just need to find your data and uh to push the model into the right direction.
00:47:12 Nathan Lambert: Um as a I was going to think that we would make this a group of six with a US company by then like if we do this in late November I would guess that a US company pro most likely Nvidia thinky or Reflection mo does stuff that gets us to say that there is an American company in this like top cluster which would be a first time for a while.
00:47:43 Florian Brand: Yeah, I I I think that is realistic. My my one wild card is Tencent, which I think we might see something by end of the year. Uh they got some new leadership. Uh they released their Hunyuan model under Apache this time. Wait, so Tencent always had these custom licenses which disallowed anyone in the UK and South Korea and the entirety of the EU to to use their their model and also had acceptance use policy and so on. Um, and with Hunyuan and their new leadership, they got a really competent model at 250ish billion parameters. Um and I think by end of the year we might see a big model release which will surprise the people not following the ecosystem.
00:48:36 Nathan Lambert: Yeah, I I am also sure we will be in for some surprises. This is always the thing with AI and especially open models. It’s very very unpredictable. Okay, I I think this is a good place to stop. We probably should really do this quarterly. It’s not that hard and people will enjoy it. Um, but good to see you and we’ll talk soon. Hopefully in person soon.
00:49:02 Florian Brand: Peace.
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.interconnects.ai/subscribe
More Science podcasts
Trending Science podcasts
About Interconnects
Audio essays about the latest developments in AI and interviews with leading scientists in the field. Breaking the hype, understanding what's under the hood, and telling stories. www.interconnects.ai
Podcast websiteListen to Interconnects, Hidden Brain and many other podcasts from around the world with the radio.net app

Get the free radio.net app
- Stations and podcasts to bookmark
- Stream via Wi-Fi or Bluetooth
- Supports Carplay & Android Auto
- Many other app features
Get the free radio.net app
- Stations and podcasts to bookmark
- Stream via Wi-Fi or Bluetooth
- Supports Carplay & Android Auto
- Many other app features


Interconnects
Scan code,
download the app,
start listening.
download the app,
start listening.
Interconnects: Podcasts in Family
























