559 episodes
- Contest links: General $50k https://www.chinatalk.media/p/50k-chinatalk-submission-hiring-contest we're looking at submissions on a rolling basis!
AI evals: Due
Over the past few years, we’ve seen hints of policymakers and national leaders using AI models in their actual policy decision-making. The Prime Minister of Sweden said he uses it for second opinions on policy; the German Chancellor is testing it to draft legislation; even Trump said he had it re-write a speech at some point. It’s fair to assume that senior leadership across the world, and in Washington as well, have started using AI not just for tactical or operational tasks, but increasingly for broad strategic decision-making.
While that’s exciting — there’s the promise of uplift and smarter calls on some of the most consequential decisions leaders face in foreign policy and national security — we’re also flying blind. Enormous effort and energy goes into benchmarking and evaluation for things like coding. The experiments you can run to make models better at software development are much easier to execute and much lower-stakes than running a real experiment when you’re deciding whether to invade a country or sign a treaty.
That’s why we at ChinaTalk are trying to kickstart a field aimed at helping researchers and policymakers understand exactly what they’re working with when they ask these models to support some of the most consequential decisions they may make in their lifetimes. We’re launching an evals/essay project contest to explore this theme.
I’ve brought on two expert AI eval creators to discuss why the field is important, what interesting work has already been done on how models approach broad national-security and strategic questions, and how you — as an eval professional, semi-professional, or just a concerned person — can contribute new ways to poke and prod at these models and see what they can really do.
Joining us today: Florian Brand, research engineer at Prime Intellect, and John Chen, professor at the University of Arizona, who’s done some pretty wild things getting models to start nuclear wars with each other in Civ V.
Our conversation covers:
How frontier labs are hitting eval-building limits — they’ve gone from undergrads to PhDs to field experts, and the models are now catching the experts’ mistakes.
How Civilization V exposes AI’s strategic blind spots (terrible second-order reasoning) and models’ distinct strategic personalities (Claude’s really into science!).
Why ethical prompting in Civilization V still doesn’t stop models from launching nukes.
Jordan's PresidentBench eval, where a Chinese model was nonchalant about a Taiwan invasion, while Claude wanted to keep Taiwan free and independent.
Advice for designing better AI evals and ChinaTalk’s new essay/evals contest!
Learn more about your ad choices. Visit megaphone.fm/adchoices - Six months into the war with Iran, the United States has fired 850 Tomahawks, emptied its LRASM and JASSM-ER inventories, expended essentially every ATACMS and PrSM in the arsenal, and burned through the SM-6 and PAC-3 interceptors it had been stockpiling for a fight with China. The administration's position is that the munitions are fine and the people saying otherwise are committing treason.
Bryan Clark is a senior fellow at Hudson Institute and a former submarine officer who runs war games for the Navy and its allies. Justin McIntosh is a retired Army Special Forces soldier — twenty years in, sixteen of them as a Green Beret — who now writes Mind of Things on defense, technology, and foreign policy.
We discuss…
Why 850 Tomahawks in two months emptied the stockpile that was meant for China
The leak hunt at Camp David, and why it should be a very short list
How "lethality" became a testosterone prescription instead of an ammunition line item
Japan's new defense strategy, which finally names China out loud — and bets on one-way attack drones
Whether AUKUS Pillar 1 is quietly consuming the rest of the Australian military
Why manned-unmanned teaming keeps failing in the war games, and what the CCA program got wrong
Learn more about your ad choices. Visit megaphone.fm/adchoices - The FCC has spent the past year quietly rewiring the American import market, first routers, then drones, next undersea cables, and as of last week, robots.
Adam Chan is national security counsel and senior advisor to FCC Chairman Brendan Carr, and directs the agency’s new Council on National Security. He previously clerked on the Second Circuit and served as a national security legal fellow for the House Select Committee on China. ChinaTalk’s Aqib Zakaria co-hosts.
We discuss…
How the covered list works
The case for “simple and stupid” rules over big-brained ones with endless carve-outs
Onshoring plans, allied carve-outs, and the 65% domestic content test
Why robot vacuums got swept in alongside humanoids
How AGI-pilled you have to be for the robot rule to make sense
Learn more about your ad choices. Visit megaphone.fm/adchoices - Don't sorry I put two ChinaTalk Records songs at the end!
their show notes:
On this episode of The Spillover, Sebastian Mallaby sits down with Jordan Schneider of ChinaTalk to unpack a frantic month in artificial intelligence. The conversation turns on whether the U.S. can actually “win” the AI race, with Schneider arguing that the most durable competitive edge will come from compute rather than model quality. The two debate China’s open-weight model strategy as a commercial weapon and question whether it will be possible to sufficiently harden systems against threats like cyberattacks and AI-designed bioweapons.
Anthropic’s Mythos model has led the Trump administration to take an AI-safety U-turn. Mallaby notes that as late as March 2026, “if you had said to the Trump administration that they would be trying to suppress an American AI model, decelerate the progress in the name of safety, they would have said, ‘You’re nuts.’” Schneider notes that a similar shift has not yet arrived in China.
AI models are now escaping their sandbox and raising false alarms about foreign hackers. In remarking on OpenAI’s models hacking AI company Hugging Face, Schneider comments, “They think it’s the Chinese. . . . And they’re freaking out. They’re calling the FBI.” For Schneider, the episode is a warning that “we’ve really crossed a threshold with these models,” which now have “the potential to do really dramatic harm just on their own, because we like can’t even physically watch them.”
The real AI scoreboard is compute, not model quality. Schneider argues that there is too much of a focus on the gap between Chinese and Western frontier models. Mallaby adds, “In other words, I shouldn’t be asking about how far is China behind in terms of the quality of model. It’s more a question of like, how much can China deliver the AI to users within China, given their lack of computational resources?”
China’s open-weight models are a weapon without a business model. Mallaby frames Chinese open-weight as a “counterweapon”—good-enough models pushed out cheaply to erode the economics of U.S. frontier labs. Schneider notes that DeepSeek’s CTO is pitching investors a Manhattan Project-like vision in which profits should take a back seat to the pursuit of AGI.
Learn more about your ad choices. Visit megaphone.fm/adchoices - Roel Konijnendijk of Oxford returns to talk Nolan's Odyssey, ancient PTSD, siege warfare, and the choices we'd all make in adapting.
Justin McIntosh cohosts.
YoungBoy Never Broke Again as well as James Blake each deliver an outtro song!
Learn more about your ad choices. Visit megaphone.fm/adchoices
More Government podcasts
Trending Government podcasts
About ChinaTalk
Conversations exploring China, technology, and US-China relations. Guests include a wide range of analysts, policymakers, and academics. Hosted by Jordan Schneider.
Check out the newsletter at https://www.chinatalk.media/
Podcast websiteListen to ChinaTalk, Londongrad and many other podcasts from around the world with the radio.net app

Get the free radio.net app
- Stations and podcasts to bookmark
- Stream via Wi-Fi or Bluetooth
- Supports Carplay & Android Auto
- Many other app features
Get the free radio.net app
- Stations and podcasts to bookmark
- Stream via Wi-Fi or Bluetooth
- Supports Carplay & Android Auto
- Many other app features


ChinaTalk
Scan code,
download the app,
start listening.
download the app,
start listening.
ChinaTalk: Podcasts in Family





























