Skip to content
PodcastsBusinessTraining Data

Training Data

Sequoia Capital
Training Data
Latest episode

109 episodes

  • Training Data

    Box's Aaron Levie: On Reinventing Yourself in the AI Age and Enterprise Diffusion

    15/09/2026 | 1h 5 mins.
    Starting a company is hard. Reinventing your company for AI as a public company with quarterly earnings results is even harder. Aaron Levie has pulled off the transition with Box and offers hard-won advice for founders. The cofounder and CEO of Box argues the value isn't only in the model; it's in the bridge from a model's raw capability to the actual workflow inside a bank, a law firm, or a pharma company. That's the case for the application layer, and Box is building it: an agent harness tuned so tightly to its own file system, permissions, and search that it beats handing the raw API to Claude or ChatGPT on both accuracy and latency. Aaron explains why token subsidies from the labs can't last, why you want a model-agnostic company routing your tokens rather than the one selling them, and why coding diffused fast while the rest of knowledge work won't. (There's no "give us your GitHub" for a sales rep.) His prediction: within five years, 90% of enterprise tokens go to work no human user ever initiated.

    Hosted by Sonya Huang, Sequoia Capital

    0:00 – Introduction

    1:55 – Are application companies the hottest neolabs?

    6:56 – Will the labs move up the stack?

    12:34 – Box and betting the company on AI

    16:50 – Hero use cases: reading a million contracts and long-running agents

    18:42 – Work slop: why AI code is embraced but AI content isn't

    24:08 – Building Box's agentic harness and the evals that matter

    27:23 – The state of the model race

    29:25 – Open-weight model adoption in the enterprise

    32:34 – Memory, continual learning, and what belongs in the weights

    37:29 – Box Labs and systems of record in a world of agents

    44:55 – Will chat be the dominant UI for enterprise AI?

    48:00 – Why coding diffused fast and the rest of knowledge work hasn't

    54:31 – Staying wired in, making a company AI-first, and what it takes to win
  • Training Data

    Making Cities Awesome: Peregrine’s Nick Noone & Ben Rudolph

    01/09/2026 | 52 mins.
    Most public safety technology companies grow by collecting more data. Peregrine inverted the model: no sensors, no new data, a business built on connecting the data and information cities already own. Co-founders Nick Noone and Ben Rudolph received more than two dozen no's before San Pablo PD let them in the door in February 2018. Today, Peregrine powers law enforcement, emergency medical services, fire and rescue, and other services in more than 400 cities and communities globally. Nick and Ben explain their north star for data sovereignty, and discuss how Peregrine's philosophy and privacy-first approach to data access and ownership preserves individual privacy and cities' sovereignty. They walk through how AI and long-horizon agents are being deployed: a cold case agent that reproduced an exoneration detectives had reached by hand, a Wisconsin county that placed a suspect using cell records buried in 300GB of evidence, identifying threats to a synagogue, root-causing an escalation in weather-related incidents, and more.

    00:00 Introduction

    02:07 What Forward Deployed Engineering Means

    03:58 What Silicon Valley Gets Wrong

    05:23 UNHCR, Dimagi And Downstream Data Problems

    08:25 Why Cities, Why Safety

    10:45 Two Dozen Nos And San Pablo PD

    14:19 Building Through Defund The Police

    18:16 The Inversion Of The Collection Model

    21:20 Data Ownership And Governance

    22:57 From Nice Search To Deep Analysis

    29:59 Agents Writing The Integrations

    31:45 The Cold Case Agent

    35:02 The Anti-Network-Effect Proposition

    38:40 Facial Recognition And Hard Decisions

    40:48 Technology For The Underdogs

    42:54 Trusting The Individual Contributor

    48:50 Ten Thousand Cities
  • Training Data

    Parallel’s Parag Agrawal: Building a New Web for AI Agents

    25/08/2026 | 55 mins.
    Parag Agrawal is making a bet that goes against two decades of web search: agents will query the web a thousand times more than humans ever have, and the infrastructure built around human clicks is wrong for them. The former Twitter CEO, now founder and CEO of Parallel Web Systems, explains why Parallel treats human click data as a bug and trains on agent feedback instead. He unpacks the counterintuitive choice to ship a search agent before a search engine, building an index incrementally, and how the new Turbo product cut agentic search to 200 milliseconds. But the problem Parag keeps returning to is economic: the ad-supported internet collapses when agents show up instead of people. His fix draws on Shapley values to pay content owners for the value their pages provide agents, with real dollars reaching publishers, he predicts, within 12 to 24 months.

    Hosted by Sonya Huang and Andrew Reed, Sequoia Capital

    00:00 Introduction
    03:25 What Is Web Search
    05:17 Why Start a New Index
    07:52 Search Agents First
    10:17 Not a Neolab
    13:14 Agents vs Google Search
    19:38 Inside the Search Stack
    28:59 Search Multipliers With Agents
    30:21 Meeting Prep Agent Workflows
    31:46 Quality Cost Latency And Turbo
    32:42 Are Agents Overtaking Humans
    34:28 Ads Model Meets Agent Web
    37:20 New Incentives For Content
    40:48 Shapley Values Attribution
    47:46 Parallel Web And Future Vision
  • Training Data

    Rich Sutton and Khurram Javed: Why AI Models Stop Learning, and How to Start It Again

    18/08/2026 | 53 mins.
    Rich Sutton, who helped pioneer reinforcement learning and wrote the seminal AI essay The Bitter Lesson, has now cofounded Oak Lab with his former student Khurram Javed. Their goal: to build agents that continuously learn from their own experience rather than from us. Rich doesn't think he holds a radical view: "I'm not weird. The field is weird." He says all learning is continual, and the field is the one that needed a new name for it. Rich and Khurram argue synthetic data is "a big mistake." Their "big world hypothesis" is that the world is massively more complex than any agent or simulator, so approximations have to be updated continuously rather than frozen at deployment. Rich calls LLMs an unanticipated scientific breakthrough, but says they represent roughly a quarter of intelligence. He says catastrophic forgetting is "totally curable" with the ideas behind their continual backprop algorithm. Khurram explains why the frontier labs can't follow: they sit in a local minimum where a new paradigm gets worse before it gets better. Their target, five to ten years out, is a trillion-parameter mind that keeps learning, stays coherent, and runs on 20 watts.

    Hosted by Sonya Huang and Alfred Lin, Sequoia Capital

    00:00 Introduction
    02:10 An AI winter, a cancer diagnosis, and the move to Alberta
    07:07 Writing "The Bitter Lesson," and what people get wrong
    09:53 Are LLMs a positive or a negative example of it?
    11:03 Synthetic data is "just a big mistake," and the Big World Hypothesis
    18:01 AlphaGo, human priors, and why prior knowledge and learning should be friends
    22:37 "Their weights never change": do LLM assistants actually learn?
    26:09 Babies, squirrels, and why no animal learns by supervised learning
    32:02 Rockets, imagination, and where paradigm shifts come from
    36:42 The Alberta Plan and its 12 steps
    38:53 Catastrophic forgetting and the cure
    43:43 Oak's biggest ambition: a self-maintaining mind
    47:56 Why the big labs are stuck in a local minimum
    49:13 If everything goes right: LLMs, many minds, and hiring
  • Training Data

    Chai Discovery's Bitter Lesson: Drug Design Is Another Scaling Problem

    04/08/2026 | 47 mins.
    Most people treat biology as a bespoke, messy science. Josh Meier and Matt McPartlon, co-founders of Chai Discovery, treat it as an engineering problem. They make the case that drug design obeys the bitter lesson: scale data, models, and compute, and the model can learn what a hand-built pipeline simply couldn't capture. The results are concrete: Chai-2 pushed de novo antibody design from a sub 0.1% hit rate to 16%, turning a needle-in-a-haystack search into something more like designing a key to fit a lock. Josh argues, counterintuitively, that biology is more verifiable than code, and explains why the goal should be more lab experiments, not fewer. Their bet: a design suite that collapses drug discovery from nine months to nine days, and arms the pharma industry rather than competing with it.

    Hosted by Pat Grady and Sonali Singh, Sequoia Capital

    00:00 Introduction
    01:52 From Discovery to Design
    03:25 Protein AI Breakthroughs Timeline
    06:04 Why Start in 2024
    10:13 Diffusion Models Intuition
    11:41 Building the Avengers Team
    15:22 Hit Rates and Scaling Laws
    25:01 Molecular CAD Vision
    25:24 Faster Design Loops
    26:32 Future Drug Discovery
    28:37 Platform Business Model
    31:14 Partnering Reality Check
    33:44 Data Flywheel Explained
    37:16 Staying Ahead at Scale
    39:44 Culture and What's Next
More Business podcasts
About Training Data
Join us as we train our neural nets on the theme of the century: AI. Sonya Huang, Pat Grady and more Sequoia Capital partners host conversations with leading AI builders and researchers to ask critical questions and develop a deeper understanding of the evolving technologies—and their implications for technology, business and society. The content of this podcast does not constitute investment advice, an offer to provide investment advisory services, or an offer to sell or solicitation of an offer to buy an interest in any investment fund.
Podcast website

Listen to Training Data, The Diary Of A CEO with Steven Bartlett and many other podcasts from around the world with the radio.net app

Get the free radio.net app

  • Stations and podcasts to bookmark
  • Stream via Wi-Fi or Bluetooth
  • Supports Carplay & Android Auto
  • Many other app features
Training Data: Podcasts in Family