Cover art for Invest Like the Best with Patrick O'Shaughnessy

Neil Movva - Making AI 10x Cheaper

Invest Like the Best with Patrick O'Shaughnessy

Published
August 25, 2026
Duration
1h 18m
Summary source
description
Last updated
Aug 25, 2026

Discusses agents, inference, investing.

Summary

My guest today is Neil Movva, founder of Sail. Sail is building what Neil calls a token factory, an inference company designed for a specific kind of future, one where AI agents run in the background for hours or days at a time rather than answering a human in real time. In that world, latency matters less and cost matters more, and Neil has built the who…

Sail Research founder Neil Nova breaks down his 'token factory' vision—building the cheapest possible AI inference for long-running background agents by squeezing maximum efficiency from undervalued chips, creative data center deals, and low-level GPU software.

Key takeaways

  • Sail Research is building a 'token factory' optimized for long-running background agents (hours/days) rather than real-time chatbots, betting that throughput-optimized inference will dominate over latency-optimized inference as AI shifts from interactive to proactive workloads.
  • There is a fundamental, unbreakable hardware trade-off between latency and throughput on GPUs—low-latency inference requires tensor parallelism via NVLink (Nvidia's strength), while high-throughput batch inference can leverage cheaper, underutilized chips like AMD, TPUs, or Trainium where software expertise creates arbitrage opportunities.
  • The next frontier of AI data is not internet text (already exhausted) but RL environment 'gyms' where models self-improve on verifiable tasks—coding, math, and cybersecurity—enabling recursive capability gains without human feedback.

Why this matters

As AI workloads migrate from human-in-the-loop chatbots to autonomous background agents, the competitive advantage in inference infrastructure will shift from latency minimization to cost-per-token at scale, creating a new market structure where software-driven hardware arbitrage—not chip access alone—determines who wins enterprise AI economics.

Entities

Related reports

Intelligent Report

Report loads when you expand this section (one request).

Show notes

My guest today is Neil Movva, founder of Sail. Sail is building what Neil calls a token factory, an inference company designed for a specific kind of future, one where AI agents run in the background for hours or days at a time rather than answering a human in real time. In that world, latency matters less and cost matters more, and Neil has built the whole company around driving the cost of a token as low as it can possibly go. What makes this conversation special is that it is one of the most

Themes

  • agents
  • inference
  • investing