
Neil Movva - Making AI 10x Cheaper
Invest Like the Best with Patrick O'Shaughnessy
- Published
- August 25, 2026
- Duration
- 1h 18m
- Summary source
- description
- Last updated
- Aug 25, 2026
Discusses agents, inference, investing.
Summary
My guest today is Neil Movva, founder of Sail. Sail is building what Neil calls a token factory, an inference company designed for a specific kind of future, one where AI agents run in the background for hours or days at a time rather than answering a human in real time. In that world, latency matters less and cost matters more, and Neil has built the who…
Sail Research founder Neil Nova breaks down his 'token factory' vision—building the cheapest possible AI inference for long-running background agents by squeezing maximum efficiency from undervalued chips, creative data center deals, and low-level GPU software.
Key takeaways
- Sail Research is building a 'token factory' optimized for long-running background agents (hours/days) rather than real-time chatbots, betting that throughput-optimized inference will dominate over latency-optimized inference as AI shifts from interactive to proactive workloads.
- There is a fundamental, unbreakable hardware trade-off between latency and throughput on GPUs—low-latency inference requires tensor parallelism via NVLink (Nvidia's strength), while high-throughput batch inference can leverage cheaper, underutilized chips like AMD, TPUs, or Trainium where software expertise creates arbitrage opportunities.
- The next frontier of AI data is not internet text (already exhausted) but RL environment 'gyms' where models self-improve on verifiable tasks—coding, math, and cybersecurity—enabling recursive capability gains without human feedback.
Why this matters
As AI workloads migrate from human-in-the-loop chatbots to autonomous background agents, the competitive advantage in inference infrastructure will shift from latency minimization to cost-per-token at scale, creating a new market structure where software-driven hardware arbitrage—not chip access alone—determines who wins enterprise AI economics.
Entities
Related reports
- Etched - Building AI Hardware to Make Inference Faster and Cheaper
Invest Like the Best with Patrick O'Shaughnessy
- Gavin Baker - Watts and Wafers
Invest Like the Best with Patrick O'Shaughnessy
- Dylan Patel - The Infinite Demand for Tokens, Claude Mythos, and Supply Constraints
Invest Like the Best with Patrick O'Shaughnessy
Intelligent Report▼
Intelligent Report
Neil Movva - Making AI 10x Cheaper
Invest Like the Best with Patrick O'Shaughnessy
August 25, 2026
Report loads when you expand this section (one request).
Show notes
My guest today is Neil Movva, founder of Sail. Sail is building what Neil calls a token factory, an inference company designed for a specific kind of future, one where AI agents run in the background for hours or days at a time rather than answering a human in real time. In that world, latency matters less and cost matters more, and Neil has built the whole company around driving the cost of a token as low as it can possibly go. What makes this conversation special is that it is one of the most
Themes
- agents
- inference
- investing