· 55 min
Audio streams directly from the publisher. Open the file · Show website
Every AI model in production today has the same hidden tax: doubling the context window quadruples the compute. That's what quadratic compute complexity means in practice, and it's the reason enterprises are spending most of their AI engineering budget on context management rather than on the actual problems they're trying to solve. Alexander Whedon, co-founder and CTO of Subquadratic, joins Craig Smith to explain how SubQ's sparse attention mechanism eliminates that tax, achieving 40 times faster inference and 64 times less compute than standard attention at one million tokens, and what becomes possible when that constraint disappears. The conversation covers striking benchmark findings: 86% of what frontier coding agents do is "read steps," just trying to gather and organize context before the actual problem-solving begins, and frontier models drop well below 50% accuracy on financial document analysis at 500,000 tokens, revealing how asymmetric long context capability actually is across industries. The most commercially important argument in this episode is about enterprise data. Most large organizations are sitting on hundreds of billions of tokens of data they've never been able to put to work in an AI product, told they need a $10 million data transformation project before they can even start building. Alex's core claim is that SubQ's architecture makes that barrier no longer necessary, enabling enterprises to process far more of their data with far less curation, at a fraction of the cost. He closes with what he describes as the most important and underexplored frontier in AI right now: we are still very far from understanding what users actually want from models reasoning over millions of tokens, and the product and alignment work needed to answer that question has barely begun. Subscribe to Eye on A.I. for weekly conversations with the people building and deploying the future of AI.
Source: Episodes via the show's public RSS feed · as of Sep 9, 2026 08:00 UTC