AI Can Write Code. Why Isn’t Software Better?
Diogo Almeida, founder of TypeSafe AI, discusses Jev, a new primitive that embeds AI decision-making directly into software rather than automating software engineering itself. The conversation explores why current AI tools haven't delivered on automation promises, and how probabilistic programming could fundamentally change how software is built and what it can do.
Summary
In this A16Z podcast episode, Ben Horowitz and Martin Casado interview Diogo Almeida about TypeSafe AI and its flagship product Jev. The central thesis is captured in Almeida's opening question: "Where the fuck is all the automation?" Despite AI's remarkable capabilities, basic automation tasks remain unachieved, which Almeida views as a missed opportunity rather than a technological limitation.
Almeida contrasts two approaches to AI in software: coding agents that automate software engineering (making developers faster at writing the same code), versus embedding intelligence directly into software itself through Jev. Rather than replacing software engineers, Jev expands what software can do by providing a new primitive—essentially a library that accepts natural language descriptions of intent, processes state information, and outputs probabilistic decisions with confidence levels. This represents a fundamental shift from generating code to generating decisions within code.
The conversation touches on why automation has stalled despite AI capabilities. Almeida argues this stems from misconceptions about AI's current limitations, and partly from industry focus on impressive demos rather than reliable production systems. He emphasizes that OpenAI's customer service automation efforts since 2020 have failed, suggesting the problem isn't technical capability but rather the gap between what AI can do in controlled settings and what's needed for real-world deployment.
Almeida shares his personal journey from competitive mathematics through Kaggle competitions to OpenAI, where he worked on early GPT models and RLHF. He describes his disillusionment when GPT-3 didn't lead to AGI, which prompted a fundamental rethinking of AI's role. This led to the insight that the real opportunity isn't in replacing humans but in making software dramatically more capable.
The discussion covers reliability as a core differentiator, with Almeida distinguishing between uptime/SLA, determinism, and robustness (intelligent behavior every time). He argues that true reliability means developers can eventually program against Jev without constant example queries, trusting it to run autonomously.
Regarding market implications, Almeida predicts an inverse "SaaS-pocalypse." While coding agents initially sparked fears that SaaS companies would become obsolete (since software would become cheap and easy to replicate), Jev actually makes existing SaaS products dramatically more valuable by adding new capabilities. He envisions SaaS companies as ideal partners since they best understand customer workflows and user problems.
The conversation also explores deeper architectural possibilities, including how AI could be distributed throughout systems at different cost and speed trade-offs, potentially enabling a new era of probabilistic programming. Almeida references the historical failure of probabilistic programming in the 1970s but sees Jev as a practical path forward. He mentions Eric, his co-founder with a Bayesian/biology background, though emphasizes his own brand as pragmatism over biological inspiration.
Almeida addresses the relationship between Jev and coding agents, noting that coding agents excel at syntax but are poor at architecture. He sees no fundamental conflict—agents could handle syntax while humans (or AI) handle architecture, with trade-offs depending on project needs. He also discusses potential applications like voice control for computers, where Jev would constantly decide whether input is a command or text insertion.
About this episode
a16z’s Ben Horowitz and Martin Casado sit down with TypeSafe AI founder Diogo Almeida to ask a simple question: AI has become remarkably capable, so where is all the automation? Diogo argues that coding agents may help us write software faster, but the software they produce still largely works the way software always has. TypeSafe is taking a different approach with Jev: putting intelligence inside software itself, so developers can build programs that reason about intent and make probabilistic decisions rather than simply generate text for a human to interpret. They discuss why reliability is the key to making AI genuinely programmable, how this could open a new era of probabilistic software, and why established SaaS companies may be particularly well positioned to benefit. Ultimately, Diogo’s goal is straightforward: technology that can reliably “do what I mean.”
Key Insights
- Almeida argues that the central problem isn't that AI lacks capability, but that the industry has optimized for impressing humans (through demos and chatbots) rather than actually automating economically valuable work
- Jev represents a fundamental architectural shift: instead of using AI to generate code (which produces the same software humans would write), it embeds AI as a decision-making primitive directly into software, expanding what the software itself can do
- Almeida claims that coding agents make developers faster at producing the same type of software, while Jev makes software fundamentally more capable—a distinction he sees as critical to understanding AI's real economic opportunity
- He contends that OpenAI's failure to automate customer service since 2020, despite having the necessary intelligence, demonstrates that the gap between impressive capabilities and production automation is a systems and design problem, not a raw capability problem
- Almeida's view is that reliability—understood as consistent intelligent behavior over time rather than determinism—is more valuable than speed and will ultimately be the differentiator that enables new applications
- He predicts an inverse 'SaaS-pocalypse' where SaaS companies become the biggest winners of AI advancement because they already understand customer workflows and can distribute software improvements to massive user bases
- Almeida believes the distribution of intelligence in future systems will follow economic patterns similar to TCP/UDP—with 'many nines' of AI calls happening in the guts of systems (deep in infrastructure) rather than at the user-facing layer, but that builders must aim for the guts from the start
- He argues that past attempts to embed AI in software failed because software doesn't naturally accept natural language inputs, forcing developers to either output to humans or loop outputs back to other LLMs (agents), making AI 'ships in the night' with traditional software until now
Topics
Transcript
Where the fuck is all the automation? AI is so unbelievably smart, and yet it's so useless at all other stuff. It doesn't matter how much AI coding agents you use, the software actually isn't getting better. Maybe you're running it faster. It's arguably getting worse. OpenAI has been trying to automate customer service since 2020. What I want instead is smart software. I want to expand what software itself can do, such that things that should be automatable can then be automatable. My favorite thing that you guys say is, we built prod, not god. So good. Because if we had any other kind of like big lab leader, even if they had joy, they would cover that. And…
Full transcript available for MurmurCast members
Sign Up to AccessMore from The a16z Show
The $1 Trillion AI Buildout | State of Markets
A16Z partners discuss the $1 trillion AI infrastructure buildout, arguing it's driven by real earnings growth rather than inflated valuations, with adoption still extremely early at the enterprise level. They highlight opportunities across consumer agents, robotics, autonomous vehicles, and enterprise diffusion, while noting that the market's 90% gain since ChatGPT reflects fundamental business performance rather than speculative excess.
The Personal Agent Race Is Here | Anish Acharya & David Pawlan
A16Z's Anish Acharya and Assistant Benchmark creator David Pawlan discuss the explosive growth of personal AI agents, exploring their capabilities across email, travel, and finance use cases, the infrastructure needed for agent-to-agent interactions, and how these systems will reshape commerce and consumer behavior.
Building a Team at AI Speed | Harvey’s Maggie Landers
Maggie Landers, VP of Talent at Harvey, discusses how the legal AI company scaled from 340 to over 1,000 employees in one year while maintaining startup culture and values. She explains Harvey's approach to rapid hiring, emphasis on progress over perfection, founder leadership, and the specific traits they seek in candidates.
Why Companies Are Becoming a Series of Loops | Anish Acharya on Lenny’s Podcast
Anish Acharya, a16z general partner and former founder, discusses how AI is transforming company building through 'loops'—automated processes where AI handles repetitive work while humans provide judgment and new ideas. He argues fears about an AI-induced permanent underclass are overblown, and that the real opportunity lies in consumer products focused on human connection, creativity, and ambition rather than just productivity.
What It Takes to Build a Startup | Andrew Chen & Matt Perault
Andrew Chen discusses A16Z's Speedrun program, which invests in earliest-stage startups (typically 2-3 person teams working from kitchen tables) and explores how regulatory complexity and policy decisions impact where founders choose to build companies. Chen emphasizes that early-stage founders lack time and resources to engage with policymakers, creating a representation gap where "little tech" voices are absent from policy discussions.