Why AI Made Your Buyers Harder to Sell To

 

After leaving a long sales career, David Rubinstein spent a year meeting more than 300 founders across 41 countries and turned what he heard into a framework for how companies sell now. In this episode of Founded & Funded, David sits down with Madrona’s Anna Baird and Eric Wong to dig into why AI left buyers with more information and less clarity, why deals stall out of fear rather than disinterest, and how founders can cut through.

The conversation covers selling against build-it-yourself AI objections, the return of in-person selling, and why the founder’s own brand is more important to drive pipeline.

This was originally an AMA with Madrona portfolio companies, but we wanted to share David’s full sales playbook with founders, GTM leaders, and operators for a grounded look at how buying behavior shifted and where to focus when the old tactics stop paying off.

Listen on Spotify, Apple, and Amazon | Watch on YouTube.


This transcript was automatically generated and edited for clarity.

Anna: David had this amazing career in sales and as a sales leader. And then Dave, all of a sudden you’re interviewing founders, like over 250 of them in the last 12 months. Talk to me, what happened?

David: I’ve been on this crazy journey over the last 12 months. Having been in sales for a long time, I was starting to feel a little bit burnt out last year. And I’m like, “I’m going to take the summer off. I’m going to play a bunch of golf. I’m going to live my best life, and I’ll figure out what I want to be when I grow up in the fall.” And that was my plan. And what I realized really quickly is that I’m not good at doing nothing. And this FOMO with everything that was happening in the world, not being in the game was really hard, but I didn’t want a boss. And so I had to create a project for myself. So I went on LinkedIn on June 9th, and I said, “I’m going to meet 30 founders in the next 30 days.” I wanted to meet people and learn and see what was out there and add some value.

And what I learned really quickly was the majority of founders that I was talking to, so probably 80%, had never sold before, and they were thrilled to have an opportunity to talk to a go – to-market leader. They’d tell me about their business, I’d provide some perspective, and at the end they’d say, “Hey, this is great. Can we meet again?” And then they’d say, “Hey, I’ve got colleagues who’ve also started businesses. Can I introduce you?” And then the aha moment for me came when someone said, “Are you documenting your journey anywhere?” And I thought, “I’m not, but I record every call, so I will.”

Why the Old Go-to-Market Math Stopped Working

Fast-forward to today, about a year later, I’ve met with over 300 founders across 41 different countries, and it’s helped to codify what I believe is a different way of selling. There are three things that I see driving change in go-to-market that make the world radically different. The first is that the math doesn’t math anymore. I think many of us grew up in a world where go-to-market was a math exercise. You hire this many reps, they make this many calls, it generates this many meetings, it generates this pipeline, you hit the number. But over time, the same tactics got weaker. Tools made them even weaker. Scale imploded the math. And so now inputs go up and outputs actually are going down.

“The math doesn’t math anymore,” — David Rubinstein

The second thing is competition. You come up with a product, and you’re competing with an incumbent, and you used to know where that incumbent had gaps, and you’d build a solution to fill those gaps. The challenge is that the incumbent is evolving so fast, you actually don’t know where the true gaps are. And the second piece on that is that there are new entrants coming to the market every day that you’ve never heard of, and you’ll never be able to keep track of all the people that you’re competing with. That was never a problem.

And the third thing is that the buyer is now talking to more people than they’ve ever talked to before. Before, where you had a really unique idea, there are 10 other people that do something kind of sort of similar, but not exactly. That’s tough. And the buyer has less job security than they’ve ever had. So they are really, really nervous right now. They’ve got more choices, they’re more nervous. And most people will tell you that the change right now is that buyers are more informed due to AI. I don’t think that’s quite right. Buyers have more information, sure. But more information without a way to tell vendors apart isn’t more informed. It’s actually more confused. And that’s what’s going on right now. That’s the shift.

So the people that you’re talking to are not slow because they’re not interested. They’re actually slow because they’re scared. This isn’t about running the old playbook harder, this is about finding what’s the real constraint. And that’s what my SPRINT Framework is designed to diagnose — what’s really holding you back.

The SPRINT Framework for Diagnosing a Stalled Deal

So with that, we’ll start off, and you’ve got to have an acronym if you want people to remember; it’s a minus SPRINT. S = speed. Speed creates attention. When it’s missing, you get great first calls that go nowhere. So the first call ends, we’re all high-fiving each other. We crushed it. It was great. It was an awesome first call, and it goes nowhere. The fix is making sure that you’re showing the prospect a mirror. They have to feel seen. That’s what’s going to turn an education call into a business call. So when you think about speed, do you want to have calls that make buyers think that was interesting or do you want to have calls that make buyers think? And that’s a fundamental thing about speed is how quickly can they feel seen in the call? Because when they feel seen, they lean in. When they don’t feel seen, you look like everybody else. So that’s speed.

P = Problem, the problem creates urgency. When it’s missing, the buyer agrees with you and still does nothing. So think about, are you a painkiller or are you a vitamin? Most founders I talk to can understand what problem that they solve. Very few can answer the question, what’s changed to make solving that problem now actually matter? The majority of problems that exist have existed for weeks, months, years. There has to be a catalyst to make someone act now, so understanding what’s changed is really critical.

The third part of the framework is Results. Results create belief. So when they’re missing, your case studies sound like marketing. The fix is getting the how before the number. When you lead with the number, you’re marketing. If you show your work, you’re providing proof. So I’ll give you an example. We helped a customer reduce churn by 30%. That’s a marketing line. Our agents surfaced the most at-risk accounts 90 days before renewal and designed a customer program for bottom quartile accounts. Our CSM activated the plans; that’s what drove 30% churn reduction. When you have the how before the number, results are believable. When you have the number first, it’s marketing. Everyone you talk to is looking for real results that are believable.

The fourth thing is Implementation. And this is the one that no one really focuses on, but it kills so many deals. Implementation creates confidence. And when it’s missing, deals are going to go dark with no stated objection. Every founder has had a deal that went dark for reasons that probably aren’t put in the CRM. So you had a great demo, you had an engaged buyer, you talked pricing, you scheduled the next call, and then silence. So you go, “Okay, I’ll call that close loss, no decision. Close lost, budget froze.” But those deals died for a reason the buyer never said out loud. And you’re losing more of them every quarter because the buyer’s job security is worse and the noise out there is louder. The number of people that they’re talking to is significantly greater than it’s ever been.

The Question Every Executive Buyer Asks Late in a Deal

So the one question every executive buyer is asking themselves late in a deal is what happens if this doesn’t work and my name is on it? They’re all thinking that. It’s not about your product; it’s about their career. And until you answer it, which means surfacing it first, they don’t move. So the right move, and this is really powerful, has two parts.

The first is you’ve got to surface what they’re thinking before they say it. So it goes something like this, “I want to address something most buyers are thinking, but don’t say out loud: what if this doesn’t work?’ That’s when you’re controlling the call. That’s what you’re doing to make someone feel seen. When you do that, you are different than every other vendor.

Then the second thing you do is you say, “Look, there’s risk in trying this. There’s also risk in not trying it. The difference is one you can see and one you can’t.” The idea is being upfront about the things that they’re thinking about. If you say that you’re great at everything, you have zero credibility, but if you can be honest about how they’re feeling and the risk that they’re experiencing, you look radically different than everyone else that they’re talking to.

There are three, what I would refer to as hidden fears that exist within implementation. The first is headcount: I don’t have anyone to run this. They’re feeling bad about sunk cost. And so that’s where you’ve got to be able to say, “Hey, the tool that you picked was actually the right tool at the time that you picked it, but what got you there won’t get you here. And a lot of the data that we can get from that other tool, we can import it, and it will actually allow this to start a lot faster.” But you’ve got to get them confident that they didn’t make a dumb move and that making this choice doesn’t make them look bad.

And the third thing is speed to value. Everyone is looking for quick wins, whether they articulate it or not. They need to be able to show their leaders, their board that they are leveraging technology to run their business more efficiently. So whatever your product does, I would strongly advise that you have a win that you can provide in weeks, and you can define what that win looks like. But along the journey, the person who put their credibility on the line needs to be able to say, “No, no, no, I’m going to show you that this is going to work and I made the right choice and we’re going to see it in a matter of weeks.” Most founders treat implementation as a checkbox at the end of a deal. I think we’ve got to move it further up the conversations and sooner.

In 2026, implementation is a bigger part of the deal than it’s ever been. You have to remember the buyer’s default is not moving. Every minute that, as a founder, you’re spending talking about the features that your product offers without addressing the risk of moving, the default of do nothing is going to win.

The Trigger Most Founders Leave Out of Their Customer Profile

The next one is niche, and niche creates fit. So when it’s missing, your pipeline is full of maybe accounts. The test is, do you know industry, segment, role, and trigger? Most of the founders I talk to know one or two, maybe three. Industry and segment, I think, are pretty obvious; role would be: what’s the title of the person that you’re calling on? Trigger is one that most don’t have. And trigger is what is the situation that your prospect is experiencing that is perfect for your business? When this happens, you should be licking your lips like this is the best prospect we could be into because that’s what you have to understand. And most people can’t necessarily articulate what the ideal customer profile is. If you can’t articulate those things, really what you’re doing is you’re hedging. And there are a lot of businesses out there that are hedging. “We’re not really sure who we sell to yet, so we’re going to have this huge ICP that we go after, and we’re hoping we figure it out.”

The last thing that I want to highlight is trust. And trust is an interesting one because founders can walk into a room with a level of credibility. You can talk about the rooms that you’ve been in and the people you’ve engaged with and the things that you’ve seen over the course of your career, and that builds trust, and that’s a tremendous amount of credibility. The challenge that happens is trust doesn’t transfer. So when you hire a salesperson, and they watch you sell, and they try to repeat everything you do, and they can’t sell, it’s because the trust didn’t transfer. Providing them with the tools in order to be able to do that is really important.

Anna: I love the addressing the fear because… And we just had a CFO conference and the CFOs were saying a couple of things, and they’re looking at all the products they’re buying. And they’re like, “We’re trying multiple things because we’re not sure what’s going to work and which one’s going to have the fastest result.” And the other thing you’re saying is I need to see from the business cases exactly what ROI we’re going to get here or what is that time to value? So just know that, we just heard from, we had almost a hundred CFOs at a conference and this was part of the conversation, they’re the ones who are approving obviously the deals that we’re talking about here. So I love the highlights on what does this mean and the fear that people have of losing their jobs is absolutely real.

Guest Question: One of the things I wonder about is relationship selling. I’ve met a number of people who probably could use what we’re doing, but I also don’t want to be the person in the room who’s like, “Hey, I just met you. Buy my stuff and let’s get married.” How do you approach that?

Why the First Seven Minutes of a Sales Call Are Wasted

David: And a lot of what I’ve learned about my own style and what I’ve seen out there is people are worried too much about the relationship. My view, you don’t hear a lot of salespeople say it, but I believe, a typical meeting, you have 30 minutes with someone. Is the extra few minutes you spend in the beginning talking about the weather or something like that — does that actually matter? Or when someone gets off the phone with you after 30 minutes, do you want them to think, “I really like her, she’s really interesting.” Or do you want them to think, “Wow, she made me think differently about my business. She could really help me. I’d like to talk to her again.” I’ve seen hundreds of conversations. People are wasting so much time on those first seven minutes of nothing, of literally nothing.

So when I do a call, I do a first call with someone, I get right to it, and I say, “Hey, I want to be respectful of your time. I know it was hard to get this scheduled. Is it okay if we jump right in?” It was a little shocking when I first started doing it, but people were really receptive because they are busy and they did come to learn. And what I found was that if people find that they get value out of the conversation with you, you get more second conversations by providing value and making people think than you do by making people like you. And so that’s one of the biggest mistakes I see people make is over-indexing on being someone’s friend and not really on focusing on diving into their business and making them feel seen.

How to Sell Against “I’ll Just Build It With Claude Code”

Guest Question: One of the things that’s changed, I think for us in the last year is this irrational belief that everyone can build things themselves with Claude Code. The default is no longer doing nothing. It’s, “Hey, I can actually build this internally. I don’t even need to get engineering time. I’ll throw this sharp junior I have on this,” we’re fighting a build versus buy, but in a totally different dynamic. I’m curious, how do we manage that belief that someone has without disqualifying ourself in the process?

David: Yeah, it’s very hard and very real. I look at the analogy with build versus buy. If you think about Salesforce, there’s nothing that Salesforce actually built that anyone else couldn’t have built, but they continued and they continued and they continued and they had something that it’s harder today to follow that model. And that’s kind of the point of your question. I think the speed piece is going to be the first part that’s most important. So before you talk about how what you build is better, you have to make sure that they feel seen. And so what most founders do that I see is someone says, “Hey, well, I could build it.” And your response is, and I say you, but anyone’s response is, “Well, no, you couldn’t. It’s harder than you think. Let me show you why it’s so hard.” That’s the normal response, which with all due respect is wrong.

And so the way I take a situation like that is, “Really? Okay, tell me more about what you’re building.” And “Wow, that’s really smart. Let me understand how you’re approaching that. Sounds like you’ve got it all figured out. Out of curiosity, what made you decide to even take a call with me?” And sometimes when you compliment someone or where you share what they’re doing really well and that they don’t have any challenges, they’ll go, “Whoa, whoa, whoa. See, actually, I wouldn’t say it’s perfect.” “Really? Tell me more.”

And so rather than go head-to-head, because there’s a lot of scenarios where the founder’s natural instinct is, “I’ve been thinking about this problem all day, every day for this amount of time. I’m smarter at it. I can do it better.” Let them talk about what they’re doing. And then the more they talk and the more you ask some questions, you’re actually going to see some chinks start to pop up and you say, “Well, really? Okay, how are you handling that? ” “Well, that’s actually something that we’re stuck on. Can I share with you something that we’re doing in our business?” So it’s a lot of asking for permission. So you ask them, you talk to them, you find an opening. Can I share with you?

Anna: Yeah, Dave, the one thing I’d just chime in on that too, one of the things that I’m also starting to see is there is a keeping up. So everybody started like, “Oh, I can build my own. I’ve got a Claude Code. I’m going to show my bosses how great I’m at. I can Claude Code this, and we don’t need to buy any other tool.” And it’s amazing, except then the next model comes out and then there’s something else that is even faster. And that also is the other component to this to be able to say part of what we’re going to do for you is be the expert sitting on the edge of technology to make sure you’re getting the best every single time something changes because it is going to change so fast. So just one other angle that has been a little bit helpful is to be the partner that’s going to be the expert for them so they don’t have to keep staying on top of this all the time.

David: But that’s a really important point, Anna, because a lot of the companies that I talk to and see, they’re betting on the jockey, not on the horse. They’re making a decision on selecting you not because of what you do today, but because of their belief about where you’re going and are you the right person to take them there? And so that’s really important. If I decide to buy from your company today and six months from now you have the same offering, I made the wrong choice. And so making sure people believe where you’re going is important and that you’re the right person to take them there is so important than making a decision.

Eric: Dave, you want to take one off of chat? The question says that, is it better to attach to a category of technology or an old-school competitor? Because the world’s used to trying to figure out, hey, if I’m going to buy you, what am I replacing? Or if it’s truly horizontal, you take a, hey, the world is changing approach in position a little more broadly. So what’s your thoughts on positioning specific or more broadly?

David: I feel very strongly that the more specific you are, the better. Every category is getting more and more crowded, and it’s getting really hard to compete. There are lots of call recorders that exist, and everyone is a little different. I worked with a company that had a call recorder designed for regulated industries. And so it did all the same things that all the other call recorders did, but they only targeted regulated industries, working with financial services, pharma, and others. And that allowed them to build out the business specifically for those use cases. And when you start to think like that, it’s very easy today to copy features and do that. It’s very difficult to pivot your company to go after a completely new vertical.

And so that was a really good example of someone who started narrow. The companies that I talk to that I’m seeing have the most success, they are really clear about who they sell to. They’re, in many cases, building a wedge, which is the narrowest group that you can go after. And they are so intentional when they have those conversations. When we think about speed and niche, they’re often kind of linked together, as if you have a really tight niche, it’s much easier for the person that you talk to feel seen and to feel like you understand their space. And those companies are the ones that I’m seeing that are growing the fastest.

Why AI Sales Prep Gives You the Average, Not the Edge

Eric: With the top-of-funnel math not mapping anymore, what are you seeing to be the most effective ways to break through early and earn the first conversation where there is so much noise and buyers feel overwhelmed by similar messages? There’s another question that’s similar to this, with all the AI tools out there, how do sellers differentiate? Everyone comes prepared and highly customized. So I think those have similar connotations. I think it’d also be interesting beyond the full sales cycle in that first call, how do you differentiate and cut through the noise?

David: As far as differentiating, so everyone walks in, they have the same prep tools, they’re using the same AI to do the research, and they’re walking into the meeting the same. I think that’s when everyone’s experiencing. And oh, by the way, sometimes the solutions that you’re selling may be similar too. So it’s just very difficult to stand out.

What most people may or may not be aware of is when you’re going in with most of these AI tools that I’ve seen that are doing the research and preparing, the AI is typically an average of all the information that they’re seeing. So the AI that it’s providing is, this is the safe way to engage with this potential buyer. This is the safe way. It is not, in my experience, the sharpest way. And the way that you stand out is not taking the research, which oh, by the way, is better than anything we’ve ever had. How do you leverage maybe some of that research to come up with sharper insights and use that as a starting point but not the endpoint?

And this is something where it’s going to take a whole set of training and coaching for your sellers. I’m spending a lot of time with founders doing that, which is kind of saying, “Hey, how do you become sharper? How do you walk into a meeting and help someone see that mirror, see themselves?” And my goal when I have a conversation is I actually want you to feel a little bit uncomfortable. I want you to wiggle in your seat a little bit and feel a little bit of tension and feel like, “Oh, he’s onto me. He knows that I actually don’t know this answer or that I don’t have the data to make this decision.” That’s the part I think. It goes toward, am I your friend or am I somebody who’s going to show you a mirror and be honest with you and let you know what I’m seeing about your business, but also be the person that can help you solve it?

So in the situations where the product’s the same, the research walking into the meeting is the same, the way you show up for that meeting and the difference between the person that spends seven minutes talking about the weather and the person who, a minute and a half in, says, “Hey, thanks so much for making the time. Is it okay if we jump right in? What I’m hoping to accomplish today is X, and this is how we’re going to do it. Is that fair?” And asking for permission, that feels radically different. And at the end of the day, the person that they’re going to put in front of their boss is the person who’s going to make them look good, not the person that they like. So-

Eric: If you move beyond the deal, the first call, a similar concept, but a different part of the stack, if you think about all the channels out there, indirect, direct, prospecting, all the ways we build a funnel to generate deals, what’s changed the most and what types of adjustments do we need to make in those channels that have changed and what do we need to do to see the meaningful improvements?

Why the Founder Brand Now Drives Pipeline

David: One of the things that’s interesting that I’ve seen over the last year plus is that the founder brand has never been more important. What I see is the founders that have a brand, that have a perspective, that put themselves out there, that are engaging publicly, those companies, I’m seeing a significant degree of success because that’s creating pull. There are a lot of places where if the founder doesn’t have that pull, that all falls on the salesperson.

So where I’m seeing success, the first place is founders building their own brand, founders having a strong perspective about where they’re seeing success, founders being willing to talk about things outside of their own product. If you’re a founder and you’re having product conversations, again, I’m not intending to offend anyone, but you’re lessening your gravitas. You don’t have the same feel, the same impact when you’re talking about the features of your product in your public posts.

When you’re talking about bigger things, bigger industry challenges, like regulations, other things that have macro implications, or just bigger topics that your prospects are wrestling with that have nothing to do with your product, you’re in those conversations, you’re in those rooms, that’s powerful. And so what I look for is I look for founders that embrace who they are, that embrace the problems that their company is solving, but are willing to be part of bigger conversations, of tangential conversations, not in the everything has to map exactly to my product.

It doesn’t work for every category, but for a lot of the categories I’m closest to, I’m still seeing LinkedIn while there’s more and more noise there every day. I’m still seeing wild success for the people that do it well. I think that in person has never been more important than it is today. The concept of selling in your pajamas has been… The 2020s are the pajama decade for sales. And I think as we get further into the 2020s, I think we are moving away from pajamas and back into selling in person. So those are some of the things that I’m seeing.

Anna: Those are great highlights, David. I will pile on: I see founders who are a big presence who are talking about these macro issues in their industry and seeing a ton of inbound coming from that. So I’m going to double down on that because I think it is a great highlight. It’s thought leadership. They want that partner who’s going to be with them six months from now who keeps thinking about what’s next, not just the pain you solve today. You show them the pain you solve today, and then you tell them you’re thinking more broadly about this, and I know the next pain and the pain after that that we’re going after as a company.

Eric: Let’s shift gears a little bit and do another one that I always like. Dave, what are you hearing, what are your thoughts on free pilots versus paid pilots versus no pilots? What are your thoughts on T for C or term for convenience? How do you find that balance between getting new folks signed up and creating actual real durable revenue?

Free vs. Paid Pilots and Why Skin in the Game Wins

David: There’s a lot of tension there. I think the most important thing, whether you’re talking money, not money, is some form of skin in the game. And if I have an executive that’s involved or something that’s specific, I think that feels better. I tend to like some amount of money changing hands even if it’s very small because someone had to go through some hoop in order to do it, and the buyer takes it more seriously. So free, I don’t love. It’s not to say that you can’t or shouldn’t, but free usually means you’re dealing with someone who’s junior because someone who’s senior, if they want it, they can find a few dollars. So I would strongly look for opportunities where someone is willing to put some skin in the game. Even if it’s just to cover your costs, they should have enough respect for you. I think if that’s too much friction to get a small amount of dollars, I think it’s… I haven’t seen as much success on the free pilots.

Anna: I think just be careful we are hearing where people are trying multiple things because they can because they’re free. And if they didn’t have intent to buy something, they were experimenting. And so we’ve talked about, is there intellectual curiosity or actually a budget-paying curiosity? They’re coming to you because there is a problem they’re trying to solve, to all the points Dave made earlier. So please make sure when you guys are watching your funnel, I’ve seen bloat of pipeline with people thinking they had deals in play, and it was just people trying to understand what was out there versus any intent to actually buy.

This was phenomenal, Dave. I think you’re spot on with so many of the things that Lauren and Eric and I are also seeing in our portfolio.

The Best Infrastructure Moment Since Cloud

 

Joe Beda and Craig McLuckie co-created Kubernetes, the infrastructure standard that became the default for cloud native computing. Now running Stacklok, they’re watching enterprises hit the same identity, permissions, and security problems with AI agents that took the container ecosystem years to resolve, and they’re building tools to compress that timeline.

In this episode of Founded & Funded, Madrona’s Tim Porter sits down with Craig and Joe to talk through what AI adoption actually requires from the organizations building and deploying it right now. They cover why the governance frameworks most enterprises have today weren’t built for agents, how MCP provides the controlled access layer that makes autonomous agent work possible without slowing everything down, and why the architecture decisions being made right now are significantly harder to undo than they appear. They get specific about sequencing, including why developer posture has to come before knowledge worker deployment, what the FDE model looks like when it’s working, and what consistently goes wrong when a platform gets deployed without staying to close the loop. And they go inside Stacklok to talk about what’s changed about building software, what makes someone valuable on a team when agents are doing an increasing share of the work, and why the lock-in risk everyone is focused on is the wrong one.

For founders making infrastructure and talent decisions, and for the CIOs and CTOs responsible for making AI adoption stick at scale, this conversation is one of the more practically useful things you’ll watch this year.

Listen on Spotify, Apple, and Amazon | Watch on YouTube.


This transcript was automatically generated and edited for clarity.

What Are the Lessons from Cloud Native That Apply to AI Agent Adoption?

Tim: I mean, putting it lightly, you’ve both been at the forefront of two different infrastructure eras now, the cloud-native era, we’ll just call it that, and now the AI era, the LLM era. Most founders today are making foundational decisions without having that historical perspective and having lived through it before. What are some of the big lessons learned from the first time around that now you’re applying to helping founders who are building in this fastest-moving era that I’ve ever seen in my 30 years in technology?

Joe: I like to say that don’t underestimate what can be done when the people doing it don’t know what they’re doing is impossible. I think the spirit of a startup is that you have to ignore conventional wisdom in the lessons of the past to some degree, and that’s the only way to move forward and really, truly innovate. But that being said, I think that there are some lessons that we are taking away. I think open source is an enormous enabler of communities is one of those things, and it’s something that we’re investing in with Stacklok. Things both move faster and slower than you think they do, especially when dealing with enterprises.

One of the things that I did before Heptio was start this open-source project called SPIFFE around identity. That was about 10 years ago, and only now are we starting to see it come into its own with the AI era. And it’s also one of the things that we’re looking at building into Stacklok products there. So it sometimes takes a lot longer for these things to fully mature, and the old adage of it happens slow, and then it happens fast definitely comes into play here.

Tim: I think it’s a great point that thinking with a blank sheet of paper in some ways helps solve problems the right way, but having some first-principle historical perspective always helps as you solve those things. So I think getting that balance is absolutely right. Craig, what would you say are the lessons you learned from helping to build cloud-native to now helping companies address this AI moment?

Craig: Well, I think it’s definitely a moment of chaos. When you start to see a new pattern emerge, like the Kubernetes and container opportunity, it’s very difficult for enterprises to know what’s real. I think that, as we saw the evolution of the cloud-native ecosystem, there are some amazing new technologies that emerge, but there’s also a lot of flash-in-the-pan moments, and it’s very difficult for an enterprise to have that kind of discernment. I think the thing that we really focused on in the past was effectively being the field guide or the jungle guide that would help an organization bring a real authenticity to the table and help folks not just reason about the technology that you’re working on, but also reason about the broader ecosystem, understand the state of the art and really deal with the impedance mismatch that exists between a very fast-moving innovative community and the relatively slow-moving but incredibly significant enterprise tech space.

So I always think of this moment and opportunity to slip the clutch to be able to engage with an organization and bring change at a pace that they can absorb is really important. And I think trust is key. I think having that awareness of what it takes to actually make change happen in an enterprise environment is really important.

Why Did Joe Beda Rejoin Craig McLuckie at Stacklok?

Tim: Joe, you rejoined Craig to help build Stacklok. Say more about this moment. You touched on it a bit about why this problem right now working again with Craig is the biggest and best thing to go spend your time.

Joe: Well, I think there are two factors that made me decide to jump in. First of all, it was just being ready, I think. Doing a startup and being part of a startup as a full-contact sport, you’ve got to be ready to be in it, and being able to take some time off and recharge was really important, and it means I can bring my all to this. But beyond that, I self-identify as an engineer after all this time, and the world has changed so much that if I want to continue to consider myself an engineer, I have to really learn how to use these new tools and be effective with them. And the only way to do that is to jump in and get your hands dirty.

How Is the Developer Workflow a Preview of Knowledge Worker AI Adoption?

Tim: Enterprises are developing and deploying AI in a lot of different ways. Coding is at the forefront. Everybody in the organization, from information workers to developers, wants to use agents. The number of agents are multiplying by orders of magnitude, and that creates a lot of security concerns, a lot of compliance concerns, a lot of trust and safety concerns, and other governance-type things. And so you guys at Stacklok are trying to make that manageable, but say more about what you are hearing from customers that problem and how you think about the framework to go best address it.

Craig: I think there’s a lot to unpack there, but I mean, the way I’d start with it is I think we’re starting to see a hint of what the future looks like for knowledge workers through the lens of developers. What I mean by that is if you look at the way that a modern development team operates right now, you look at the way that the Stacklok development team operates, we are generating tremendous value through the use of agents. And the way that we work has changed. There’s no longer a single developer operating in an IDE. It’s a single developer taking responsibility for the work product of a pretty broad cross-section of agents that are all specialized to perform certain tasks. And the developers are really becoming the center of orchestration for these agentic systems, curating them, making sure that they can understand the work product that’s being produced, being accountable for their behavior, and ultimately making sure that they’re productively engaged in a direction that the company cares about.

And so when we think about where this is going over the next few years, I think what we’re seeing in the development space will naturally start to cascade into the knowledge worker space, but there’s going to be a lot of challenges for knowledge workers to be able to embrace that pattern of engagement. Developers are pretty flexible in their thinking. They have a relatively high-paying threshold when dealing with the configuration of their local environment. They can be educated in certain patterns, and their very nature leads them to this kind of curiosity with the tools that may or may not exist in the space with knowledge workers. And I think when we look to the patterns and practices, the most important thing is being able to trust that agent to do work for some amount of time to create value.

And for a developer, that might mean setting up an isolated environment and standing up a whole bunch of tools, and then making sure that the right controls are in place, and they’re willing to actually make those investments for themselves. You cannot ask a knowledge worker to do that in a realistic way. And so I think the starting point for us is really like, “Let’s figure out how to enable AI to be a natural part of a knowledge worker’s work. Let’s make sure that they have access to the set of systems that they need to be able to actually generate value. Let’s enable them to start providing tools that these AI systems that can actually start to do work. And then, over time, it increases the richness of those capabilities so they get to that experience where agents are a natural part of the workforce.” And it really starts with how do we borrow from some of the ideas that are coming out of the development organization and make that accessible and democratize it for knowledge workers?

What Is MCP and Why Does It Matter for Enterprise AI Security?

Tim: So that’s the big picture. Let’s zoom in. One of the places that is an important standard that enables this knowledge worker and developer experience is this MCP. So people want to use these LLMs, they want to connect to their data, they want to connect to other tools, and they want to connect other services. Enterprises aren’t going to let either their knowledge workers or their developers just to download random code off the internet or connect to it. Talk about MCP as a standard, where that’s going, and why that’s an important initial attach point to help address this broader challenge.

Craig: I often draw analogy, and I think back on the first time I saw Docker. And, I mean, Joe, I don’t know if you remember this, but I remember just kicking myself because when I looked at Docker, I could see two things almost coexisting in the same space. One was a technology that unlocked containers in a new use case, enabling productivity for developers, creating portability for their applications. But you could also squint at Docker, and you could see Kubernetes on the other side of it. And I think I had a very similar experience when I looked at MCP when the specification was introduced. It was kind of like, “Oh, wow, why didn’t I think of this?”, because it really promised two things. It suggested what the future of AI native applications might look at. If you have the LLM being the presentation layer and the view model for modern app, what’s that middleware?

What’s that middle tier that enables you to actually start accessing data? And MCP is suggesting that. And I think that’s where people start to think about these experiences, but it also represents this opportunity to control the aperture whereby agents start to access systems. And having a formalized protocol means you can now have the selectively permeable membrane that you can wrap around your existing systems and make sure that value flows in both directions, but harm does not. And I think that starting point of MCP as a gateway that enables controlled access to resources is incredibly important. And for organizations, when you’re starting to think about these workflows, people are very poor at scrutinizing the work of an agent.

“Having a formalized protocol means you can now have the selectively permeable membrane that you can wrap around your existing systems and make sure that value flows in both directions, but harm does not.” — Craig McLuckie

You think about the YOLO mode, or you think about developers being asked to reason about every access request to a resource. Humans don’t work that way. And so the gateway to productivity is having the right guardrails in place so that you can set that tool or that agent off to go and do work, knowing that it can’t get up to any harm and that it can’t egress data that you don’t want it to egress. And so I think that is the gateway to productivity. It’s not just about asserting control and slowing things down, but by actually creating those guardrails, you enable a lot more velocity and productivity in organizations. It works well for developers. And then, as we see knowledge workers starting to have to asking the same questions, it will work just as well for them.

Tim: People listening that might not know, we mentioned that Craig and Joe co-created Kubernetes. I mean, maybe after Linux, the most successful open-source project in history. Maybe that’s debatable, but it’s just become an absolute standard across enterprises. You all saw this problem around orchestrating cloud workloads. You just mentioned you saw Docker and containers, and looked ahead to what’s next. Joe, talk about that a little bit. There became lots of different pieces to Kubernetes, and I don’t want to overextend the analogy. Maybe MCP is one piece here of this broader puzzle to get to Enterprise’s ultimate goal.

Joe: Yeah, definitely. I think what we’re seeing is MCP is a great starting point. It’s both a great place to control risk, as Craig was talking about, but also when we talk to enterprises, you can give your knowledge workers a chatbot, and that’s great, but a chatbot that’s connected to your internal systems, that’s even better. As we talk to customers and they try and make this stuff work, I think they start off with just rubbing some AI on and hoping for the best, and it turns out that’s not good enough. You really do need to connect it to your external systems.

What Is an LLM Gateway and What Is “Left of the LLM”?

We’re seeing three overlapping concerns here that are driving the ecosystem of tools. The first is safety, this idea of governance and making sure that bad things don’t happen because you’re essentially coupling these stochastic systems that have built-in randomness. That’s one of the ways they work with things that are critical to your business. And so you have to be careful there.

The second thing is that you need to make them actually work. They have to actually deliver on what they’re promising. And then, as we see more and more of this, cost control ends up being a piece of the puzzle, also. And so, just like we’re seeing MCP as being one lens and one place where you can exert control both around safety and ability, I think the next piece of the puzzle is the LLM gateway. Essentially, having something standing between your workloads, your agents, whatever an agent means in that context, and the upstream LLM providers, that’s a place where definitely cost control can come into play. But there are also safety concerns there around egressing sensitive data, and also maintaining the level of flexibility so that as we see new models come out as self-hosting of some of the open-weight models becomes more tenable and more realistic, depending on the use case, that LLM gateway becomes a very strategic choke point across all three of those factors that people are worried about.

But then, if you zoom out and look at the larger picture, I think when Chat GPT first hit the scene, so much of the focus was on the LLM itself and everything that’s happening there. But if you draw a diagram of what happens between a user’s intent, let’s take a knowledge worker using a chatbot, for example, and that LLM, there’s a bunch of stuff in the middle there, and it’s that stuff in the middle that we’re seeing a ton of innovation. We’re actually seeing a ton of advancement in terms of the capabilities. When you look at something like the coding agent, something like Claude Code, the LLM is a big part of it, but also what’s happening as part of Claude Code itself is a significant piece of the puzzle. It’s going to be a similar thing as we start seeing knowledge workers start using these systems in a more in-depth way.

It’s not just a direct line to the LLM. There’s a bunch of stuff happening in the middle there, around context management, memory management, being able to upload and organize documents, and share those. How do you actually build skills and share those skills with your teammates? That peer-to-peer enablement, defining what the agent is. All of those things are happening left of the LLM in that diagram, in my mind at least, of that journey from the user to the LLM. And that’s a place where there’s enormous appetite, I think, to create something that is both open and aimed at enterprises for them to run behind their firewall. And I’m really excited about exploring that and partnering with users around figuring out what it makes to really enable them to provide those Claude code-like moments, but for the rest of the company.

The Mainframe vs. Open Platform Question That Will Define the AI Infrastructure Era

Tim: That’s really helpful, and it mirrors in security broadly this notion of shift left, get closer to the developer, designed for security and safety earlier in that process. You mentioned LLM gateways, maybe just in terms that another founder can appreciate. Everybody agrees on, “Hey, I want my team to use more of the LMs. Hey, I want them to use agents. Hey, we’re going to get more sophisticated around harnesses around these things.” But I also hear, “Does an LLM Gateway solve the security and safety? Do I need to have an observability and monitoring solution and just look at what happens either on the network or around the workstation? Is this an identity problem? Is it a registry?

There’s a white list like, “Hey, these are known good tools or other data sources, we’re okay.'” But I think to your point, there’s a lot of “stuff” that happens there. Maybe how should a founder or even an enterprise think about is it the right set of these things that works for you? Or is there some other framework to think about it, Or debate that’s even the right set of things.

Joe: No, no, I think the answer to all of that stuff is that we need all of that stuff. And I think there are parallels here with the Kubernetes cloud-native world. So if you look at the CNCF landscape, it’s this eye chart with a gazillion logos. And it’s intimidating for enterprises to figure out, “How do I even get started with that?” And I think we’re seeing a similar thing in the AI space where there’s so many choices and all of them have these impacts in terms of security and cost and the ability to get stuff done. And I think a lot of companies are looking for a starting point and a guide to help them navigate this territory.

Craig: I mean, I think about what Kubernetes offered and why the CNCF landscape map became such an eye chart is it really represented a principled interface so that someone could specialize on one thing that’s relevant to every Kubernetes environment like observability or some facet of security around eBPF or novel ways to actually drive isolation. And it created efficiencies so that organizations could start to specialize and actually deliver value across the world. I think the big question that we have to ask ourselves as an ecosystem, as enterprise organizations, as startups, is are we moving back into the mainframe era, where the world is highly vertically integrated? Or are we going to be in a world where a platform starts to emerge that enables organizations to create specialized value and enables optionality so that you’re not necessarily tied into one system?

“Are we moving back into the mainframe era, where the world is highly vertically integrated? Or are we going to be in a world where a platform starts to emerge that enables organizations to create specialized value and enables optionality so that you’re not necessarily tied into one system?” — Craig McLuckie

And it could be around inferencing. Obviously, the frontier models are running at the absolute forefront. There’s incredible innovation happening there, but open-weight models are maybe a generation or two behind and they offer incredible price performance advantages for some organizations. So being able to hold option value on that is important, but that’s not going to work if so much of the value is coupled in what Joe describes as being left at the LLM. And so I think we do need to see this platform architecture emerge in the open source. I think MCP is a seed crystal that offers an opportunity to grow this ecosystem around it, but it’s obviously insufficient.

And I actually think if you look at it, Kubernetes offers a lot of the hints, not just in terms of how to build around a community, but literally just Kubernetes. It actually is a pretty principled control pane. It has some really nice properties in terms of reconciliation-driven action. And I think we will naturally start to see this community emerge around Kubernetes itself where we start to introduce first-class primitives that unlock this class of use case and value.

Why Is LLM Lock-In the Wrong Risk? Where Does AI Lock-In Actually Live?

Tim: So Joe, you mentioned shifting left of the LLM. We talked about how a big driver of the whole cloud-native wave was allow interoperability and to make enterprises feel like they’re not getting locked into a certain cloud or a certain technology. There’s a bunch of lock-in potential considerations here, either in like a technology stack choice or even the LLM that you want to use. My goodness, every week there’s a new model that comes out that leapfrogs the other one. And I would imagine that customers want to stay very nimble about, “Hey, whatever’s the best fit for me, that’s the model I want to be able to use within a framework that’s safe and reliable, et cetera.” Is that something that you think is an important consideration here as well?

Joe: Yeah, exactly. I think as this world unfolded, I think there was a certain level of relief that it was actually so easy to switch between LLM providers, and people felt like they had that optionality there, but the stuff that is sort of to the left of the LLM, that ends up being much more tied to your day-to-day existence. Now, when we talk about developers, they’re really good at switching tools. They do it all the time. And we’ve seen like, oh, folks move between different coding agents and different IDEs to be able to do that. But I think when you’re looking at this through the lens of an enterprise delivering tools to knowledge workers, they’re not going to be as nimble as developers. And so if you train all of your sales team on how to use Claude or how to use OpenAI, you’re not just tying yourself to that LLM provider, you’re tying yourself to the entire experience delivered through the set of front-end tools provided by that.

“If you train all of your sales team on how to use Claude or how to use OpenAI, you’re not just tying yourself to that LLM provider, you’re tying yourself to the entire experience delivered through the set of front-end tools provided by that.” — Joe Beda

So really, we’re seeing in that much more knowledge worker-facing or consumer-facing space, much more vertical integration that can lead to lock-in. And we think that doesn’t really serve the enterprise as well. We think that they want to own the destiny of how they train and how they inform and how they build tools for their knowledge workers beyond the developers, and they want to make sure that that is insulated from the actual upstream LM provider that they’re using. And so that’s an evolution that I’m going to be looking forward to and something that we’re very interested in at Stacklok.

Tim: A great point. And I think Stacklok enables, “Okay, whatever’s the best model for you, plug it in.” But also if, “Hey, today MCP is a great way to,” but that might be a different standard down the road, and you’re also abstracting away from that as well. It’s like, “Let’s take whatever the best interoperability standard, the best model, and make sure that enterprises can access that as the best-of-breed in that scenario.”

Joe: Yeah, exactly. I think MCP is a great starting point, but this world’s moving fast, and we want to make sure that we have all the right touchpoints so that we can enable our customers to stay abreast of the latest and get the best out of it. So one of the things that we see is that, let’s say, you have an enterprise that’s doing the AI-enhanced banker. So it’s a human, and they want to give that human a tool so they can do stuff faster, more effectively, access data, and join all that information together. That’s really powerful. And traditionally, oftentimes we’ll have corporate controls, but there’s an overlay of your company handbook says, “Don’t do X. And if you do this other thing, we can actually maybe arrest you, and you go to jail.” There are all these sorts of human-level constructs that keep people from doing things.

Well, those don’t apply to AI, really. An AI can move faster than a human. It can access a lot more information before you understand what’s happening. Also, you can tell it to follow the handbook, but it doesn’t always follow the handbook. It is a stochastic system, and it’s not afraid of going to jail, so the set of controls that we have in AI that we need to advance those things. So, for example, if you have this AI-empowered banker and they’re on a call or in an interaction with a customer, you want to make sure that as they enable tools to act, say, sensitive customer data, they have some sort of proof that there’s a live customer on the other end and that they’re accessing this on behalf of a customer. So that’s another piece of information that there’s a transaction with a customer that you have to carry from the agent interacting through whatever interface that they’re using, through maybe an agent calling an agent calling an agent calling an MCP server calling a database.

“An AI can move faster than a human. It can access a lot more information before you understand what’s happening. Also, you can tell it to follow the handbook, but it doesn’t always follow the handbook. It is a stochastic system, and it’s not afraid of going to jail.” — Joe Beda

And then, as you get to that database, you may want to restrict to say, “Oh, only return information for this single customer.” And so that entire chain, you have to carry that information through, and we don’t do that typically today. There are a few places where enterprises can do that type of transaction token, but it’s pretty rare. And being able to plumb that through that entire stack is what we think a key thing to be able to really build trust that you can give knowledge workers access to sensitive information, which they already have, but through agents in a way where they can trust that things aren’t going to go wild and do something that is unsafe.

Where Should Enterprises Actually Start With AI Agent Adoption?

Tim: You’ve started with banks, different customers have different priorities, and where you start. And that’s maybe what I’d like to talk about next is to try… It’s all complicated, it’s going to evolve, but where to start? So, from just a user adoption standpoint, do you see patterns within your customers about where they’re most focused initially?

Craig: Just talking about us as a company, we see the developer space as being highly contested and notoriously hard to monetize. This is true of developer technologies universally, but it also is the place where an organization can understand and learn the full value of the system, but also create leverage for themselves. So one of the patterns we see is during the days of Kubernetes, we would come into an organization, and we would like to partner what’s now known as FDEs. We used to call them field engineers back then. Now, they’re called forward deployed engineers with a platform, and they would help an organization get over the learning curve. Our FDEs are now outperforming our expectations massively because they’re not showing up and just doing work. They’re showing up and bringing skills. They’re showing up and helping to educate the team around how to actually operate their own agents.

And so getting the development posture for an organization right from the get-go is really important because it also accelerates the use of the technology itself. It accelerates the operationalization of all these platform pieces and it accelerates an organization’s path to value. And so I think for most organizations it is about getting the engineering flow right first, making that transition so that you’re starting to use agents to generate asynchronous value, getting into that world where the engineers feel like they’re flying. And once you’ve there, you’ve now established a set of patterns and you can start to extrapolate that for knowledge workers. And you can also start asking questions around what’s working for developers that just won’t work for knowledge workers. If you start looking at technologies like Claude, which is really amazing, so much of the work is being done locally like the developer’s desktop is the anchor point for integration.

It’s where memory management and session management happen. That’s probably not going to work for your knowledge workers who need to create experiences on phones. So you can start to extrapolate what good looks like, but you can also ask the question, what needs to change and what needs to be built to create those outcomes for knowledge workers? And then, I think it has to be grounded in a very specific, concrete use case. I think that where enterprises fall down, and some of the things we’ve seen, is that it is kind of by modality. If we get a chance to work with an enterprise, not just around a platform deployment, but identifying the first cohort of users, the first agent, getting them to the point where they have the sophistication to actually create the outcomes, great things happen.

But if we just deliver a platform and walk away and expect that enterprises are going to be able to turn the corner on agentic development and understand the mechanics that are actually building something that’s precision-oriented to create value in the space, things don’t necessarily go as well. And so I think you really do need to close out the arc and be patient with identifying, first getting the development team that you’re working with to a point where they’re operating excellently so that you can bring in the skills so they can actually start to feed themselves, and then identifying those initial use cases and working those to completion before expanding the program up.

How Should Enterprises Configure AI Agents? MCP First, LLM Gateway Second

Tim: Within a use case, now this is sort of a technology lens on it, is the place you advise customers or other founders you talk to around securing this whole process within use cases, start with this notion of LLM gateway or start with this notion of like, “If you’re really sure that everything you’re connecting to over MCP is secure, you’re in good shape.” So from a product technology standpoint, what’s the analogous here, where you get started down this journey?

Craig: I mean, I think the MCP section is an important starting point. And then, the way I think about it is like, what are you trying to optimize around? And I think the thing that you need to optimize around specifically, when you’re working with development teams, is how long can an agent do useful work before I have to go back and intervene?

Joe: And the factor is trust there.

Craig: And the factor is trust. So the first thing you have to do is get to the point where the agent is able to run for some amount of time without you having to constantly watch it and approve every tool access and integration. And so starting with some very coarse-grained controls around, “I’m going to configure this thing to run in a relatively isolated environment, I’m going to be judicious around the tools it has access to, knowing that they have read-only access and the ability for the system to start egressing sensitive data, et cetera, is limited,” is the logical starting point. Because once you’re at that point, you are now creating efficiencies for yourself.

“The gateway to that outcome is MCP. It is a natural starting point because it gives you the guardrails that enable you to go fast.” — Craig McLuckie

And so I think you always start there, and the gateway to that outcome is MCP. It is a natural starting point because it gives you the guardrails that enable you to go fast. Once you’ve got to that point, you can start asking questions around, well, LLM Gateway would be another is the immediate next question. A lot of organizations that make that first transition start to experience a little bit of sticker shock with token consumption. They want to be able to start doing intelligent routing and say, “Hey, these types of requests, you probably shouldn’t be using Opus 4.7. Maybe a cheaper model makes more sense.” And then, you start to work into creating the more complete system over time.

How Is Stacklok Building Software Now? What Has Changed About Engineering Teams?

Tim: Another topic is you referenced how forward deployed engineers from field engineers, how that has evolved, and just using AI in everything you’re doing internally. Maybe just talk about that. How is Stacklok building software? Are you in the “Hey, it’s basically the same number of developers,” and everyone is doing 10 or 50 times more? Is it a mix where your team is mostly directing other agents to do things? How are you driving efficiency and productivity using AI within Stacklok?

Craig: It’s wildly different. It’s unrecognizable. The profile of what constitutes a successful developer is unrecognizable. You always maintain as a startup fund, you always have in your head your life raft exercise of who your most valuable engineers are. That’s become completely Freaky Friday’d up. Everything is different. And I think the starting point is recognizing that we treat engineering now as a performance sport. You don’t need a hundred football players to win the Super Bowl. I mean, I’m not a sportsball guy, so I don’t know how many you actually do need. I’m pretty sure it’s not a hundred. So you’re able to actually be far more effective with the resources you have. I think there’s almost this Ballmer Peak. We used to joke about how much your code performance improves with drinking. But I think there are performance elements to having smaller teams, smaller, more senior teams that are able to reason about not just the code. Joe describes it as product engineering.

“The profile of what constitutes a successful developer is unrecognizable. You always have in your head your life raft exercise of who your most valuable engineers are. That’s become completely Freaky Friday’d.” — Craig McLuckie

The wisdom to be able to balance understanding of the customer with what you’re driving is very different. PM’s role is very different. Their ability to scale is disproportional. We have systems where we’re obviously very careful about what conversations are transcribed as we are interacting with larger organizations, but we will make sure that we take great notes and then put that into a channel, and then make that available to all the engineers so they don’t have to speculate. They can ask questions about which of the 60, 70 customers you’ve talked to over the last month want X, Y, Z. That is immediately available to the engineers when they’re doing research on features.

And so everything has changed. The scale of management has changed. The tools that engineering management use are fundamentally different. The way that we handle planning and the set of skills we’ve built around planning and design are completely changed. The way that our designers work is just foundationally different. We no longer work in Figma. Our designers produce working systems. And then, they take those working systems to the engineers and then they throw them away. Or they’ll take subsets of it, but our designers-

Joe: Code is not nearly as precious as it used to be.

Craig: It’s not precious, and free code is free. And it’s so much more evocative to just have a working system that your designers produced than to try to extrapolate on Figma. We just don’t do that anymore. We’re able to… The FTEs themselves are contributing significantly. They’re taking these skills that the product team is producing, and they’re now able to expand the product surface as they’re working with customers, because they’re just using the same tools and practices, and they’re fitting in. I mean, it’s incredibly exciting. It’s wonderful. It’s scary as well. It’s just a whole new world.

Tim: Joe, you said code is less precious. Say more about that because, obviously, you can generate a lot more code. Sometimes with companies, I hear, “Well, that just creates a different bottleneck, which is code review and code quality.” And then, I hear others say, “Code review? I don’t look at the code. You say AI reviews it, we just test it with AI.” Now, of course, it depends on what you’re building. If you’re building for pacemakers versus a consumer website, then the stakes of if it goes down are differently. But how do you balance that?

Joe: Well, I think the job has just fundamentally changed. One of the things that I joked about in the pre-AI era is that the job of somebody who’s a high-level principal engineer, architect level is that oftentimes you’ll draw a bunch of boxes on a whiteboard and then you’ll have teams go off and compile the whiteboard. And these days, you still have to actually know the structure of what you’re building. You still have to have your pulse on what’s happening, but you’re operating at that compile-the-whiteboard level. Instead of having engineers do it, you have agents that can go off and compile the virtual whiteboard for you.

But one of the lessons that you learn as you’re managing these large teams is that you can’t read all the code that’s being produced, similar to how you can’t read all the code that’s being produced by agents, but you build an intuition around what are the key places where I do want to keep an eye on things, schemas for your databases, API contracts, level of test coverage? You get this sense of like, “If I pay attention to these things and something looks off, then I go ahead, and I dig deeper.” And so I think part of engineering now is building those more senior skills of how to actually manage a lot of work and find the real inflection points that you want to watch. And that’s not something that a typical entry-level engineer knows how to do.

There’s another interesting aspect here in that if engineers now have like six coding sessions talking to an agent all going at once. I started at Google 20-some years ago now or not even, well, not quite that long ago, but like maybe 15 years ago, and there was this idea of 20% time there, this idea that if you give engineers room to play, they can do amazing things and you want to give them that freedom. Now, the idea at the time was like one day a week the engineer goes off and actually does a personal project that’s related to the company, but also is more exploratory. Well, now if you’re running five or six agents, you can take one of those agents and have them do something more exploratory. So there’s still room to bring that human element and that intuition of exploring new spaces, but we just have a much expanded tool set to be able to do it with.

Tim: You know, every successful company in my experience was good at bringing together a diverse set of people, a diverse set of perspectives, and together you can do great things. And if there’s a monoculture, you mentioned like the job of a senior engineer architect used to be put it on a whiteboard and then a team goes and compiles it. Okay, so today you can have your-

Joe: I’m being tongue-in-cheek there.

What Two Traits Separate Elite AI Teams From Everyone Else?

Tim: For sure. But I mean, just to extend that analogy, today you could define it and have a team of agents go do that. There’s an argument like, “Okay, so that means you just need more senior people who know how to do that.” But then, we also hear like, “Well, no, we don’t just need these ‘later-in-career people’ with all this experience.'” That’s great. You do need that. But sometimes the earlier-in-career people who were born in AI. They know how to do this, and we need more of that, just like, “This is a new world, and it’s completely my first language.”

So that’s maybe another form of diversity on the engineering team is like people that were really AI-native and maybe don’t have that experience, but really know how to use these tools. And then, there’s, of course, the, “Hey, I know how to build software and manage a team and now that team is part people and part agents.” How do you think about that and getting the right mix on the team, Craig?

Craig: It’s really interesting because it’s like I look across and I won’t constrain my attention here to just engineering, right?

Tim: Yeah, thank you. Yes, absolutely.

Craig: I look at my recruiter at Stacklok and the work he does. I used to think he was just a great recruiter. He would go and source fantastic candidates, build relationships, and do a lot of things. He is now able to do that job. He was massively outperforming in the recruiting context. And so he has taken his curiosity and has led him to other areas. He is now very effective in operating in our sales organization because a lot of the work that he did translates over. He has access to new tools, so he’s constantly able to build agents and run them. And he started to self-identify almost as a developer himself. He’s not writing code, but he is building agents, and those agents are creating value in a lot of different domains.

“You absolutely have to have the depth of talent to have taste … The thing that’s really valuable is the ontology that is constructed, the views that are used to generate the code, because the code can always be regenerated from the ontology. You’ve got to know what good looks like.” — Craig McLuckie

I think you absolutely have to have the depth of talent to have taste. When you’re thinking about like, “Where is the IP, where’s the value?” It is obviously encapsulated or expressed in the code, but the thing that’s really valuable is the ontology that is constructed, the views that are used to generate the code, because the code can always be regenerated from the ontology.

Joe: You’ve got to know what good looks like.

Craig: You’ve got to know what good looks like. That taste, that sensibilities, the instincts that these agentic systems don’t necessarily demonstrate now, they may over time, but they don’t have it now. So you really need that depth in your engineering organization, and you are able to start scaling that out relatively quickly. And I think the other piece you need is this kind of curiosity that comes with that generalist mindset, this willingness to be unfettered in terms of how to think about value creation, being willing to challenge the boundaries of your role, the boundaries of conventional wisdom.

And so I think you do need to bring those two capabilities together, and you need great teams are going to have the right mix of both taste and sophistication, but also that exuberant curiosity and willingness to challenge, that innovative flair. And it doesn’t have to be constrained to any one function. I think the thing that’s surprising me the most is just how versatile some people are when they have access to this class of tools.

Tim: Well, thank you both so much. This has been such a great conversation. I mean, from where I sit, this challenge of whether it’s Madrona, whether it’s an enterprise, whether it’s startups we talk about, we want our teams and our people to deploy AI agents to use it, to lean into AI every way they can use to make their job more efficient, more productive or even more enjoyable, but how do you do it safely? How do you do it securely? You two are a dream team. It’s fun to see you working together again, addressing this really critical problem, and also just these lessons learned and things you accomplished in the past, how there are some good things there, but also now it’s different in how we are adopting, just like the rest of the world is, to move super quickly to build a great business by helping customers. So thank you both very much.

Craig: Hey, thanks for having us.

Joe: Thank you, and thanks for being part of the journey, Tim. It’s always great to work with you.

AGI Needs Formal Reasoning. Carina Hong is Building it at Axiom.

 
There’s a theorem being tested about how AI reaches general intelligence. Carina Hong’s answer: through mathematics.

Carina is the founder of Axiom, and in less than a year of building, her team’s AI has scored a perfect 120/120 on the Putnam mathematical competition — a test where more than 50% of brilliant undergraduates score zero. More concretely, Axiom Prover has reached 98.93% on a Lean software verification benchmark that leading alternatives solve at 11–12%.

In this conversation with Matt McIlwain, Carina explains her central thesis: that math and code are the two pillars of the digital world, and that any AI infrastructure missing a formal verification layer is structurally incomplete. She walks through the history of verified AI research at Google, DeepMind, OpenAI, and Meta, and explains why each effort stalled just as commercial pressure mounted. She describes what makes hardware and software verification the natural first commercial market, and what Axiom discovered when they tested their prover against circuits that industry-standard formal checkers could not verify.

For founders and operators trying to understand what’s actually changing in AI capability, and for anyone building in adjacent infrastructure spaces, this is a map of where the frontier is and where it’s heading.

Listen on Spotify, Apple, and Amazon | Watch on YouTube.


This transcript was automatically generated and edited for clarity.

Matt: It’s great to be with you, Carina. To me, what’s interesting, having been an investor from day one in Axiom, is there are these debates about what is artificial general intelligence and what is artificial super intelligence. Curious how you define, and whether you choose to define it as AGI or ASI, how you define that, and then we’ll go from there.

Carina: Thanks, Matt. It’s always great to chat with you. I feel like the word AGI actually currently, it has this emotion element in it. It’s the belief that reasoning capability can be so abundant, and that you can have copies of any genius you pick in the world. And it compounds from there. You can have these geniuses work with each other, collaborate, and then ideas fuse between different scientific fields.

And then that capability then trickles down to the day-to-day life, for ordinary Joe to enjoy the same amount of capability and speed of knowledge, discovery, generation, and also applications. So that’s why I think a lot of people when they say, “We are AGI-pilled,” that’s what they mean. They believe that future is going to come.

Then there are also other people who define AGI as more practically. When that capability really flows through everyone’s consumer’s day-to-day life and enterprise workflow, because I think we’re quite ahead in the capability, and there are so many people building applications, which is great, but it’s not like a time that someone, say, in Alabama, would use whatever that is the most state-of-the-art coding model, and say UN code verification tool to build a project that is in their passion. And we’re not there yet.

Matt: Yeah. I was just in Georgia, and I think people are using the state-of-the-art of AI without realizing they’re using the state-of-the-art of AI. And of course, as you know, and we’ll talk about, the state-of-the-art is moving rapidly.

Carina: That’s right.

Matt: So you’ve made this provocative statement that math is AGI. What does that mean? And what should we take about the statement?

Carina:Yeah, I think that is the DNA of Axiom. We’re very clear that, in the future, we’re going to see so many promising markets. The first one is best versus verification. In the future, we’re going to see more markets, but the DNA of the company remains math. And in our view, we think that from math and code, which, if you understand it in a sort of Curry-Howard way, are actually twins. You can then do a lot of things in the kind of use software and generation verification to control things that are in the real world, like physical engineering stuff. And then you can go to everything else. So from our worldview, math is AGI because that is the starting point of reasoning.

A lot of other people have different definitions of AGI. I would say that if you take the view that math is AGI, you have your action plan, you execute that a lot faster, because math is a lot cleaner. So, I was on this AI for science conference in San Francisco. It’s a really great panel, and everyone is doing experiments with the physical world. And obviously, the latency of iteration is quite long. In math, just like in coding, the digital world, your iteration time is a lot shorter.

And the other thing that’s really interesting is that because math is data-scarce compared to the other domains. So you hit some of the roadblocks a lot faster, which allows you to debug those roadblocks. So that’s, I think, one interesting difference between math and coding. With coding, we were able to make a lot of progress at the start locally, just because there’s a lot of coding data.

Matt: There’s a ton of data, yeah.

Carina: That’s right.

Matt: Repositories everywhere.

Carina: And that’s just massive amounts of data compared to, say, informal math data, or formal math data, which is even smaller.

Matt: Yeah, say more about this math scarcity, data scarcity issue, because I think that that’s one of the areas that you all are creatively trying to address.

Carina: That’s right. So, math is such a data-scarce domain. I think a lot of researchers at the beginning were kind of blindsided by the idea that, well, we just take whatever has more data in math, and take that research path, which is informal math. So what they do is they start scraping textbooks, and papers, and chain of thought reasoning data through expert data labeling companies. And they try to basically build a math model that way.

It worked for a bit. So, it worked for, say, high school level, where you really just have a lot of data. And it’s a lot easier to say find human experts around the world at the high school, international math Olympia level, than say at the postdoc, or junior researcher, algebraic geometry level. We thought a bit more first principally. We were like, “Well, we think we need structured data.” And even if that means there isn’t a lot of structured data right now, how do we invent to create structured data? Instead of going to the unstructured route, because today we have a lot more unstructured than structured. And the structured data looks a lot like code. I mean, it’s mostly written in the theorem-proving language called Lean, which has its own history, that’s separate from the deep learning development timeline.

And just wasn’t any Lean data. And people had to hand-code the undergraduate textbooks. They start with algebra, and then they did analysis. I remember this was 2020. I was at MIT. Fast-forward to, I think we started July last year. And we started under the belief that we can try to make progress on formal proving pretty fast. We were unclear what was the moment that formal math can catch up to informal math.

Matt: So take us a little bit, because there’s, I think, two interesting things here. One is this idea of formal and informal AI. Let’s unpack that.

Carina: It’s almost like a religion.

Matt: Yeah, go ahead.

Carina: People just debate about it. And I cannot convince someone who does informal math to look at the formal math. I’m like, “You’re missing out.” Really are. And it’s usually the people who are working on reinforcement learning for other domains, like coding. They know that if you do RLVR, it’s incredible, they know that math is most interesting when it’s proofs, and it’s most interesting if you assign reward for the proof instead of just a numerical answer. So, they tend to join this movement.

Matt: Some people ask the question, “Why now? Why is this all important today?” And I usually start with, well, in a world where we’re building these LLMs, which by definition are non-deterministic, they’re predictive systems, you need some kind of system that’s a prover, that’s a verifier. How do you think about that on the why now question?

Carina: So, first of all, we could not have been early. I think that’s right. Because we’re seeing the… I mean, informal is doing pretty well. It’s at a time when we needed to get started to play catch-up in the data-scarce domain.

I think around the time of 2024, there’s this amazing achievement of AlphaProof. And they use something called Monte Carlo Tree Search. That is quite an expensive system for a startup, expensive method for a startup to do. It was not clear to me that was the only way. Now, in terms of why the urge to make it happen, both from my perspective and a lot of others who join us, I think it was at a time when people had more faith in structured data. So, I think if we talk about the latter half of 2024, people thought about coding as just another enterprise application. Coding is no different from developing an AI financial analyst. It’s no different from developing an AI McKinsey consultant.

Now, obviously, we know now coding’s a lot different than that. I don’t want to talk about recursive self-improvement, but perhaps if you can agree on, at least, that is like autocatalytic. So, coding does have that advantage. We’re not far, I think, from building an AI that can do AI research. Now, the definition of that is very loose. You can define that as getting publications, or even spotlight papers, and ICLR, ICML, NeurIPS. You can also define that as being as good as one of those researchers out there at frontier labs.

Matt: Yes. And some of the great researchers in academia, I was talking with our mutual friend, Carlos Guestrin, and he and some of his teams have been working on some of this area of continuous learning with reinforcement learning.

Carina: Yeah. But I’m sort of fortunate and grateful that the RL coding community, they do strengthen humanity space in this other religion. And I think that’s one reason. So the code gen development. The other is the informal reasoner getting stronger, which actually also helps us. In the AI for math community, there’s this paper called “Draft, Sketch, and Prove.” The idea is after you produce an informal proof, and you try to produce a Lean sketch, and you try to fill in the stories that are the voids of the Lean sketch, it works better. Now, the first part is obviously beneficial if the informal models are doing better. So, September 2024 was 01.

Matt: And so in a sense, it’s iron sharpens iron. They’re making each other better.

Carina: That is correct. And I think that is something that Axiom can talk about a lot and a lot, right? And we’re not talking about the only formal Monte Carlo Tree Search rigid stiff approach. I do think that we are able to get the benefits of both systems. And I think the third thing is just like after September 2023, Lean4 becomes a lot more workable from an engineering perspective. It is still surprisingly not workable. That’s why we built AXLE, Axiom Lean Engine, which is this sort of infrastructure. There are a dozen tools, and we’re adding tools to it on a weekly basis. I believe there are people around the world, literally all across the world, using it. And people are using it interestingly with their favorite model. We see the choice of literally any model. It can be a large language model that’s off the shelf, it can be the specialized model. Even some of our competitors’ models.

Matt: I love this.

Carina: But they use all these models with the infrastructure of AXLE because it does solve a pain point that Lean is slow.

Matt: We were together the week you launched AXLE. And it’s incredible to see the progress.

Carina: Yeah. The growth curve of the user metric is quite high.

Matt: And I think it’s at building that ecosystem out where the iron is sharpening iron that is, I think, going to help the whole community benefit. I love that framing. And I think for some of the folks in the audience that are saying, okay, how much is this getting applied, or beginning to be applied? This tension as it were between math proofs and coding. There are companies, big companies in some cases like the Amazons and Googles that are already working on this area. You’re a little bit familiar with that. Maybe talk about what they’re doing, and then maybe where you think the world’s going.

Carina: So, I think, I certainly draw a timeline of verified AI. So 2016, Google made a pretty important bottoms-up strategic investment. There are good papers coming out of it. Playing with a language, they’re improving language called HOL. HOL List is one of the papers. And then there are white papers, I think in 2019 by Christian Szegedy, a friend of mine, when he and Tony Wu, who later became xAI’s co-founder, that was like 2016, 2019. And then that’s Google, right? Mothership.

And then 2021 was DeepMind. Also bottoms up effort. One intern decides to apply this geometry. It’s just personal hobby. So, when this person was a high schooler, he has no interest in the other Math Olympiad problems compared to his interest for geometry. So he just said, “I just want to do geometry.” And that’s such a smart idea. I sometimes think of it, because when I was a kid, I could never solve the geometry problem on the Math Olympiad in the geometric way. There’s something about my brain, and I think it’s still like now, I don’t have any sense of direction, I can’t envision spatial objects. I can’t do topology in the visual way. When I was studying geometry, I couldn’t picture those curves and surfaces. It’s almost like my brain is inhibited in that way. On the other hand, I overcompensate by being the strongest algebraic brute force person you can find. I mean, when I was a kid, I would eat all the inequality problems.

So, from a strategic point of view, I just figured out that I can express all these lines, points, intersection points, circles, triangles, in complex coordinates, and this is a brute force method. So then I just throw away that picture. I just look at work with this complex coordinates, and I do it the algebraic way. And it would took me three times the time.

But it tells you something philosophical. That’s not quite exactly the AI corresponding thing, but it’s like, if you have a geometric figure, you can try to express it in a symbolic language, for both that human kid, and in this case, this intern of Google DeepMind, who then that was a wildly successful project. They expressed the geometry problem in the vector language, not Lean, even more smaller domain specific vector language. So, all these pictures suddenly become computer code.

I think they later gave up on it. I don’t really quite know. After Alphaproof, it was such a success. And so somehow the informal math team, Gemini had the mainstream voice and the non-lean approach, which is Gemini post-training, took over, they started develop evolve, that is discovery, construction not proving. So we didn’t quite hear much from them.

Matt: Well, I think they got pulled a little bit given some of the new model companies that were making such good progress, and they had to focus for a season on the here and now, which of course creates the opportunity for Axiom Math. So, where do we go from here? Where does the field go, and then where does Axiom Math go?

Carina: I still think sometimes it’s just so interesting to think about why things stopped at each of these players.

Matt: Yeah, that’s an interesting question.

Carina: And we just covered the Google story. The OpenAI story was just like, when Elon was there, with Stan Polu and other folks, they had so many papers. That was frequent within a couple months. Meaning F2F was the high school benchmark they established. That was later sort of really trained on, contaminated by everyone, but it was one of the initial benchmarks pointing at least a direction to go. I think there was GPT-F. They tried to have early version GPT to maybe stand for mathematical functions, I think. And then they just didn’t… I mean, Elon left. They just wasn’t supported.

Facebook was doing it. Facebook was doing it. The Llama team, they spun out to be Mistral, and that was the end of story. I mean, there was Guillaume Lample, there was Fabian. You still see some of the bottoms up Meta researchers, but not really feeling like AI from X was supported. So they came to Axiom. Yeah, great opportunity.

But you just see kind of how short-lived each of these excellent projects are in the big places.

Matt: I think that probably in some of these places the research had to be in service of what ultimately became the commercial pressures. And that perhaps, a few years back, the commercial opportunity to have provers that are the complementary tension to systems that are predictors, was not quite ready for prime time. But now it is.

Carina: But why did they start it? So that’s kind of my…

Matt: Well, I think they knew that it would come.

Carina: I see.

Matt: And it was maybe a timing question.

Carina: Oh, they did it too early?

Matt: They did it too early. That’s my theory. But now we are at this time, where people realize, especially as we go from things like chats that, okay, if it gets a little hallucinated, okay, whatever, you can go find another way to check it, check your work. But in the world of all kinds of commercial use cases, I mean, think about hardware verification, think about software verification, think about a transactional verification, or a security use case, you want some certainty that’s going to counterweight the predictability of the, as we think of them today, LLMs.

Carina: It’s shocking to see how many people are using AXLE to verify things in blockchain.

Matt: Interesting.

Carina: We’re just seeing, wow, all these interesting use cases coming in.

Matt: And a blockchain supposedly is provable. I mean, in and of itself, the system is supposed to be provable.

Carina: I do think it’s interesting, because it’s like the benefit of being in a startup, it declares DNA, hopefully it’s doing well, it’s that you have actually that research security. You’re the core engineering group, you’re pushing on something that is on the critical path, and that critical path does not favor. We’re talking about years of good work, and at a sort of prime stage of this researcher’s career, and just compound from there. After a win, unlike in some other company, there’s a rifting company, you have finally reached that. So your point of too early is so funny. They have outdone themselves.

Matt: They have outdone themselves, and I think-

Carina: Should be forced to pivot.

Matt:… the researchers that I have loved getting to know the most are doing research for impact. And I think at Axiom Math, you are bringing people onto the team that are excited about research for research, sake, and impact. It’s an and equation.

Carina: Yeah.

Matt: So tell us about the Putnam competition from last December and how this came about.

Carina: That’s right. So I think early stage startup, people get excited by a common goal. And for us, that goal, that sort of mission was Putnam. Putnam exam is very hard. So I think if people choose Putnam, that generally mean, one, they have taken that exam, which we found a lot of people whose Math Olympiad trajectory stop at the high school level, and two, they’re not totally hopeless yet. I mean, so a lot of people go to Putnam and more than 50% will get a zero out of 120.

Matt: That’s discouraging, because these are very, very bright people, and to get a zero.

Carina: Yeah, partial credit. Still, more than 50% of about 4,000 participants each year, who represent their respective universities, the Putnam teams, got zero.

Matt: And so many, many people are getting zero out of 120, and there’s only been a few people that have ever gotten a perfect score.

Carina: Out of, I think, 99 years of Putnam, there are five human perfect scores. We got a perfect score, and this is an AI perfect score. And this is just quite, I think, striking, even from our perspective. I personally did not do that well. I got like 39, so not even 40, and got an award for it.

Matt: So how did you prepare for your systems to take on the Putnam test?

Carina: Surprisingly, a pretty hard part was aligning on the school, just because people were like, “Well, this is quite hard.” I mean, you probably hope to start with a high school real time tournament. But timing wise, there just weren’t any. So we started on the day of the international Math Olympia, when we had no money. So the company literally just operated that week. I mean, we had campaign chairs, foldable tables, we didn’t have a line of software written.

Matt: Perfect startup dynamic.

Carina: Yeah. So we didn’t have such a high school tournament to help climb before Putnam.

The other thing was, later we realized that this year’s Putnam, the 120 perfect score is hard for two reasons. First, top performing human of this year is 110. And the best informal LLM, sort of general off-the-shelf models is DeepSeek — 103. So this is the first time a formal AI beats an informal one in the data scar sort of disadvantage kind of starting.

Matt: I think related to this is this whole idea that the pace of innovation is… I’ve never seen it. And being in venture for over 25 years, the pace of change. You’ve been so deeply involved in this area for the last several years. How are you seeing the pace of development in applied AI and where it’s going, and the research?

Carina: So, I think generally progress happen for a couple reasons. Some are random, some are not. The sort of random reasons are like, somehow people start believing buy-in the religion. So I think if you think about how many companies are pitching continual learning after the Ilya circus for the podcast, you can call that like an Ilya surge.

Matt: Yes.

Carina: Continued learning, people have been pushing it since forever. Before the AI people got interested in it, the neuro AI people were quite interested in it as well, just because of a meta learning sort of in the cognitive science literature. Well, interestingly, I think they were studying why it was hard. I did my master at Oxford Neuroscience, and one of my master dissertations actually on continual learning, but from the angle of why it’s hard, from a theoretical perspective, obviously that’s not going to make things easier, better on the practical side. And there are some other not so random reasons, which is, people just realize that hallucination’s getting bad and you need a reasoning model. So I think around the time of 2022 or 2023, I looked into formal theorem proving. My main motivation has never been about hallucination. It just has never been, because I know that’s going to be a controlled problem. Just from a law of physics perspective, humanity’s not going to let this level of BS hallucination happen forever.

Matt: Yeah. It’s almost like the benefit of having systems that can hallucinate, is that in a sense they can be creative, but the risk of a system that hallucinates is that you can’t prove it to be true.

Carina: Back then it wasn’t even creative, but I think it’s like also you can choose to build contractually, kind of like controlling the downside. You can also build expensively. And I’m always the kind of people that’s in favor expanding the universe choice for both myself and others. I kind of jump around and academic disciplines, I do whatever I want. I don’t question why I’m suited to do it. And in that sense, what my interest in formal theorem proving has always been expensively. It’s like, what if we can understand the human brain? And what if we can understand the universe?

Matt: And as you pointed out, I mean, this field, if we’ll call it artificial intelligence, actually, it’s almost the exact 70th year anniversary of that. There was a gathering at my alma mater, Dartmouth College, in the summer of 1956.

Carina: Wow.

Matt: And these different pathways are now in some ways in contention with each other, but I really believe, and I think you believe, that they’re each pushing each other at a pace that we’ve never seen before. And that leads to lots of debates: reasoning systems, continuous learning, where do we get most of the benefit? There’s been a mindset of, well, pre-training was dead, that it was all over. It was all going to be in post-training, and then it was all going to be continuous. And now we’re kind of coming back around and saying, “Wait a minute, some of the newest models.”

Carina: I had that line in our initial seed round pitch deck.

Matt: I think you did. I think you did.

Carina: I think something like, “Scaling is dead, but long live scaling.”

Matt: Yes. I like that framing.

Carina: I think that the point was, it wasn’t scaling instead, is that we want to scale promo data. And it was fascinating.

Matt: And that was a year ago.

Carina: Yeah.

Matt: Right?

Carina: That was so bad. That’s my point. I want to put it in today.

Matt: Well, it’s a fun way to catch people’s attention.

Carina: That’s right. It’s like clickbait.

Matt: So people are curious, I’m sure, how you’re thinking about, okay, so we’re building these really interesting models, we’re succeeding at the Putnam competition. What are the commercial use cases you’re exploring? What might a verification system be able to do?

Carina: I think a very specific call to action is to call everyone who wants to reimagine the EDA space to just come and join us. That’s one. I would say that we have learned a lot in the hardware verification space thanks to a lot of subject matter experts, both in terms of potential customers and in terms of people who have been working in the field, and wanted to be done differently. I think that a lot of the SMT-based formal checking tools are widely used. They are widely used. But they cannot make it autonomous in a way where you need humans to do some handholding, adjusting the counter value to a smaller amount where it does not suffer from the complexity explosion. So we have learned a lot in the hardware verification space.

I think the dream really is to… I mean, if you think about verification cycles being three to four times design cycle in time, and three to four times number of people in headcount, then designing, then surely there’s something that needs to be improved. And then a big sort of question here is, are these existing formal checking tools good enough for the more elastic demand future? And what we have found is that Lean through Improver could verify circuits that a formal checker could not.

Matt: Wow, that’s a big learning.

Carina: Yeah. Not an entire gigantic chip yet, I want to clarify this, but it is a small local part of the circuit. That is still too large for those formal checkers. So this is some sort of interesting margin or gain that we are observing.

And another thing I would say that’s also interesting and counterintuitive is the proof of that case is hundreds of lines of link code. When most of our platinum problems are thousands of lines of link code, and in some of our research problems, tens of thousands of lines of link code.

So, I think it is possible that we could have not one Putnam and be able to serve this customer. So that is interesting, right? It is the sort of center capability overhang that we observe, and just across AI. But it doesn’t mean we shouldn’t win Putnam, because it’s frankly quite fun. And I think that there could be a lot of reasons why you want to push on the math self-improving capability side, because of software verification. So, what we have found is I think we have one of the strongest math provers in the world. And that also corroborates with we saturated the code variant benchmark, a Lean software verification benchmark. I think DeepSeek Prover was at 11, 12%. Goedel-Prover’s at 11, 12%. We are 98.93%.

Matt: That’s incredible. Yeah. And I think this is part of what you’re beginning to learn is that being able to build verification systems, prover systems, can be generalizable at some level. And not only generalizable, but then order magnitude better.

Carina: Yes. And I could argue something that’s interesting, which is, if you think about, “Hey, I just don’t have an affinity for math,” supposedly. “I just don’t care about math.” This is a very sort of extreme brainstorming exercise. And I want to train a model that can do all these things. So, I will need to collect some hardware data. I will then need to collect some software data. And all these data have commercial value attached to it, versus math is kind of free. So, if you train on math data and it transfers quite efficiently to these domains, then this is sort of generalizable, not universally, because you can’t make a good writer or poet out of it, but it’s generalizable in the best way in that you are taking a short, efficient path to tackle some of the most valuable domains where you might not have a data mode if you just pursue those domains.

Matt: I think that makes sense.

So, let’s come back a little bit then to this question of what does this mean for the definition of artificial general intelligence/artificial superintelligence? How do we put all these things we’ve been talking about in the context of math provers systems into that broader world?

Carina: So, think about the word GenAI applications. And so there’s this internet meme of, they’re like a huge building, a lot of bricks, all this stuff. And then there are only two pillars, literally just two bricks here and there. And there’s all this stuff on top of it. And these two pillars, in my view right now, the infrastructure layer, are LLM and predictive analytics. Nothing else. The whole world of genAI is built on probabilistic LLM, which we don’t understand, and predictive analytics.

So it’s just from a symmetry standpoint, Matt, just from a symmetry standpoint. If you have such a stochastic system, it makes sense for you to have some deterministic component. It’s just from a symmetric. So then we want to basically be, we think of ourselves as this sort of edition, an inevitable edition in all possible futures in the AI in front layer.

Matt: Yeah. You need determinism in a world of non-determinism.

Carina: Yeah. I’ll add something that is another… I like symmetry, I think it’s a beautiful concept. In group theory, you have symmetry everywhere. Another symmetry here is, isn’t it beautiful that you can express math in code, and you can verify code with math?

Matt: I love that.

Carina: I think that is just beautiful, because it’s like, in a way, there is a very clear path in front of us to unify the two pillars of the digital world. I kind of agree by so many experts. Eric Schmidt talk about it all day. Two pillars of the digital world, math and programming, math and programming. Programming gives you output; math gives you property. And output is meaningless if you don’t know what property is satisfied. My age is a number. That property, it is my age. I mean, in that sense, you have logic, and you have computation, and they’re unified, I think, in what we’re doing, and what we’re set out to do.

So I think that’s something that is oriented to today’s AI landscape, but there is a future. There’s a future one.

Matt:

Which maybe brings me to one last question, which is, in less than a year, you’ve already built this incredible team with complementary coding, and math, and infrastructure, and a whole set of capabilities. How have you been able to do that? And what are you learning from being a founder, and company builder, and team builder?

Carina: There are a few theories I have. One is, it could be that just formal math was so data poor. It’s like this plant is almost dead, but then you just water it, and it has beautiful flowers coming out of it. That’s one theory, right? Which means very good sample efficiency. That’s kind of the technical idea, is that if you have verified output and pairs, then grounding the natural language, have formal logic ground the natural language, you’re going to do really, really, really well.

In terms of the other sort of team building aspect, I think sometimes, and this is quite encouraging, when I was talking to folks like Francois Charton, for example, and Ken Ono, I felt like it was like there is a shared dream, and there is this shared yearning. Some people don’t state it; some people state it. That is that we all yearn for a superhuman mathematician. And this is very interesting because if you think about in the context of job replacement, say if you have a super human AI lawyer, that means some lawyer is going to be out of job, and some lawyers might be averse to that.

Math is a field where there is a common goal that people are aligned around, and that is the pursuit of knowledge. And so it’s like, if that is a common goal, there isn’t a lot of ego around who’s going to make it happen. And the feeling of being able to clone yourself, you’re a great mathematician, and extend yourself, and clone someone who has passed away decades and centuries ago, and talk to that, these are, I think, beautiful concepts that, somehow, because of technical reasons we have talked about, are coming into life. And what are we seeing every day? And this is maybe why people feel motivated to work even harder. We’re seeing every day is like, we will get an email about a math problem from, say, a professor in Germany that is in Ken’s network at 10:00 AM. By 2:00 PM, four hours later, Axiom Prover generated a whole link proof without human intervention, autonomously checked, and sent it back to that professor. Here is a proof for your unsolved math problem. And that is daily mathematical Pokémon hunting in a funny way, but there’s something deeper, and we’re still slow to react to that fact.

Matt:cYou’re learning your way through that.

Carina: We’re so to react to that fact that there is hardware and software verification. I think there is endless market that is in the category of mathematical superintelligence.

Matt: Well, that shared mission about mathematical superintelligence, the incredible team you’re building, and this ongoing curiosity to discover what the use cases are for that to impact the world, is what’s inspired us to be partners and colleagues with you.

Carina: I don’t know if the last time I saw you, I told you this news, but just their mathematical results, that after I tell people, they’re like, “I don’t really know how to technically prove it, but it sounds beautiful.” And here’s one. It’s something recently autonomously proved by Axiom Prover. So, the Ramanujan tau function is tau, and it is a function that appears frequently in Ramanujan’s lost notebook. It’s one of the fundamental functions in number theory. Axiom Prover, the AI, autonomously proved that the Ramanujan tau function misses 100% of the primes.

Matt: Wow.

Carina: And this is assuming the ABC conjecture, of course, which is something like the Ramanujan hypothesis; a lot of people, or generalized people, just agree or assume all, but this is something that we don’t know. And so that is a beautiful result. I mean, just when we shared it, people felt like that is… I like to be this math evangelist, because I just like talking about it. Relatedly, there are mathematicians like Ramanujan in this case and Maryam Mirzakhani. They left behind bodies of work in terms of Mirzakhani, there is the dynamical systems. And recently, there is a paper that we published, in which Michigan Professor Alex Wright posed the problem. And in that case, we were able to also formalize the solution. So, verify it is correct, and actually optimize the proof and correct the proof mistakes. So, comes in different forms, some autonomous proving, some auto formalization, but we are at a point where we have, I think, seven papers on archived now, where the main engine is done by Axiom Prover. I think five for proving, and two for other formalizations.

Matt: Well, I’m excited about this idea of the beauty of math, and the applied beauty of math that Axiom, you, Carina, and your team are going to help us discover in the years ahead.

Carina: Thanks, Matt.

Matt: So thanks so much for being here with me today.

Carina: It’s really great.

Matt: I really enjoyed the conversation.

Carina: Thank you.

 

Palantir Alum Explains How AI is Used and Bought by the Federal Government

 
Nick LaRovere spent years at Palantir before co-founding Pryzm with friends, including a Lockheed Martin alum. And in this episode of Founded & Funded, Nick shares his inside account of how AI is actually being adopted inside the federal government, and what it takes to sell technology into that market.

Pryzm is an AI-powered intelligence engine for government business development: it aggregates data across CRMs, email, Slack, and public procurement sources to help companies win contracts. Nick’s argument is that by the time an opportunity appears on SAM.gov, the deal is already decided. In this episode, Madrona Partner Chris Picardo and Nick cover: why there’s still no purpose-built CRM for government buyers, what Pentagon AI adoption actually looked like from the inside (workers hadn’t touched an LLM as recently as two years ago), and what forward-deployed engineering at Palantir taught Nick about building close to the customer.

If you’re selling into government, building for defense, or trying to understand how AI procurement works inside federal agencies, this is a useful map of how the terrain actually works.

Listen on Spotify, Apple, and Amazon | Watch on YouTube.


This transcript was automatically generated and edited for clarity.

Chris: Nick, thanks so much for agreeing to do this and coming out here, and it’s going to be a super fun conversation, and it’s an interesting time to be having this conversation about defense.

Nick: Indeed. There is a lot going on in the world right now. But yeah, I’m excited to be here. It’s always great to be back in Palo Alto. I was here with Palantir for a long time and am always nostalgic to get back.

Chris: Yeah. Well, let’s dive right in. I mean, as we just said, super interesting time in defense, really timely to be having a conversation about what’s going on, especially in the digital world. But why don’t you start by just giving us an overview of the state of defense and public sector innovation? You see both sides of the market of Pryzm. What’s really going on right now, and what’s changed if you look back, maybe even since you started the company?

Nick: One of the things I remember was when I was at Palantir, and I was walking into work, and people were protesting because we worked with the military, because we were working in defense. And at the time, it seemed to be a very partisan issue; whether you engage with the military, you support defense. And I think I’ve seen over the last several years, thankfully, a very bipartisan movement to get around and to support defense, just like it’s a critical national security issue to support the warfighter and support those missions. I mean, fundamentally, that’s what keeps us safe at home and safe abroad. And so I think when you look at the state of defense and defense tech, that’s not by happenstance. How we got there, I think, has been one; unfortunately, there is danger in the world. There are a number of conflicts, which I am not qualified to speak about, but there are a number of conflicts going on in the world.

And so we need to be able to meet those. We need to meet those as a country. And so a couple of things have really supercharged this defense tech movement, even beyond some of that conflict. One is, I think, there’s a focus on commercial first technology, which has really erupted. I think for too long, the department has not always seen the best innovations of Silicon Valley finding their way into the department’s hands. And sometimes that means you’re building these really custom exquisite systems versus, hey, there might be these low-cost alternatives that are already mass-produced in the commercial market. I think if you look at what has happened in the drone market, that’s a super interesting comparable.

But hey, can we bring those in at a better unit cost, bring those in and ramp those up? I think in parallel to that commercial movement, there’s been a heavy emphasis on acquisition reform. That’s something that we at Pryzm are really excited about and trying to help lead the charge on. To not just go commercial first, but to use AI, to use first-class software systems, to train and trust our procurement officials and professionals to go out, take risk and bring in the best technologies.

For a long time, a lot of defense procurement was, I think, more focused on the process than the outcome. It was, “Hey, can we be compliant?” Versus, “Hey, can we actually achieve the mission and achieve the outcome?” And so I think across these realms of commercial first, better acquisition processes, and whatnot, along with just a surging defense budget, for the very first time, a $1.5 trillion defense budget. Now, there’s a little bit of color that’s within that of what it goes to and how much is base budget versus reconciliation, but you see all of those things coming to a confluence.

And that’s pretty interesting, especially when you see the Palantirs, the Andurils of the world starting to grow. That’s when I think venture really starts to say, “Hey, there’s something going on here, and we should be supporting this, and there’s an opportunity here.” So I could talk about this for a long time, but clearly it is an interesting time.

Chris: One of the things we were just joking about was on the venture side, we care a lot about market size and in your world, you could just say, “Hey, I mean, the TAM has expanded from $800 billion to $1.5 trillion just in the last five years,” through both administrations, too. So, talking about the new budget and expanded TAMs and all these things, one of the things I’m hearing a lot about is flexibility in the defense budget now. What does that mean? Talk us through how we should think about that and the opportunities there.

Nick: So to boil it down first upfront, one thing that a lot of people don’t realize about federal contracting versus your traditional enterprise sale or enterprise SaaS is the idea of colors of money. Have you ever heard of this before?

Chris: I have not, no.

Nick: It’s pretty interesting. So, the colors of money basically say that, let’s say you’ve got $100, that’s your total budget, you actually can’t just go spend, as a government official, you can’t just go spend that $100 however you want. There’s particular compliance attached to different segments of that, so you can spend $20 on mashed potatoes, and you could spend $40 on entertainment.

Chris: Sure.

Nick: I don’t know why I said mashed potatoes.

Chris: Why not?

Nick: So you can align it to different things. And so, within the government, that’s things like RDT&E, research and development, procurement, operations and maintenance. But there are basically different things that the money is attached to.

And why that matters first, from a congressional and from a taxpayer perspective, is that it’s about oversight. It’s that, “Hey, we want to make sure that the money is being spent in the right way,” right? We’re not having corruption or fraud. It’s going to the things that the taxpayer actually cares about. So that’s the idea behind color of money is trying to set the priorities and make sure it’s spent on the right things. But the downside of that is that it also creates a lot of red tape, right? It creates a lot of overhead into how you do business.

And so one thing that the department is now moving into and exploring, obviously, this is a partnership, right? The department may want this, but Congress needs to approve it as well, is the idea of colorless money or more flexible pots of money. And so one example of that is what’s being done now with what’s called PAEs, but they’re basically now saying that for a given capability area, like unmanned systems, you basically have one person in charge of that set of budget, and they’ve got a ton of flexibility with how they can spend it. So rather than needing to stick to spend it on this, spend it on that, they can just make those decisions.

Now, the downside of that is, well, is there oversight? Is it going to be spent in the wrong way? But the upside is, well, maybe if a vendor’s not performing or if we see an emergent need evolving in a conflict overseas, we can reallocate money and put it towards where it needs to go, like any other business would do or any other organization would do.

So I think it’s pretty exciting to see this movement towards that colorless money, towards that flexible money. I think, in turn, the piece that we’ve heard from Congress that people would want in response would be transparency. So it’s, “Well, if you want to flexibly decide where you spend it, you at least need to tell us what you’re spending it on.” And I think that’s a give and take that’s still emerging, but I’m optimistic about it.

Chris: And that’s pretty well situated for the Pryzm approach and infrastructure.

Nick: We help with so much of that transparency. Yeah.

Chris: And one thing I’m curious about before we dive deeper into Pryzm, too, but just because you have such a unique perspective, broadly, defense has been dominated by the primes for a long time, and they have deep relationships, they’re super embedded, strong businesses. And we’ve just seen this wave of innovation, right? Palantir and Anduril have maybe been the tip of the spear on some of that innovation, but what else is driving that? Why now, in the last 5 to 10 years, have we seen this acceleration of just, call it Silicon Valley or just call it American technical innovation, going after this space, which was for a long time, really focused on a couple of large players.

Nick: I think one thing that I just want to call out, just it’s one of our mantras, and I think it’s important to keep in mind here is, first, we exist, we say as a company, we exist to get the best technology to the missions that matter the most. And the reason I say that is that mission statement is regardless of company. We just want to level the playing field and allow everybody to compete. And then we’re believers in the free market and fundamentally, whether it’s one of the stalwart primes or an upstart out of university, whoever can make the best thing and scale it, that’s what should get to the mission.

But I think what’s so interesting about the time right now, why now? Why is there this opportunity? I think one has been the tenor of this administration change. There has just been a very real focus to get this right and get this fixed. People have been talking about acquisition reform since the ’90s, even further back. I mean, it’s been going on for decades, but for once, we’re actually coming in with such a radical approach to change some of how we do acquisition, to change some of how we do procurement. So that’s one piece that is going on here.

The second is that we just don’t have time to wait anymore. There was a time when, hey, maybe we could be okay with stasis, or at least we thought we could. But the reality is, there are adversaries. I think we see that the time is now. We do need to step up and not just to be on the aggressive, but even to defend ourselves — the time is now. And so I think people in the administration are trying to make change. There are these real pressures that are coming that say we can no longer delay, “competition” is here.

And then the last piece, which obviously we’re beneficiaries of is AI. I think what AI is able to do around autonomous systems, around, for us, intelligent procurement workflows and fusing together all of this data — those three forces are all really come together to say, okay, there’s a really unique window now where we can actually make change on this.

Chris: Yeah, so it’s the convergence of there being a lot of technical innovation that has been happening, and AI has really accelerated it, and just the geopolitical realities are causing, for lack of a better term, a customer pain point, right? And a lot of demand for new solutions, innovations, and products in the space.

Nick: Yeah, definitely. And to give credit where it’s due, SpaceX, Palantir, and Anduril — they didn’t just get here. They laid a lot of the groundwork for everybody else. Palantir and I think even SpaceX had to actually sue the federal government to get their right and get looked at as a solution. And I think what that’s proven out, those folks who did that hard upfront work is like, “Okay, well, maybe there is commercial technology that we can bring in here. Maybe there are some innovative players that we can bring in here,” and it’s opened up some of the eyes of the government as well and their appetite to say, “Hey, maybe we can accept a little bit more risk here and we might get a good reward as a result.”

Chris: Yeah, that makes a ton of sense. I think that is … Credit where credit is due, right? Spending a little bit of time on Pryzm, which is your company, which is super fun to talk about. I mean, you guys are at first, maybe just explain a little bit what Pryzm is, and then we can dive into what’s so interesting about where you are in the market.

Nick: Yeah, definitely. So we built Pryzm as the AI-powered OS for all things government BD. If you’re a vendor and you’re trying to sell into the government, we will supercharge your workflow. You will win contracts, you will see your revenue grow, and your technology will get to where it needs to go.

Now, what does that actually mean? How did we even get there? So as I mentioned, I was an engineer in industry for a long time at Palantir. My co-founder was coming from Lockheed Martin, and we saw from both the new prime and the stalwart prime how hard it was and how difficult it was to actually work with and sell into the government. And so when we spun off Pryzm, we’d seen what good BD looked like, but we knew that that was not something that was accessible to everybody, and so we built out tooling to really help supercharge some of that workflow.

And if I’m to boil it down to a few things, it’s one: data pipeline. So if you’re trying to understand market intel, we pull from all these sources, and we’ll tell you who has budget, who that person is, how much budget they have, what opportunities they have, and what companies they are already working with. I mean, we’ll give you all of that intel.

But we don’t just stop there. One of the things that we’ve seen repeatedly is that by the time an opportunity is publicly posted, it’s already too late. So by the time you see a solicitation on sam.gov, you’ve already missed the boat. And so what we’ve built out is really tooling for … It’s basically a CRM stack. This intelligent CRM stack basically emancipates all of your data across your organization, your Outlook, your Gmail, your Slack, and your Salesforce. I mean, we’ll pull all of that together, and then we’ll use that to help you actually go out and shape and get ahead of these deals.

Chris: It’s an intelligence engine.

Nick: It’s an intelligence engine. And then the public data that we have is really just an enrichment layer that just supercharges it even further. And that’s a bit of a contrarian approach to be frank. I think if you look at a lot of the existing stack that people had been having to suffer through, it was these either highly generic tools or just these very low lift AI proposal writers, but to really say, “Hey, we’re actually going to focus on the harder problem, which is the people and who they are and how you shape them and your conversations, your workflows with them.” That non-linear synthesis type workflow, that’s a harder problem, but that really is actually how you win a deal.

Chris: Yeah, it fits into this idea that what you really want, especially for a lot of these AI-enabled workflows, is as much context as possible, right?

Nick: 100%.

Chris: Yeah, and so you have context on the government side, you have context on the customer side, and that allows you to build this intelligence and the matching between customers and opportunities and vice versa.

Nick: 100%. Yeah. I like to think, I don’t know how you guys think of this at Madrona, but I’m sure it’s probably similar. I think humans, where we really can thrive in this new age of AI, is as the curator, almost like an editor, a museum curator. And so, I’m an engineer by trade, and so I’ve got Cursor running for me, or Codex running for me, writing code, and I’m obviously reviewing it, synthesizing it, pushing it to be as impactful as possible. Why can’t our federal BD people have that same level of tooling, that same level of stack to pull all that context together and then take action with them, really at the helm of this suite of power?

Chris: I’ve thought about it a lot in the context of discernment. You have a ton of information in front of you. Something like Pryzm does an amazing job of pulling that all together, putting that in an actionable format, but you, as the human, have the discernment to ultimately make the decision and tell it what to do.

Nick: Yeah, exactly.

Chris: How do you think about being a digital-only company in a space where most of the names you hear are hardware-only or very hardware-focused? How does that work as you think about positioning your company and you’re in this broader space that people are super excited about, but you’re the digital player?

Nick: Yeah, I enjoy it. Yeah. I enjoy that niche that we fill. One thing that we consistently hear from, we partner with a lot of venture firms and one thing that we consistently hear from venture firms is that when they’re investing in any sort of company that is going to market into government, you need someone already on the team, or you need a plan to hire someone who knows how to sell into government. It’s just such a funky sales world that-

Chris: We’ve thought exactly the same thing.

Nick: Yeah, and you can have the best hardware in the world, you can have the best stack in the world, but if you don’t know how to sell it and get into the hands, unfortunately that may be the end of the company, you may not be able to sell and then you may not have revenue and then you might have the best technology, but it just goes caput.

And so fundamentally, we saw this emerging, right? There were so many of these amazing hardware companies and software companies that are trying to sell into the government, but had firsthand seen and felt how difficult that pain was. And so for us, what’s I think needed is a software native solution, a really AI native solution to, again, help those teams to, one, if they’re new to this, to play up and if they’re already first class at this, to meet them where they are and help them sell even more and help them really go out and hunt.

And so that’s a little bit of how we think about it. I think we want to be an enabler, so that fundamentally, I see a world where eventually every company in America is part of the defense industrial base. I know the department shares the same view, but how do we pull in some of these non-traditionals? How do we pull in really every sort of company and engage them, almost like you saw in World War II. You would have all these automobile manufacturers actually stepping up and creating airplanes, and it’s like, “Well, maybe we need something similar now,” and maybe Pryzm can be a part of that with our software platform.

Chris: Yeah. Being actually personally from Seattle, right? Boeing and their role in World War II is something you grow up hearing a lot about.

Nick: Oh, yeah. Amazing. Yeah, critical.

Chris: We’ve talked about the AI on the customer side, right?

Nick: Sure.

Chris: And you might be working with customers who are much more comfortable with that or who are born AI native now or are trying to get there really quickly. How does it work on the government side? Because you have a view of that too, right? So how does the government think about using AI, embedding it in the workflows, meeting customers maybe on that, or suppliers right on that level as well? Are you seeing innovation there?

Nick: Definitely. Although I will say that it’s been a recent but dramatic change. We’ve talked to some folks who are either still in government or recently out of government, and anecdotally, I mean, these were folks sitting in the Pentagon, and even as recent as one to two years ago, they had never even touched an LLM. They just didn’t have access to it within their job, which is obviously starkly different than what you see, I think at the commercial sector. And so there has been dramatic change around that. I know XAI is even being brought in as genAI.mil being brought in to be used across the department, but really just writ large across the department. I think you’re seeing innovation in those workflows, innovation, and pulling AI into procurement and beyond.

I think if you look at the status quo of how a lot of these workflows were occurring, and we’ve seen some of this firsthand is a lot of workflows for these really talented, whether it’s program manager or contracting officer, a lot of the workflows that they were having to do were really on spreadsheets and systems that were coded maybe 20 years ago and were not really built for these modern, intelligent, automated workflows.

And now even on some of our contracts, we’re on some contracts with DIU and a few other spots within the Department of War, we’re being pulled in to actually apply some of what we’ve already done on the commercial space and pulling that in so that from one, a market intel side, they can tell very quickly what companies should we be working with. I guess a slight segue there is that a type of market research workflow today is very manual. People will go offline, they’ll go to VC’s websites, they’ll try and look at their portfolio, and be like, “Which companies do what? Who should we work with?” It’s another level to say, “Well, can we build an intelligent workflow that says, ‘Hey, these are X, Y, Z companies with such and such capabilities and they fit exactly what you’re looking for.'”

And so we’re being pulled in that regard, and I think you’re seeing that really across the department, just because the time is ripe and you really need to, I think, fix both sides of the equation. You need an industry that is able to sell and engage, and then the government, having first-class AI-enabled, accelerated procurement workflows.

Chris: Yeah, I love that you’re able to actually get some of your own AI tooling into those workflows for these really specific but high-value use cases, right? Finding the right companies to partner with and vendors to work with.

Nick: Yeah, 100%. And one thing that we’ve realized too is, and this is some of where our unique edge comes in. Government is a little bit funky. And so even if you’re evaluating a vendor just to go industry lingo for a minute, oftentimes you’re trying to evaluate things like TRL level, it’s things you wouldn’t typically find in a company database, but it’s technology readiness level, which is basically how mature is this technology on the maturity wave? The highest level of TRL is, “You can put this into the war fighter’s hands now.” The lower level is, “Well, this is an R&D project that we’re building out and trying to productionize.” And so there’s just all these little nuances, and that’s a place where I think having a first-class workflow that meets that person and feels fit to what they do, really matters.

Chris: Yeah. This may or may not be the right time to ask this, but I’m going to, which is that I think it goes to this digital and AI theme, which you mentioned to me as we were prepping this idea of the digital thread and the digital thread challenge. As I told you over email, I’m super excited to hear what you think about that. So explain to me what the digital thread is and how you guys are approaching that.

Nick: One thing that we’ve repeatedly seen, especially as we’ve gone to enterprises, is that oftentimes they’re operating out of anywhere from 5 to 10 different CRMs across their different business units and business lines. That typically comes together because of acquisitions or really just business units operating independently and spinning up their own stack. Now that’s just the CRM. Then you have however many different messaging instances and email servers, then you layer on compliance, whether that’s FedRAMP high or impact level 4. And so what that ultimately creates is a truly large enterprise, you’re just operating pretty independently.

And that actually, we’ve heard this verbatim from several of the largest defense contractors, that causes issues, then when you have this really important federal official, someone like a PEO or a PAE who oversees acquisitions for a particular capability, and they’re being hit by five different business units across your organization in an uncoordinated way.

So you’re seeing this behavior happen, and that leads to lost sales, money left on the table. That gets even more complicated when you then look into the CRM tooling that a particular government organization is working. And here’s a hint: there is no gov CRM, there’s no federal CRM. It’s everybody will take something that’s one of your traditional legacy enterprise CRMs. They’ll spend, and we’ve heard this verbatim, hundreds of thousands, but typically millions, to customize the thing to work for the government workflow, and then it’s immediately brittle, it breaks, you can’t update it, and it’s just not a good system.

And that’s really where Pryzm comes into play — through all of that mess, we’ve built out this actual digital thread where we can plug into some of your existing CRMs or operate wholesale as that CRM stack. We are built for GovBD. So the appropriations process, understanding the Hill, understanding the federal buyer, all of that sits natively; those workflows sit natively into what we do. So there’s no customization; you could pull it in today. We work with the rest of your stack so we can pull in all of that data, all of that technology.

And so, one, your teams, we talked about making decisions in context, as they’re going to a meeting on the Hill or they’re going to meet with this really important military buyer, they have all the context they need immediately to go in and kill that meeting. But even more than that, at the senior level, at the executive level, we then allow you to actually roll up, and for the very first time, you actually get a view into your business. One, you can get a view into your government business or we can roll up and you can get a view across, if you’re a dual use company, across all facets of your business and be like, “Okay, well, this is how much revenue we’re actually in line to make,” at a bit more of a real time basis than again, some of these legacy tools.

Chris: It really merges the idea that you started with of, “Hey, you’re building the AI OS for this space.” But also the real vertical workflow intelligence, and I feel like in lots of verticals we talk about workflows and all these things and they’re unique and they probably are, but I mean in this world, it is a very unique workflow ecosystem in how things are done and that allows you to connect all the pieces having that deep layer of context and intelligence.

Nick: A hundred percent. Yeah, just giving you that context to make an informed decision.

Chris: Yeah. So I want to ask you something totally not about Pryzm, although you could talk about Pryzm, but more about you, right? Which is, you started at Palantir or you were at Palantir for a while at a critical time of Palantir’s development. I’m curious what you learned there and especially why you think there are so many ex-Palantir founders who have been so successful, and how that impacted your own journey as a founder.

Nick: Yeah. Can I talk about Tesla and Palantir for a moment?

Chris: Sure. Yeah.

Nick: Two of the foundational experiences that I think I had and saw were, I mean, first, when I was at Tesla, this was just an internship in the summer of 2018, and this story has been told by many others beforehand, but seeing it firsthand was crazy. I was over in Fremont, and we were aiming to build 5,000 Model 3s a month. That was the goal that Elon had said. That’s the goal that we had hit that investors were expecting. And unfortunately, we were falling short on that goal, and in just my second or third week on the internship, Elon built a tent behind the factory. He basically canceled a lot of what the rest of our internships were and put us on the factory floor. And I saw him sleeping on the factory floor. I saw him doing that level of work, and I was moving boxes and installing windshields on Model 3s with oversight and approval and whatnot. But that was-

Chris: You learned a lot about building cars?

Nick: Yeah, I learned a lot about building cars. I mean, I’m a software engineer, but I learned a little bit about some of the physical world there. But it was just such a … I think that was one really formative experience, and I just want to call that out because that’s something I repeatedly tell our team is like, “Yeah, I was working hard there, but Elon was working harder. He was sleeping on the floor, and he was putting in that time.” And so that was just an unbelievable lesson that the founder needs to work the hardest, you always need to be pushing, and things may seem impossible, but if you really just eliminate distractions and focus, you can do incredible things. That’s what I learned from the Tesla side, just because there’s a bunch of amazing founders coming out of Tesla and SpaceX, there as well.

And then on the Palantir side, as we get into some of the government pieces, I think the thing that Palantir does extremely well and has led to this diaspora or mafia of Palantir founders and Anduril founders coming out of this was leaning into the customer. I think forward deployed engineering is a term that’s very en vogue now, but if you actually look at it and you unpack what that term actually means, it’s, “Hey, engineers shouldn’t just be sitting in a room.” I’ve seen it at other companies before where your engineers never leave the office, you’ve got sales out in the field, and you’re almost passing the ball over the fence, right?

Chris: Yeah.

Nick: You’re, “Okay, well, I’ve got the deal, now let’s throw it over the fence.” And then an engineer is going to have to make decisions on product and make decisions on what they’re building, but they’re going to make it without the proper context to really make those decisions. And so the thought that I think Palantir really had that was pretty innovative was, how can we get as little latency as possible between the customer and the engineer who’s building it? And so I really saw that firsthand at Palantir.

I mean, an example was, I think I can tell this publicly, but over COVID, over the holidays, most people were on holiday break, but an opportunity came up to go and surge on some work with the CDC and responding to COVID. And so I ended up leaving my holiday, leaving my typical 9-to-5 job at Palantir, and kind of surged on that for a few weeks to a few months, and was daily on calls with folks from the CDC as they were real-time trying to work through problems.

And that’s the type of thing where I’m like, “Yes, there’s friction when you first start that,” because it’s like, “Well, we have to build the thing and we have to understand their need, and we have to move quickly here.” But in turn, you work through that friction, and you actually build something that solves their problem and is fit for their problem. And so Palantir’s done that tremendously well, that level of, again, forward deployed engineering, really customer listening and not being afraid to do the hard thing. Again, work hard, talk to the customer, and ship hard. And I think if you do that, I think those happen to be skills that then translate well into being a founder as well.

Chris: Yeah. You can’t really solve the customer problem unless you actually know what it is, which requires knowing what the customer is actually doing.

Nick: Yeah, 100%. And so why try and do that in isolation? Get out, go talk to the customer. And that’s something that we really pride ourselves on now as well. One of the things that I think is one of the greatest competitive advantages of a startup is its speed. I don’t think that’s groundbreaking, but it is your speed. And especially when you first have that first cohort of maybe three, four or five customers, can you make them raving fans? Can you go to them, hear their need? And whereas at a big company, that feature might take months to ship. I mean, you’re just serving them so you can move quick and you can ship fast, and then they become an evangelist to what you do.

Chris: I think the other interesting thing about that point around getting close to your customers is the types of customers that people were getting close with were not traditionally the ones that your software business would be saying, “Hey, that’s an easy sale.” Or, “I can go work with them.” It was a very different type of customer. It’s so interesting that that is the type where people really embedded closely … Probably were able to solve the problem better because of that.

Nick: A hundred percent. Well, especially for these really hard problems, right? We’re not just building a hacky algorithm to go viral. It’s “No, no, no.” There’s something to these really hard problems. And ultimately, if you can solve it, it’s probably hard for a lot of people; if you can solve it, because it’s hard for a lot of people, people will probably pay a lot for it. There’s probably a lot of people with it, so you have a big TAM, and because it was so hard to solve, you probably have a moat. And that just happens to also then build, you’ve got a big TAM, a high willingness to pay, you’ve got a moat. Well, that’s probably the problems you should be going after — the hard ones.

Chris: Yeah, so work on hard things sort of holds-

Nick: Exactly. Yeah, it holds. It holds. It’s way better for your business.

Chris: Yeah. So one thing I’m always curious about, and we always have a bunch of founders who are listening and watching in, and I think trying to understand your journey, especially, and we’re talking about hard problems in the space that’s historically been hard. As you’ve built up Pryzm just from the founder journey, I’m curious, what’s been harder than you thought it was, and what’s been sort of most valuable from an experience perspective as you become more of a leader, CEO, running a much bigger company.

Nick: As a founder, the first thing I want to do is just give thanks to my team and our team. We’ve just got such a high talent bar and such a high talent density, and some people who work incredibly, incredibly hard. And they could be working at so many places across the world, and they choose to work at Pryzm, I think, galvanized by the mission, the technical aptitude, what we do. And I think what has surprised me the most is, I mean, if you’re persistent at it, I think you can find luck, you can create luck, but honestly, so much of your job shifts into hiring. And I think that’s the one thing that really surprised me and that I would call out. It’s definitely not; it’s something you hear over and over again, but it’s one thing to say it, and it’s another thing to do it. Hiring and keeping the talent bar high and your team motivated and all rowing in the same boat, that is a hard thing to do. That’s an inherently hard thing to do.

And I think I do a good job of keeping the team motivated and aligned in the same direction. There’s this saying that Steve Jobs had about engineers that, “The difference between the best engineer and the second-best engineer is not an order of magnitude of 2X or even 3X, it’s like 10X or 100X.” The best engineers are that much more valuable. And I think that that really percolates through across the talent landscape. And so I think that’s the thing about the journey that I would encourage everybody, and I would always be thinking about, am I keeping the talent bar insanely high? How am I attracting and retaining the best talent? And I think a lot of that is, honestly, even doing things like this, it’s getting out to sometimes talent that might not typically know you. Maybe they’re not already in the defense circle, but they’re amazing engineers somewhere else, and they’re looking for a mission just like this, a rocket ship just like this that they want to come in.

And so a lot of that is really, actually a lot of branding and brand awareness to get your message out there, your mission out there, and resonate with people who might want to join the cause.

Chris: I can’t think of a much better place to end than on that note. And Pryzm is such an awesome company, and you’re building such an impressive platform, and I really appreciate you coming in and having the time to chat with us.

Nick: Yeah. No, thank you so much for having me on. A long way to go, a lot of work to do, and excited to keep talking about it. So thanks, Chris.

Zapier Has More AI Agents Than Employees. Here’s How That Happened

 

Zapier is doing hundreds of millions in ARR, has 800 employees, and has more AI agents than people. That ratio isn’t an accident.

Wade Foster, CEO and Co-founder of ⁨Zapier⁩, built one of the most capital-efficient software companies in history on less than $1 million in venture funding. When GPT-4 launched in March 2023, he called a company-wide “code red” (a term he’d never used before) and shut the company down for a week-long hackathon. What happened next reshaped how Zapier hires, operates, prices its product, and thinks about the future of software.

In this episode of Founded & Funded, Karan Mehandru sits down with Wade to unpack:

    1. Why ChatGPT didn’t trigger urgency at Zapier, but GPT-4 did — and the specific signal Wade used to make that call
    2. How Zapier went from 10% AI tool adoption to 90%+ across the company
    3. The pricing overhaul that simplified Zapier’s model around task-based usage, and why agents made seat-based pricing structurally broken
    4. Why Zapier’s head of HR became the Chief People and AI Transformation Officer, and what that reveals about who actually leads change inside organizations
    5. The “build first, run always” framework Wade uses for deploying AI agents safely inside enterprise workflows

For founders and operators navigating their own AI transformation, this is a practical, unfiltered look at what it actually takes from a CEO who’s in the middle of it.

Listen on Spotify, Apple, and Amazon | Watch on YouTube.

You can also read Karan’s tactical advice for founders here.


This transcript was automatically generated and edited for clarity.

Karan: It’s always special for me to chat with my good friend Wade here. I never miss an opportunity to try and get his pearls of wisdom, and Wade and I go back a long ways. I first met Wade when he came out of Y Combinator. It’s been just a joy to watch the company scale. I’m lucky enough to be an investor, but more lucky to be Wade’s friend. So, thank you Wade for everything. One of the things I’ve always respected about Wade above many of the things that he has is he’s always been a first principles thinker. And I think this market requires first principle thinking more than any other market that I’ve ever seen in my career.

So maybe we’ll start there, but before we go into the first question, Wade, maybe you can give a little bit of context on yourself, where Zapier is today, how big you guys are, to the extent that you can share, how many human employees you have. I guess at some point I’m going to have to ask how many agent employees you have as well, but any color and context on you and Zapier would be super helpful.

Wade: Yeah, so founded the company in 2011, so it’s been coming up on 15 years now that we’ve been working on this started as integrations. That was the key thing we were trying to do is help non-technical folks build integrations between all the SaaS products they were using, and quickly realized that as valuable as integrations were. Workflow was where a lot of the value was really lying is, if you could build these deterministic workflows, you could actually automate real work. And so, in a lot of ways the zaps of the time were kind of like these mini agents. They were this deterministic workflow that would do work for you. Now, they don’t have all the power that you know an AI agent has these days, but at the time, that was the best way to do automation. Fast forward to today, the business is doing well, several hundred million in ARR. It’s got 800 employees working here, fully distributed, far more agents than there are people. I think that’s going to continue to just grow and grow and grow. But we’re still hiring people too.

I do think that there’s still, like humans still are valuable too. I think they’re both pretty uniquely useful. I think that kind of catches folks up to where we’re at today. We’ve had to do a lot of work to retool the company for the AI era. Both from a product perspective, but also an operational perspective. And you know, it’s still very much a work in progress.

Karan: One of the things you didn’t mention, which I think some people know, but some people don’t, is that not only are you in hundreds of millions of revenue, which makes you special. I think the other part that makes Zapier special is that you’ve done that with less than $1 million of venture capital ever raised in the history of the company. I think that puts you in like very rarefied air. So congrats on all that.

So let’s go back to like when ChatGPT came out, and I remember you and I had a conversation around the time and you knew immediately that this was something that was worth paying attention to That was October 2022 That was Walk us through that moment when it came out, what that represented for Zapier, how you felt, what your first few actions were, and then how did that sort of permeate across So walk us back there.

Wade: Sure. We had been dabbling with AI a little bit before the ChatGPT launch. My two co-founders, Brian and Mike, had moved into IC roles in the company and they were like, “Hey, this AI thing looks like it’s going to be important enough that we should go figure out what this means ourselves.”

I was like, okay I’ll go manage the entire company. You always go figure out what the next part of the company needs to become.

you know, that was like two months before ChatGPT. And one of the things they built was a a text bot. And so it was a phone number that you could text and ask questions. And, you know, I would ask it things like, “Hey, what do you think I should get my wife for her birthday?” Or like, “i’ve got this weird toe pain what’s going on with that?” And it would answer. And if this sounds familiar, it’s because well that’s what happened two months later is ChatGPT launches. And that’s the kind of things you could do with ChatGPT. W hen ChatGPT launched, it was impressive, but because we were already experiencing the, like the texting thing, it felt like, okay, this is cool. And we didn’t actually, at the ChatGPT launch, respond with much ferocity. Instead, we were very just kinda like suggestive to the company. So, you know, post in general in Slack and say, here’s a cool product. I like cool products, you like cool products, maybe you like this too. And that was very much the vibe at the time. We encouraged a lot of product managers and engineers to be thinking about how you should be injecting these capabilities into our product.

And, we launched a handful of things over the coming months. Probably the most interesting was you could insert a ChatGPT step inside of a Zap. But again, that was all fairly organic by watching our customers respond to them, things like that. The real tide shift for us was the GPT-4 launch, and that happened about six months later in March of 2023. And the reason that was a big tide shift was we saw that the improvements over 3.5 and 4 were so large. The cost decrease was so large and then it only took six months. And so we started to just say like, well play that forward the next six months. What will we see then? And if it’s even a fraction of the improvements that we saw in these six months, that is massively disruptive to the entire ecosystem.

And so we called the code red right then, and we said, Hey, we need to rethink our entire product roadmap. We need to rethink how we operate the company. And, at the time, we’d never even used this like code red language internally. So people were like, okay, Wade, you’ve got our attention. What the heck does this even mean? What do we do about a code red? And truth be told, I was like, I don’t know, it’s code red, let’s figure it out. And we did a bunch of things at the time, but the thing that was most impactful was we stopped the company for an entire week and called a hackathon and said, “Hey, everybody’s going to go build with this. Not just engineers, but marketing and sales and support and finance. We want you to just go play around with the tools.” And so we worked on making sure we got procurement lined up so people had access to ChatGPT. We made sure that there was a couple Loom demo videos of here’s how to call the API, here’s the types of things you can do, here’s how it works. Like a normal API, here’s how it’s like maybe a little different from an API. We made sure we had like our, data privacy and all our compliance I’s dotted and t’s crossed. So no one was doing anything like willy-nilly. But we mostly just turned people loose.

And what the result of that was pre that week we had roughly 10% of the employee base was using AI with some frequency. After that week it was over 50%. We had massively changed the companies understanding of what is happening in this moment? And then from that point on, we just rinse, wash and repeated. Every four to six months we do another hackathon just to keep people’s mental models up to date on this stuff. And, we still do that to this day. As things advance, as things improve, we’re constantly trying to get the team to be on the cutting edge, both in like how they use the tools, but also what does that mean in terms of how we build our own products.

Karan: You already had hundreds of employees at the time. You knew this was significant and it needed to be addressed. It could be a challenge, could be an opportunity. Don’t know exactly what, and you were throwing a bunch of stuff at the wall. I remember a statement that you made to me back when we had our regular catch up and you said, “yeah, a lot of the people inside the company feel really uncomfortable because they feel like we’re throwing a whole bunch of stuff at the wall without really having a clear strategy.” And I remember you said to the team, which I still remember. Is that, “I know this feels uncomfortable, but trust me the alternative is worse.” You know, it’s a lot of times when you’re trying to make this big change inside a company, there’s, it’s not just the making the change. It’s managing the psychology and the emotions and all of that stuff and getting past that. So maybe just talk to us a little bit about this because some of our companies are at scale, and they are seeing this sort of threat or opportunity, whatever you want to call it. And it does require a step function change, not an incremental or linear change. So how did you get the team corralled with the sense of urgency that you wanted and not have people think we’re chasing our tail, but in fact trying to find the right way forward?

Wade: It took some time. I’d say know, maybe six months to a year of just beating the message into folks’ heads But also like a lot of show and tell, a lot of hackathons, a lot of just like builder opportunities. And some folks didn’t make it like some folks left the company of their own choice. Some people we asked to leave the company. But I think the mindset that I had just locked into was basically what you repeated is like, Hey, this might be uncomfortable, but the alternative is worse. And the way I saw it for these folks is this is going to be the defining technology probably of my career and probably of everyone in the company’s career. And if you are in a position and in a company that is going to help you and encourage you to yield these tools for maximum benefit, there is no better thing for your career than that. I get that you might be comfortable in the way you’re working. I get that may be scary because you see all these headlines about how AI’s taking jobs or how, you know, it’s going to, kill companies or all these sorts of things. And even if you believe that to be true, the best thing for you can do, if you’re going to remain in tech, and this is going to be your career path, the best thing you can do is learn how to use these things.

And so I am going to make sure that Zapier is the type of embracing these tools maximally. So anytime, I would bump up against friction inside the company where, a manager, a leader would be like, “Hey, there’s a lot of change going on right now. I’m not sure that people can. I’m not sure they’re ready for this. I think they’re pretty stressed.” I’d be like, “yep, I’m sure that’s true. I guarantee they will be more stressed if we do not do this.” And so I just kept coming back to that push and saying I know it stinks. I know it stinks. I know it stinks, but trust me, it’s going to be worse if we don’t do this. You know, In some ways it’s a lot like exercising. No one wants to get up and go running every single morning. Some people do. I personally don’t. But the reality is, if you do it, it is going to be much better for you in the long haul.

Karan: It’s interesting because Eoghan, at Intercom wrote an essay. Intercom by the way, is customer service SaaS company. Was troubled. It grew really fast and then basically hit a wall. Then ChatGPT came out and customer service became the first application that was outcome driven versus user seat driven. And he had to reinvent. He came back being the CEO. He reinvented the company, launched this product called Fin. So, in some ways it was a existential moment for the company. And one of the things he says in his essay is, “I was willing to destroy the past to build the future.”

And I wonder if that, the way Zapier kind of reoriented the whole company around AI was a little bit of where you just had the mentality of: I’m willing to burn the boats to make sure that we are relevant in this new world that we’re going into. ’cause it’s that significant, or was it shy of that?

Wade: You, I think it was. I think one thing Eoghan calls out in that post he shared is that, he says “partly this was easier for me because we had no choice.” And he actually shows his growth rate. Like you can just see what is the company.

just see what was going

Karan: went down and then it went back up.

Wade: And I do think that in some ways that’s probably a gift to him where it’s like, well look what choice do we have, we’re a dying company anyway? And so it probably made it easier. Think the thing that is maybe scarier is there’s a lot of companies out there where your numbers look okay. They’re not great. You’re not, they’re philanthropic going from $1 to $14 billion in a year or whatever. But like you’re doing okay and it’s like, yeah, you wish you were growing. Okay, maybe I missed the quarter, but I still got to 90% of my goal. I think that’s a really tough position to be in because you can very easily convince yourself that, Hey, we’re actually okay. We don’t have to have this sense of urgency. We don’t maybe need to make that tough call. And, I think for those people in those positions, you have to have more conviction. You have to push harder. You’re going to have an employee base that is going to have more pushback. Because the reality is when they say hey “it looks okay to me.” They might not be wrong. It’s like it’s harder to convince that, where Eoghan can literally just look at the scoreboard and go,

Karan: yeah.

Wade: It’s not that it’s okay.

Karan: Status quo is not

working. Yeah.

Wade: you know, I do think for companies in that like middle ground, like right now is a particularly important time for CEOs to show leadership and say, Hey, let’s get ahead of this and let’s really make sure that we will be able to meet the moment.

Karan: All right. So that’s great. And so you went

through this moment, you realized you did start doing these hackathons. Some people were part of the solution. Some people were part of the problem and the people that were part of the problem exited the company either by your behest or by them choosing to leave. How has that evolved into your hiring practices now? And as you think about now, you’ve on the other side of it AI is such a big component of it. I think you’ve put a post out there that you have more AI agents than humans today. And I know Brandon, who’s your Chief People Officer, is now the Chief People and AI Transformation Officer. So talk to us about how all of that led to the changes that you made operationally, tactically, and related to hiring and recruiting.

Wade: You know, I think there’s a handful of things we’ve been doing. First. We were really focused on initially raising the floor, let a thousand flowers bloom. We wanted everyone in the company to be working with and using these tools. And we’re somewhat agnostic to how. We buy technology from all the labs. We buy all the applications that are doing these things. Like we’re very promiscuous with our tool usage. Because if someone finds something that’s interesting and enables them to do something, we want to empower that productivity. We want to like cross pollinate those learnings. And we also want to be inspired by like these other products and go oh, they figured out a novel like interaction pattern or way to use the data in a unique way. So like, how do we bring that into our own technology? So we’re very, we are very promiscuous in terms of pulling stuff in. And that I think cross pollinated just like a lot of ideas going on inside the company. Now, one of the limiting factors there is that approach doesn’t necessarily raise the ceiling. You get a lot of stuff, but you get a lot of things that one person can do on their own. Hey, I can go make this happen. And so one of the next bottlenecks we observed was we have these, like bigger bets we want to go take, but it needs like coordination. An example of this, like in engineering might be, okay we’re generating a lot of code right now, which is great, but now the bottleneck is code review. And so we have to go reinvent our entire like code review process. So that an agent is actually doing all the code review. And then how do we get it to where our time to ship is like happening in minutes like for code review to be improved versus like days. Which was what was happening before. And that requires you to rebuild the entire system from the ground up. And it’s not just one engineer that can show up and generate more code. It’s like your whole engineering org has to say, we’re going to go do this a little bit differently. And you can rinse, wash, and repeat those types of problems across various functions in the company.

And so we felt like, okay, now we need somebody who is very senior who can help us look across all these functions and drive a lot of the change management required. And it turned out, in our case at least, our HR team, funny enough, had been the function that was doing the best at this. They were racing ahead and doing a lot of these things. Brandon is really good at change management and so I just said, “Hey Brandon, can you help all of these other functions figure out how to take a swing at some of these more ambitious projects?” We weren’t struggling with the technology side of the house. We were just struggling with, Hey let’s be bold and let’s try some of these harder things.

And so that’s been the progression for us is going from let’s get everybody just using the tools bottoms up adoption to now let’s systematically find the big levers in the company that we can generate massive ROI and go work on those on a sort of project by project basis.

Karan: When I said in the beginning that you’re a first principles thinker, that there’s constant examples of that. And the one most recently was the one that you just mentioned, which is the fact that your chief AI transformation officer inside the company is in HR And usually companies have this preconceived notions that innovation comes from either product or some other group in the company.

But in your case like, look, some person in the company is amazing at change management. They’ve embraced it more than anybody else and they’re driving their group to perform at a level that you want every group to perform. So that’s one of the things I really liked hearing and also just respect about you, so that’s awesome.

Wade: I think one of the interesting things about that with AI is that there’s this weird thing that I’ve observed, which is that if you are an expert at a particular domain, you often have these like preconceived notions about how something is to be done. You can find somebody who maybe doesn’t do anything in that domain but is like really excited about using AI and they will come in and they will achieve things that you’re like, why did you do it that way? I didn’t. , It wouldn’t even dawn on you to do it that way because you’re so used to doing this in a very comfortable way. And I think this is what’s exciting about this technology wave is it really is tearing all the gates down. And so now you’re just seeing this huge influx of people who are achieving things that just felt like it wasn’t in their skillset or wasn’t in the range of outcomes for them now. And I think it’s exciting to be able to see people just even inside Zapier, like seeing how many people are able to deploy code or how many people are able to contribute to this area where they had interesting ideas, but for one reason or another they just weren’t able to jump into the fray. But now it’s just so much easier with these tools.

Karan: Yeah. By the way, you mentioned code review as one of those problems that had to be re-engineered. Just outta curiosity, how much of the existing code of Zapier had to be refactored or rewritten? If you had to ballpark a number. And I asked this because this was literally a conversation I had with one of one of my other founders where we talked about that the, we’re going to have to refactor our code to be written for agents as opposed to for humans.

Wade: Yeah. I don’t know the exact percentage, but this is a very real problem. We’ve got a legacy monolith that’s been around since 2011. And it’s like that is the area where there’s the most challenge for, ironically it’s hard for a human, but it’s also hard for an agent to get in there and work on those things. Whereas like greenfield stuff, very easy for an agent to jump in and go, Hey, i’m going to go build you a thing. And so yeah, there’s a fair amount of work in those areas to make sure the code base is easily able to be worked in by an agent.

Karan: Alright, so let’s move to, one of the other trends that’s happening, and it’s not necessarily a trend, it’s the reality, which is, not only is the code and the product need to be reinvented, but so does the business model for a lot of these companies. And so we’re sort of moving from workflow that was sold to departments on a user-based pricing model to now outcomes that is delivered by agents with consumption or outcome-based pricing. I know pricing at Zapier has been a it’s the most interesting story because I remember your first pricing matrix, which was the Fibonacci series. And I’m curious to hear, did you try to apply the Fibonacci series again to your outcome-based pricing?

Or like,

Wade: again.

No.

Karan: but, okay, walk us, what, how did you figure out the whole business model and has that changed and evolved since you’ve incorporated AI or have you stuck to user based pricing at this

Wade: We’ve almost always had a some form of like usage-based mechanism inside of Zapier. And so early on we, we metered by counts of zaps and counts of tasks. But I think the thing that we had to go refactor was — we did a big price change in 2024 — was the way we had rolled out our usage-based pricing was very confusing. It was confusing to humans. And if it’s confusing to humans, it’s probably going to be confusing to an agent. But in particular, the problem was that some plans, you could have usage-based pricing. Some you had to move between subscription tiers. It was somewhat arbitrary, like why some plans worked the ways that they did. And as a result you could just see in our customer feedback that people were just feeling like Zapier was nicking and dimming them and just not treating them very right. And so in 2024, we basically went in and refactored all of this stuff, and we said, okay, every plan is going to have a subscription tier based on tasks alone. And so, you could buy a bucket of tasks and if you went over that subscription amount, you could upgrade to the next subscription rate or you could go pay as you go. And now the pay as you go rate will be slightly higher than the subscription rate because we want to incentivize people to commit to us. But you can still do it. We can provide that flexibility in case you’re just need a handful of more tasks — not ready for the full bundle to come along. And so we made that big shift. And as a result, we just saw a lot more people making use of the usage-based option. And so now that’s something we’ve just leaned into all across the entire portfolio is to say, Hey, how do we make sure that anytime we’re launching something that we can boil it back to this usage-based approach versus continuing to bolt on like a feature here that costs an add-on or a new product there that costs another add-on. And it just continues to confuse the whole thing.

I don’t wanna put Zapier up as like on a pedestal say we’ve done this part particularly well, but I do think that when you think about how software will be consumed in the future. It’s very likely that you’re just going to see more and more software is used by an agent and not by a human. And if that is the future, it also seems pretty likely that the agent is going to choose what software to use. It is going to choose what software to buy. And now how it buys it. Maybe they’re just buying by like spending tokens. Right? Now we defacto give these agents a budget by allowing them to spend tokens even though we’re not giving them like a physical wallet. And yeah, I do think we’re going to increasingly be in this world where an agent is going to sort of just go out and figure out, okay, how is, what is the best way for me to go solve this problem. And they’re going to have usage-based pricing makes sense for that

Now. I guess it’s possible you could see seat based pricing remain for like certain products, but increasingly, that’s going to be a problematic approach, I think. And when you hear a lot of the, there’s been a lot of talk about the SaaSpocolypse think there’s a lot of reasons for that. But the one that resonates the most with me is how much AI eliminates the need for seats. And so the SaaS is still useful. It’s just that you just don’t need the same number of seats that you needed in the past, that’s where those business models get disrupted.

Karan: I actually just wrote an article on LinkedIn about my own view on this stuff with the SaaSpocolypse and user based pricing. And I think you’re right in that before a lot of SaaS was selling workflow to humans that were naturally delineated because of the functional roles they were doing in the functions that they belong to. Agents don’t respect those boundaries like we did in a human world. And so I think this is where the whole importance of the whole context graph comes into play for the future of AI productivity software, because AI can go across boundaries. You don’t have to have a product for sales and someone for marketing, one for customer success because the humans aren’t, you know, it’s like the agents can go across those boundaries. And you’ve always lived in the glue logic between these functional roles. Because Zapier was always about connecting these roles. And so I’m curious to hear if this conversation around these context graphs that power AI agent outcomes that go across boundaries, is a threat to the value prop of Zapier, which was doing that anyway, but still selling to humans in those departments. Or is that an accelerant?

Wade: I think it can be both. Is the an is the honest answer. If your sort of view of Zapier is like, Hey, this is a dumb pipe that pushes data from one app to another, then, yeah, that part is increasingly getting commoditized. It’s much, much easier to do these things. If the view of the future is, hey, something has to reliably, run these workflows, govern them appropriately, make sure that it’s deployed in a way that enterprise feels comfortable, those are the areas where I think it’s a huge opportunity for Zapier because we have a long history of doing this stuff very well. And if our product roadmap can innovate fast enough to make this stuff interoperable, governed, etc., we’re better positioned than any other incumbent or startup to go tackle these things. And, I think this is. True for most people right now is there’s aspects of your business that don’t look so great when you look forward. But then there’s probably aspects of your business that are like very well positioned to take advantage of these trends. And so you have to figure out how do you take some of the things that you’ve built that are valuable and repackage them up to, ride this new wave.

Karan: Got it. Actually, that, that brings me to a question that’s asked in the chat about when you mentioned governance and guardrails. So the question is how do you think about the uX of guardrails for agents? So the more conservative enterprise customers selling into regulated verticals and things like that can also trust agents with critical work.

Wade: Yeah. The thing we’ve observed, the pattern that works the best is, when you are building a workflow, this is where it’s really great to work with an agent that’s non-deterministic, probabilistic because it has the flexibility to go back and forth with you to figure out like what is the workflow that you like. And then once you get it to complete a task, you have this like magic aha moment. You’re like, holy crap, that worked. This is increadible. Then you start to go okay, I want to do this over and over again now. At that moment in time, that’s where you want the agent to go build something that’s deterministic. You want to say, okay, we figured out how to do this thing now I want you to repeat it every the same way, every single time. You know, maybe that’s writing code, maybe that’s deploying a zap. Whatever it is, you want it to operate in a very structured way. I almost think about it as this like creation moment where there’s a lot of back and forth and you have the like, probabilistic moment and that’s like a really powerful building mechanism. And then there’s the operating workflow that’s like the repeated thing in the background and that’s where enterprises really want things to work the same way all the time.

That sort of build first run dynamic is the thing I think people are learning how to go do.

Karan: The other question here is: you recently launched functions, and AI is really good at writing code. How have you seen AI change the way that teams think about building workflows and integrations?

Wade: Yeah, I think it comes down to the thing I just described. When you think about the workflows of the future, they’re very likely going to be a coding agent that’s going out and building these things. And it is hooking into the tools that you use. At Zapier we use Zapier MCP and the SDK ton to help build these tools. And so like just this morning I was building out a hiring workflow where I had figured out how to — we do these bar raiser exercises when we’re hiring and my workflow today,

Karan: Describe the bar

raise

Wade: Basically, my role in it is I still approve every hire that goes through, and I mostly just go review in Ashby as our, our applicant tracking system. And I go look at all the scorecards. And I’m mostly auditing the process and just being like, have we done a good job of holding a high standard and a high bar or do I see us falling prey to like common patterns where, a common pattern might be everyone, there’s like soft yeses across the board. And so no one’s like sticking their neck out to say, this candidate is not good enough. Or, they’re missing certain like dimensions where maybe they’ve screened them really well for like technical competency, but they haven’t really done a good job at a value screen.

There’s missing data there. Or maybe the references are very rosy and you can tell they didn’t ask like the critical questions. And these are the things I’m usually auditing to say like have we actually done a good job to really understand is this person a fit. So one of the tools that I recently built to help me with this is I have a hiring council skill that spins up a bunch of subagents. Those subagents go independently, review the scorecards, come back with their own assessment on is this candidate going to raise the bar? And each of those subagents pretends to be a different be a different persona. So one might be like a hiring expert, one might be a domain expert, one might be a CEO, one might be a like thrifty CFO one might be a wartime operator like, you know, a bunch of these different personas to say does this candidate truly do it?

And so the way I was doing this workflow is I would go into Ashby, I’d copy and paste all the data out of it. I’d come into Cursor, I’d invoke the War Council. When I get the feedback back, I’d go over to Slack or go paste it into a Google Doc and then I would share the Google Docs. So it’s just this cumbersome, like multi-step little process. And so one of the things I was doing with the Zapier SDK was I was just like, Hey, so what I need you to do is go hit the Ashby API suck down the scorecards when I give you a link, run it through the War Council. Take the output, generate the Google Doc. Go find the Slack channel or go inspect the hiring team that’s on the panel. Make sure the doc gets shares with them, and then go find the hiring panel inside of Slack and post this to it whenever a candidate comes through. All that workflow was built in code, but each step where we’re, fetching from Ashby, where we’re creating the Google Doc, when we’re posting to Slack, all of that stuff is using my tokens through the Zapier SDK. And so that to me is where I see the world is changing, is a lot of these workflows are going to be built in code by an agent that has access to a whole bunch of tools, whether it’s like an SDK or an MCP or an API, and those tools are like well governed by the organization to make sure that the agent only has access to the things that it should have access to. That people aren’t like copying, pasting tokens into these agents left and right, willy nilly. Which you might be scared to learn happens all the time right now.

So I do think there’s a lot of power that’s coming, but there also is a lot of best practices that are going to have to be built to do this. Otherwise we’re going to see some pretty, pretty scary stuff in the news, i imagine, at some point in time.

Karan: So let’s talk a little bit about you. You touched on something that kind of spurred a thought in my mind, which is that the future doesn’t affect everybody in a very evenly distributed way and nor will AI. And it’ll be absorbed and consumed and either disrupt verticals and various different sort of velocity rates. And the same is true for companies and groups within companies as well. So you mentioned, you created this process and a way of thinking where every group inside of Zapier is embracing AI, but I’m sure not all groups embraced it equally. It wasn’t as successful in certain groups versus others. And so maybe A) just describe if it worked really well in certain groups what caused it? And if it didn’t in some cases, then what caused that? So that would be super helpful.

Wade: I would boil this down to two buckets for why it maybe worked or didn’t. And the two buckets would be one, there are some domains that aI just worked better in earlier on, and so things maybe took off faster there. This would be like your engineering teams that are working particularly in like new areas. These are places where it’s so to quickly see the power of these tools and you have been able to see it for a while. Then the second thing that I think also very much influenced the teams that did well is when they had leaders that also were very hands-on with the technology. And so there you tended to just more experimentation happening and you also saw the standards be higher. Like these managers just were able to like, hold people to a higher bar because they themselves saw what was possible when people used these tools. And those were the two traits that I saw of the teams that tended to move the quickest is they had those traits inherent to them.

Karan: It’s interesting. I’m observing it even across our own portfolio. There’s probably some resonance between your comment and the fact that a lot of the founders are coming back and running these companies because this transformation is more about product as opposed to go to market in a lot of ways. And so I wonder if, maybe it, it causes us to revisit how we think about leaders inside the organization and do they have to be the best at leveraging product within their domains and then rise them to be the leaders of those groups as opposed to great people managers. So anyway, it’s an open question, but I wonder if there’s of view there. You guys have always operated on your own rule book, and so I’m curious have you changed the way now that you think about promoting leaders and hiring leaders in those groups because of the comment that you just made?

Wade: There’s definitely like a learning curve or that we are figuring out about what the best leaders look like in the future. One thing that we’ve realized is that definitely like high agency, action oriented, invention oriented people tend to do like very well. Their value is so much more accentuated. This is where the like hired gun management tends to struggle. They’re used to coming in and saying, Hey, something has worked really well, or it has worked well enough, like a founder has figured out a thing. Now my job is to take it from, some success to like lots and lots of success. And, a lot of these companies were in a moment where the job is not to go from success to more and more success. The goal is to rebuild the thing from the ground up. The old thing doesn’t work quite the same way anymore. Now we’ve got to do it totally different.

That just flexes at a like different muscle. That’s very different from the things in the past. I can just share like a thing we were talking about yesterday when we were doing a talent review is, we were noticing that there were some leaders that we have in the organization that have historically been very strong. But we’re increasingly going huh, it feels like it stalled out a little. And the realization was there was a group of people who were very good critics and they were very comfortable speaking their mind around these things. And those folks in companies in the past have tended to be very helpful because they call BS on stuff, and so they help shut things down. They help like reorient to things that are more all that sort of stuff. But one of the failing traits of that personality is that they aren’t as inventive. They’re like what should we do? What is the thing we have to go build?

Increasingly what we’re realizing is that AI is a very good critic. All of a sudden that task, that job of being like the BS detector, is increasingly getting delegated to these like council skills and things like that, that are able to say, Hey, you don’t have clear DRIs. You don’t have this, and this. Your operational plan is weak on this dimension. You need to go fix that. Then you’re like okay, how should I go fix it. That’s where like these high agency creative get stuff done people are like increasingly very valuable. And yeah, it’s just like little things like this that you just observe by using the tools more and just like trying to figure out what’s, what is valuable, what is not.

I’m sure that like over time, even some of what I just said will change. Like I, I’m noticing, there’s this like whole thing, six months ago where it was like the agent it’s not as creative as me. And like I’m finding for more and more tasks, I just ask the agent like, what would you do? And it’s pretty dang good. And so, you know, is creativity like an edge any more? For certain things, sure. But there’s a lot of jobs where you’re just like, I just need like the normal answer here. Can you just give me the normal answer so we can get going?

Karan: Yeah. Yeah.

It may not be, it may not be creativity, but I think judgment is still as important, if not more important, given that you can create a lot of things. But at some point in I think you need judgment on top.

Wade: Yep. Yep. AI’s getting pretty good at judgment too

Karan: It’s, I don’t know. I’m counting the days until it starts investing in people in companies, and I can go sail in the sunset as well.

Wade: Yeah, it’s odd like it’s, you just have these moments the more you use it where all the sudden, you start to realize this task that I’ve done for a long time I don’t know that i’m any better at this than the tool is.

So

Karan: yeah,

Wade: I

guess I’m just going to let the tool do it now.

Karan: What do, so what do you think in that sort of world? Now we talking about the future. There’s. So much, narrative right now and, as buffet said, ” in the short term the markets are voting machines in the long term, they’re weighing machines,” and so we might be in the weighing scale part of SaaS, and we’re still in the voting side of AI, although we’re getting pretty close to the weighing side as well. What do you think people are underestimating? What do you think people are overestimating in all this narrative? There’s the doom and gloom scenario. There’s also like just the greatest time in history to create jobs and tools and software and produce outcomes and help society. I don’t want to get to motherhood and apple pie, but as you think about and read all this stuff and from your purview of building workflow software to make people more productive what do you think we’re over weighing and what do we think, what do you think we’re under weighing?

Wade: I think it’s very, it’s human nature. It’s easier for us to assess loss than it is to imagine I think societally right now, like you see the headlines filled of things that AI is And so I think it’s much easier for us to go, oh my, this is what is dissapearing. But the reality is there’s just as much that’s being created, if not far more, The demand for code and software and all that stuff is going through the roof. insert something, somehting, Jevons paradox here. And I think we’re just in this magical moment where we are massively underestimating the incredible things that are about to be built.

Yes, aI is increasingly taking more and more tasks away, but oftentimes those tasks were like not all that interesting and valuable in the first place, and they’re freeing us up to go do bigger, more ambitious things. And so I don’t know. That’s how I would answer. I I think we’re probably a little too scared about what’s dissapearing and not excited enough about what’s coming.

Karan: I totally agree with you. If you, obviously Zapier is at scale and you’re trying to change engines and change things while you’re in flight, but we have obviously a lot of companies that are very early in our portfolio just getting off the ground, sometimes don’t even have a product released yet. So if you were rebuilding Zapier from the ground up today with everything that is around you: tools, technologies. You are now 15 years older and wiser what would you do differently if you were to build Zapier in an AI native world as in an AI enhanced world?

Wade: Gosh it’s so much different, right? Building is easier than ever. So that’s the exciting part. Like code is not expensive anymore. It’s much more trivial to go build these things. I think the bottleneck that I think most startups are running into is attention …distribution. It’s like, how do you actually get anyone to notice you or care about what you’re doing. And that is just like really hard these days. You’ve got a lot of saturated channels. You’ve got channels that are decaying. You’ve got new channels that are not obvious, like how they work yet. Obviously like you can go viral for a minute, but like that attention comes and goes,

Karan: Yeah. And you guys,

and I know Zapier did so well with organic content marketing in the early years to get that attention, has that evolved as well for you all?

Wade: 100 Percent.

Yeah. It still is like a really important channel for us, but increasingly we’re paying attention to what do these agents recommend? What do you, how do you find, how do we show up in ChatGPT and in Claude and, inside the and all that sort stuff. We want to make sure that they think zapper is a good solution for the various problems that are being asked of it.

Karan: Yeah. interesting. I think one of the other things we’ve noticed is how, I wrote this in my LinkedIn article, which is the land and expand funnel in AI is inverted. Which is in SaaS world, you could land and that took a bunch of time, effort, money, resources, people, and then expansion was a little bit easier. Today, to your point, it’s easy to build product, it’s easy to sell product because a lot of people have experimental budgets, but we don’t know if those deals are going to renew. And so in some ways the expansion is the land. How do you feel about that statement? Do you feel like attention is obviously an issue. And, but I think durability is probably another issue that a lot of firms, companies are dealing with, which is you get a few logos but they don’t renew because people are trying everything.

Wade: Yeah, everyone’s trying everything. They’re very promiscuous. They’re happy to try pilots. They’re happy to switch to something else. I think another thing that’s very unique is that a lot of the enterprises that are buying this stuff, they don’t know how to use it yet. And so you do have like big heavy services and the rise of the FDE, you know, it has become more and more important because they want the outcomes, but they’re not exactly sure even how to use the software. They’re like what what exactly is this? And so there’s a lot more education you have to do to land these customers. And so it is a just much more involved upfront sale. Versus in the past.

Karan: The SS have flipped on SaaS. It’s now services as software.

Wade: Yeah. It’s just a very different world. Like when we launched, we were so focused on how do we make the product so dang easy? You don’t have to talk to a person. And now I believe you can do that with AI products. But there’s also so much that the customer doesn’t know that it feels like to get the best outcome, you do benefit a lot having having a guide show how.

Karan: yeah, that’s totally true. Are you, so speaking of outcomes, are you planning, that was a question as well. Are you planning to move to an outcome-based pricing model or are you going to stick with usage and consumption?

Wade: Yeah, we’re sticking with usage for the time being. We’re a pretty horizontal product used across a lot of different areas. We don’t have any intents to move to an outcome based model. I don’t, I’m not sure it makes sense for Zapier, but we have an open mind on pretty much everything. think that’s the, maybe the meta learning is that pretty much anything we felt was a settled issue in the past we’re coming back to and saying maybe it’s not so settled anymore.

Karan: Yeah. One thing I’ve observed, at least in my career is that and speaking as a sales and marketing hat on, I’ve, I feel like 50% of your customers might prefer an outcome-based model, but a hundred percent of them want the choice. And so it might be the most the most I can offer as far as insight, and I think it’s worked in a couple of cases with folks.

But anyhow I know we’re out of time. Thank you, Wade. It’s always a joy and a pleasure to talk to you. If I spent all my waking hours listening to you, I’d be a lot wiser than I am today. Thank you for being so gracious with your time.

Twitter’s Ex-CEO: The Web Was Built for Humans — Let’s Make it Work for AI Agents

 

What happens when AI agents — not humans — become your primary customer? That’s not a hypothetical. It’s already happening, and the founders who recognize it earliest are rebuilding their entire infrastructure stacks from scratch.

In this live episode of Founded & Funded from our IA Summit in Seattle, Madrona Venture Partner Jon Turow sits down with Parag Agrawal, former CEO of Twitter and founder of Parallel Web Systems, and Nikita Shamgunov, who led Neon through a rapid AI pivot before its acquisition by Databricks.

What they cover:

  • Why Parag is building a new search index from the ground up — and why existing ones weren’t designed for AI agents
  • How to pivot an established company in weeks, not months, when your customer base suddenly changes
  • The “pagers vs. iPhones” framework for knowing when to lean into disruption vs. protect what you have
  • Parag’s two-person hiring rubric for teams operating in deep uncertainty
  • Why Nikita added the head of product for ChatGPT to Neon’s board — and what that signaled to the market
  • The “two-way door” model for giving agents real autonomy without catastrophic downside

Whether you’re building infrastructure, running an AI-native startup, or trying to figure out where your product fits in an agent-first world, this conversation will sharpen your thinking.

Listen on Spotify, Apple, and Amazon | Watch on YouTube.

You can also read Jon’s takeaways for founders here.


This transcript was automatically generated and edited for clarity.

Jon: You guys have made major bet the company’s strategic calls, not incremental tweaks on how to change your infrastructure-level products based on what you were seeing at the application level. And maybe you could just take us through what it is that you did and why you were so convinced that you should bet the company on doing that.

Parag: Well, thanks for having me here, Jon. It’s good to be here. I’m building this company called Parallel, started about two years ago, and our product is essentially a bunch of infrastructure that agents can use to use information on the web. And I think when Jon says we bet the company on some vision of the future, I think what he’s implying is that we decided not to focus on use cases that were around when we started the company, and instead decided entirely to focus on AI as being our primary customers. And we did that because I was playing around building agents myself two years ago, and I could see all the reasons that it was impossible to do so.

And it became super obvious that things had to change, having built a past generation of infrastructure on the web where you were obsessed about imagining a human on a mobile device and how they browse or navigate or search or act, and what drives them. Every decision you make in building products or technology is about the end customer. And what you could see happening was that your end customer was about to change completely because an AI was going to sit between you and the end customer. And that is what creates the biggest conviction and opportunity.

Jon: One of the things that you’ve had to do is you’ve had to build a search index from scratch, which not a lot of companies have done until now. And how did you connect that requirement all the way up to these use cases that are coming that are so early?

Parag: People have built search indexes. They’ve been doing so for 30 years. There’s two or three really good big ones out there that have been 20, 30 years in the making. I think when the customer changes, everything changes. So I think before starting the company, we went into a multi-month design prototype, idea maze, reimagine the future. Once you go do that, you realize that how and what we crawl must change because a bunch of the… Our index is a complement to models.

Now models, as much as they feel like humans when we use them, they aren’t very human-like in how they work, how they behave, and what’s good for them. The data, they already have a lot of data in their parametric memory. So the index must be a complement to the model. What you rank is different, so you now need to own the index and the ranking function. We believe over time, the agents will move from a pull-based system, as we live in today, wherein people pull information from the web for the most part. Agents will be triggered by changes in the world as they manifest in the web. And once you decide that everything’s going to look so different, you kind of have to go own all of the parts of the stack where what we’ve built to date is no longer right.

Jon: Let’s go to Nikita now, because I think this kind of dovetails. Here you are with an established business that you were already running, and a moment where use cases and customers and history told you that something had to change. And in a matter of months and even weeks, you pivoted the company. Can you talk about that?

Nikita: It was about a year ago when Replit agent launched. And every time I’m going to say agent now, I’m going to be referring to coding agents, agents that generate code. So think about Replit, Lovable, Bolt, Cursor, Windsurf, Claude Code. So when I say agents, it’s one of those. Right before that, we were telling ourselves what we were building at Postgres tier for the internet. Make Postgres serverless, make it super easy to use. And the user, speaking of when the user changes, the user is a human developer. So we divided the world into infrastructure and dev platforms. We were saying dev platforms are different from infrastructure platforms for these reasons: developer experience, taste, workflow. And then, we were building that out. We were on that trajectory, and things were working. Every day, more and more people were signing up into the platform. We were very happy because we were a bunch of nerds, twisting knobs on the backend, and then the world would respond to us with more usage and better retention.

And one of those strategies was to go and embed ourselves into other developer platforms that are broader than the database platform. So we went after Vercel, Retool, Replit, were successful with this and they would like wrap us because they wanted to offer Postgres on their platforms and we’re like, “We’re easy. If you want to wrap us, wrap us.” And so we got wrapped into Replit and we were expecting a lot of usage and there wasn’t a lot of usage. So we were like, “Ah, Replit.” And then in September last year, they launched Replit agent and Replit agents, for those who don’t know, allows you to prompt and then generate an app based on a prompt. And then we saw that our usage basically skyrocketed on that particular channel to a point that there were more databases spun up by Replit than the rest of the world.

Now, they were very different in consumption because turns out that Replit generates lots of apps and a lot of the apps is throw away. That’s where serverless is useful. And then we realized that the retention of Replit user on the Replit platform is high because the user gets a lot of value. Retention of an individual app super low because people create apps so cheap and then they throw them away. And I remember being on a Zoom call with a bunch and then we have this guy, Arjun on the team, he’s like jumping, he’s jumping on the Zoom. He’s like, “We are having a moment. We’re having a moment.” Arjun runs go-to market. But now what do you do? You’ve built a product for humans and then suddenly you got to take off with this new user. And when the user changes, everything changes.

So now when you take a pause and you think about it, you’re like, okay, somebody, if it’s not you, then somebody will build a dedicated product for the new user. Someone like Parag is probably already working on this, and then you will be left in the dust. I spent some time at Khosla Ventures and I saw some examples when a disruption happens, you either lean in into this or you don’t. And when you don’t, suddenly you’re selling pagers and the whole world moved to iPhones. I think that’s the analogy, but now what do you do? And what we did was A, we said, “Is Postgres here to stay or Postgres is going to go?” Because of that transition. And luckily, I think Postgres is here to stay, lucky for us. Because if it wasn’t the case, then we wouldn’t have a company.

A visceral example in my head was a company called Navisphere. When Kubernetes came in, they had a pretty big dilemma. Are they going to keep betting on Mesas or they need to lean into Kubernetes? Mesos was in the name. Thank God we didn’t have that in the name, a company called Neon, so Neon can be anything. But so again, what do you do? So first of all, you say the user is the new user, which is tricky because you have traffic from Replit, but you don’t have traffic for anything else. But the challenge you’re also having is an internal and external. How does the world perceive your company? So is your company this modern company that leans into disruption or they perceive your company as an older company. And so this whole dilemma of infrastructure versus dev platform changed. Are you now like we successfully doing a dev platform, but are we building an AI platform, AI infrastructure platform or not?

So then what you do is like you look at all your peers. I think Eric is somewhere either in the room or around for modal. It’s like, oh, they’re building an AI platform, AI infrastructure platform. Guillermo from Vercel, who just now raised a 10 billion, is an angel in the company. And I’m calling Guillermo and he’s telling me that they are going to lean in super hard into AI and then you blink and they launch v0. So now probably a lot of people in this room launch v0. So first of all, you identify peers that are leaning into this, as well as companies that don’t. So don’t be the other category, be that category. Then you need to play, right? So September. In December, MCP server, MCP comes out. Nobody knows what MCP was. Anthropic just launches it. Is it exciting? Is it not exciting? It doesn’t matter if it’s exciting or not. You have to play.

So they launch it in December, a week later, we launched our MCP server. That’s not a lot of technology, to be honest, but the important thing that as the world moves and changes, you launch your thing as well. The thing that I learned from Vinod Khosla at Khosla Ventures is the team you build is the company you build. And so if you want to be part of that new world, then you need to have that on your team. So we stood up an AI team, and I remember at re:Invent, that was November, I think, or early December. We were talking about this problem, and that Replit thing happened in September. In November, we’re like, “We need to have an AI team.”

Jon: And it was two weeks later you came back and said, “Done, we’ve got an AI team.”

Nikita: Yeah. And of course it’s hard to hire an AI team for a non-AI company. So modernizing your company and changing that perception is important. Then you can afford to tell a story, then you come up with a story, you stand up the AI team. The other thing that we did was kind of unusual. I was at the Menlo event and then Kevin Weil was there, chief product officer from OpenAI. I was obnoxious enough to invite him on the board. He said, “No.” But he said, “But you should talk to Nick Turley, who is the head of product for ChatGPT.” And then I went to Nick and invited him to the board. He agreed. So now we’re bringing talent into the company, we’re bringing talent to the board. What you also could do, what we also could do, what we didn’t, is to bring AI talent on the staff team.

But for that, we thought we don’t have a product yet quite to do that. But so then we kind of stopped because first you figure out what to do and then you figure out who and then you bring the who and then your team you build is the company you build and the company steers into the direction with a new set of people.

Finally, I want to say when you lean into something that’s new and uncertain, you will fire a bunch of bullets and you’re not going to score on each one of them. And so a bunch of things that we did didn’t work, but that comes to your batting average. So compare and contrast it with either being very, very, very careful choosing those bullets and then you hit every bullet, but you fire too few or not doing that at all, and then you missed the whole wave. So that’s what we did. Probably like four out of eight, nine bullets turned out to be really transformational. Two, three just didn’t work at all. And then the rest was kind of like a shrug.

Jon: There’s three points that stick out for me. And then I want to bring it back to the experience of Parallel. One is storytelling. And there’s a story that you told to me, and you were pretty loud about that hasn’t come up here, that at one point you were seeing agents create new databases at 4X the rate of human beings, which is striking and captures the imagination, and probably got Arjun excited and illustrated who is our real customer, one. Second is what I can only describe as a ferocious pace of execution. You were one of the very first companies that we saw actually have an MCB production server and that’s useful for…

Nikita: Yeah. And that was December, right? Then it blew up on Twitter in February. And now it’s like, I’m CP, whatever. And so again, that’s to your point is you have to play. Sometimes it’s expensive, by the way, because it distracts you from your other “core roadmap”, but you have to do it because if the user changes, then the user demands what the new features were.

Jon: Yeah. Parag, can you talk about how you’re building the team that maps to the customers, the end customers, and the intermediate technology you want to build?

Parag: Yeah, I have perhaps a couple things that were core to building the team. One, gently when you’re building in a time of extreme change, and for the future, we were going to… One of the harder things was that there wasn’t a market of AIs when we started working on this. We started building technology and infrastructure, but our customers hadn’t yet shown up in a very real way. And the customers that were around looking for the kinds of technologies we were building and the words we were using, we didn’t want to serve them at the time. And I think that guides the team you build. Number one, the team has to believe in this future and be persistent on it and be willing to not go take what’s available now and build toward this future. That’s number one.

Number two, the team has to be agile and adaptable. That’s why we built a very in-person team so that as soon as the opportunity presents itself, you are really, really fast in moving. Number three, the team must be… My philosophy is that you h-ave to… I agree with the idea that the team you build is the company you build. And if you want to take risk… We used to talk about go big or go home. If you want to take risk, you have to take risk on the team. So there’s basically two types of people on the team. One, people who add a bunch of risk and the corresponding upside to the company. And two, people who are great at working with people who add that kind of risk. So you kind of need both types to have a large team. And if everyone is only a risk adder without people who can actually channel the risk into a concentrated bet on the future, then you don’t actually get concentrated risk, you get diffused risk. So those are the only two rubrics of how we hire.

Jon: Got it. There are human beings at the end of this chain, and I’ll make a concrete example of why this matters in the case of Parallel. I have human beings who want to do searches on a new index, and Parallel can decide how much compute to burn, a little or a lot, based on how important that is. But Jon the human is not talking to Parallel. Jon’s talking to some app that rides on Parallel. So how do you actually get the UX right and the experience right end-to-end, such that Jon gets the information that he wants with the value and the level of compute that he wishes to burn?

Parag: It’s a really hard question to get the UX right, and that is why we leave it to the great customers we have to solve this problem, to interact and build. So we built a horizontal platform for agents to use. So we have customers who will do anything from sales to finance to recruiting to consulting to any kind of knowledge work effectively, where the web might be useful. And our customers have to do a lot of work around getting the UX right for getting work done. The thing we’ve actually… Our main prioritization rubric on what we work on is how much work or how many searches happen on the web for any single user action. And we believe that agents will use the web a lot more than humans ever have. Then you have to believe that the amount of work that happens is not limited by the number of humans.

We go initially working towards use cases where a single human action will trigger a large amount of work. Now, how might that happen? A single user action goes and triggers and fills up an entire CRM or an entire database with information collected, recent, overstructured, and insights extracted from the web. So you get a million rows in a database from one human action, or you have one human action that triggers this extremely deep research thing where an agent does what humans would take four, six, eight hours to do.

A single human action creates this always-on system that is now doing work for you repeatedly every day for a year. And so we go and focus on those use cases. Now, we’ve been lucky that there are a bunch of fast-moving companies. Some of them were even on a slide earlier today that are our customers who are innovating on the end user experiences. I know Jon is smiling because I was telling him that their lift is already still, we should be on there.

Jon: You got to come back next year.

Parag: Next year.

Jon: Well, that sounds like a good segue to a couple of minutes for questions from the group. The folks who would like to hear from Nikita and Parag, now’s the time.

Audience question: Great. Thanks for running this, Jon. So since you mentioned two calling a few times, of course, you had agents creating a bunch of database 4X -the rate of humans. I’m curious how you guys are thinking about setting the right guardrails and lives of trust on tools that can take actions that are irreversible. How are you thinking about the confidence there to invoke humans as little as possible, but yet have a level of confidence that you need to build robust systems that your customers will come back for, right?

Parag: I can start. So two things. I think when you think of confidence and guardrails, there’s two important things. One is how do customers trust that the agent’s end-to-end quality is very high? To do that, we’ve done basically two or three things, and that’s been our core focus. Number one, we build a lot of evals, which are general-purpose, but also customer-specific. In order to qualify use cases that’s going from this no longer works with agents to now it works reliably enough with agents. That’s a lot of the work. Second, we do a lot of work to… We have built models that actually evaluate the confidence of all the outputs we produce and produce evidence to them, and they’re actually calibrated against many, many evals. What that means is we are able to somewhat reliably know when we don’t get answers right. And that’s really important for enterprise use cases that we try to serve.

Now, in terms of guardrails, we’ve built a system that is read-only. There’s a reason that we’ve focused initially on being entirely read only and we do not… The words we use are we leave the web just as we found it. And that helps protect from some of the big downsides.

Audience question: What is the most trust you’d give an agent to do something for you? What is your personal level of trust to set an agent loose to manage finances or manage things in your family? I’m just curious, you guys are brilliant, right, so I want to get your own personal sense of how far you would go to be most courageous here to let an agent go into it?

Parag: I personally take a decent amount of… Get agents to do a decent amount of work. So anyone who signs up to our website, we know we run agents to figure out who they are, what they do using our own APIs. When I personally will let an agent to do all of my data analytics on our company for me, I don’t have a static dashboard. We let it fix bugs, let it even take the initial crack at coalescing and prioritize a bunch of things we should work on as a result of all the customer calls I might have done in the last two weeks. And of course, then we refine from there, but we burn a lot of GPUs experimentally.

Nikita: I have an interesting take on this reliability thing. I think Basis one time famously talked about two-way doors and one-way doors. So if your infrastructure is such that it’s a two-way door, meaning actions can be undone and verified, then I want to give agents an opportunity to just rip. From the infrastructure standpoint, when we change our user, we build a snapshot to restore functionality for the database. Now the agents can change the state of the database, but if it is not to their liking, they can go back very, very quickly. And the other thing that I think was going to happen in the infrastructure is a parallel execution. So whatever the notion of an environment is, so data and analytics, database is kind of one, but I think you can apply it for almost anything, for like CRM. So if you can fork your environment, M times, and then run agents in parallel, that will allow you to burn a lot more GPUs.

TK was talking yesterday in a different event that Google optimized the cost by 33X in a span of six months. So I think the cost of those actions are going to go down, and then the infrastructure around them will change such that you can try and go back functionality as well as try multiple things in parallel and choose the winner.

Now, there’s also one-way doors because maybe a customer interaction when you actually present results to a human could be a one-way door if it’s like the final result, or it could be a two-way door when you use a ,human for verification. So the more two-way doors we have and the better ways to, with evals or whichever other way to validate a human in the loop at the end of the day to validate that the results are good, the better it’s going to be. And now we can burn off our compute.

Jon: We could talk about this a lot longer, but I think we’re going to have to leave it there. Nikita Shamgunov and Parag Agrawal, thank you so much for joining. Thanks everybody.

 

This is How F500 Companies are Buying AI Today

 

What does it really take to sell an AI-native product into the Fortune 500? In this episode of Founded & Funded, Madrona Managing Director Matt McIlwain sits down with two founders deep in the trenches of enterprise AI adoption, Yoodli’s Esha Joshi and Gradial’s Anup Chamrajnagar. Their companies are selling into some of the world’s most complex organizations, like Google, SAP, Snowflake, Databricks, and more. And they break down what founders can sometimes underestimate about enterprise AI sales.

Listen on Spotify, Apple, and Amazon | Watch on YouTube.


This transcript was automatically generated and edited for clarity.

Matt: So Esha, tell us a little bit about the original wedge and how that original wedge allowed you to enter the enterprise market.

Esha: Yes, my name’s Esha, and I’m a founder of Yoodli AI Roleplays. What we are building is a native AI roleplay and conversational coaching platform for customer-facing teams. We started the company to help people become more confident communicators. And what we do today is we are enabling sales account executives, customer success managers, sales engineers, and support specialists to practice high-stakes conversations and get personalized real-time feedback in a judgment-free zone without a human manager physically in the room. So it’s a very fun time. Our wedge is with sales enablement, helping our enablement teams ramp hundreds, if not thousands, of salespeople very quickly and onboard them without having to worry about consistent training and the lack of consistent messaging across an organization.

Matt: Anup Gradial is also a young company, about three years old. Tell us a little bit about the original wedge for the company and how you won those first couple of enterprise customers.

Anup: Gradial is an AI-native accelerant for your content supply chain. What that means is, generally speaking, a lot of AI for marketing tools that we saw a couple of years back, first when ChatGPT recently came out, was mainly used for things like content creation. So blogs, article pages, base level, simple images, but nothing else that happens after that, which is actually 80 to 90% of the bottleneck on why things take so long to get out in enterprise marketing to begin with. So we tried to build what we call the configuration co-pilot to help software companies configure workflows faster, where you have a configuration problem at the beginning and a continuous configuration problem throughout while trying to update your site and make new designs and code new components. And so that was really the problem we set out to solve first was solving content execution at scale for some of the largest companies in the world to help people get content out faster and get more scale and get more efficient at the same time.

Matt: Well, let’s talk a little bit about what the state of the world was a couple of years ago. It was a time when a lot of people were piloting things and trying things, and some of those worked and some of those didn’t. What did you learn about buying behaviors to be able to tell the difference between somebody that was just purely a tire kicker, somebody who was serious about the pilot and then what you ultimately needed to do to be able to get that person to be a committed long-term customer?

Anup: Yeah, I say this thing around the room that we’re not selling a solution, we’re selling pain. And I think that was something that we realized early on that if we could get to the actual day-to-day pain of our customer as quickly as possible and have them recite back to us that they basically empathize with what we’re saying, then we can find ourselves in a situation where we can actually co-opt a solution together in the early days. And they started saying, “No, that’s not it.” And they give you little tweaks over time, but it was a mixture of selling the pain very accurately and then being able to basically show them a demo of something that was very consistent across the industry, and I’m sure you guys find this as well, of enablement. I was joking that I should have probably used Yoodli before this entire podcast, but it’s industry-wide, vertical wide, you try to find this similar pain and then that’s how you start achieving scale with that kind of stuff.

Matt: Esha, what was the experience for you with some of those early customers that were piloting things? I remember some of the times where you had these really small pilots and their buying process and how that has been evolving in this era of AI-style solutions.

Esha: I wouldn’t say it was easy for us to get in those pilots, but we were certainly doing more feature selling and showcasing what the value of AI functionality could look like, which in the Yoodli realm was all about these contextualized back-and-forth conversations for people to be able to practice ahead of a conversation. So a pitch, practicing skills of discovery that had never been done before. Previously, what was status quo is humans and human managers and their direct reports have these manual, awkward, in-the-room conversations, roleplays, they were inconsistent, they were a little bit stressful. And then when you took technology and put that in the mix, you then had very static talking to Smart Mirror with a very canned response. So nowhere seen before had they had these really contextualized conversations that took in the context of company collateral.

So to answer your question, it was still hard because it was enterprise sales, but getting people to really see the value and dream was, I would say, easier back then than it is now.

Matt: Now, so why is that? Have the expectations gone up now in terms of somebody being even willing to do a pilot, and how have those expectations been changing?

Esha: I think the expectations are different from different companies. So certainly IT, tech, SaaS companies, they have different solutions, different founders metaphorically knocking on their doors, outbound messaging, cold calls, maybe they have an AI bot doing outbound AI-driven emails. So there are a lot of opportunities, and there are a lot of solutions that you can look at.

So now I would say buyers are not just looking at features, they’re looking at the outcomes you are providing, and if the user experience bar is raising? I feel like it’s gotten a lot harder to make sure that your user experience is high enough. It’s very difficult to keep people delighted, and I think there are incumbents for certain solutions that have been doing this that are not necessarily an AI-native platform. I think you heard both of us say that we’re AI-native platforms, meaning that the flexibility of adding on capabilities that drive certain outcomes, whether it’s time saved, revenue attained, or whatever the equivalent is for you, it’s a bit easier, I think, for both of us to do that because we can add on AI functionality.

Matt: Yeah, so you’re moving on from the wow factor and you’ve got durable outcomes, durable impact. That is what the customers are looking for today. Is that what you’re finding as well?

Anup: Yeah, exactly. I mean, in the form of a pilot, so I mean, you either go in with preliminary engagement, or if the scope is very large for what the client wants to do, then you obviously run a pilot of a microcosm of that, and then run success metrics. But I think the thing that was successful is that many clients or prospects will initially use words like, “Let’s do a POC.” And we try to typically circumvent that to at least say, “Let’s do a targeted pilot.” Because I think the difference is everyone asks, and I’m sure you guys get this all the time, “Do you guys have a sandbox we can play with and test out a couple scenarios?”

And so for us, we know we’re solving an end-to-end problem for folks. We know we have enough reference ability; we know everyone will come in and say, “We have such a complex environment, can you do it?” At this point, we know we can. And so what we try to do is say, “Hey, here are the three things that we’re going to get out of this pilot as an objective. Do you guys agree? Can we build a mutual success plan together and say, ‘Okay, if we come out of this, we’re going to roll it into production and actually get a real thing going here?'” So I think you don’t let the pilot die on success.

Matt: Right. Right. That’s a great way to say it.

Anup:

Because I think a lot of AI companies will be successful, and then it’ll die out. Right? And so you need to make sure that you get that commitment upfront so that when the work is actually put in, it’ll result in a good outcome.

Matt: Esha, do you do a similar type of thing to validate that, “No, we’re serious about this. We’re not just doing pilots to kick the tires. We’re doing pilots today so that this company can be an impact player with us”?

Esha: Absolutely. I would say before we enter into a scoped pilot, call it 60 days, 90 days, I think anything longer than that, you’re starting to extend the time with which you can get a clear answer. We try to scope it ideally like 45 days. And we say, “During the 45 days, here’s what we will give to you and here’s what we need to have in place before we get into it.” Ideally, you align on pricing, but even that’s asterisks because procurement will get in the picture, and then it’ll become a negotiation all over again. But align on pricing, align on success metrics, be very declarative on the outcomes you can give to them within a 45-day period, what you need from them. And then also beyond that, tell the story of what this is, just the starting point of, and what this enables in the future, with this existing use case with 3x the number of people or even more use cases in the organization, depending on how you roll out.

Matt: Let’s spend a couple of minutes on that because we talked about the early wedge, the initial problem that you were each trying to solve. What are you starting to see maybe even with some of those earliest customers that you won, how they’re starting to pull you into adjacent use cases?

Esha: So I talked about how our wedge is sales enablement. It’s still a primary wedge for us, and we find that we start with sales enablement in particular. But there is go-to-market enablement. So it’s very easy to make the case of moving into customer success managers, moving into sales engineers. So that’s one direction we’ve moved in. Similar types of folks, very similar skills and competencies and even onboarding process.

There’s another area we expand into, which is not just onboarding and real-call preparation, but also knowledge acquisition, and that is definitely enabled by a lot of the advances we’re seeing in AI today. But if you think about it before you get into even a podcast like this, I’ve obviously practiced for this podcast, but in order to practice the content, I need to know the content I need to push myself.

Matt: Right. So there’s a learning dimension to it too. So you’re getting pulled that way.

Esha: Exactly. Before this, I would’ve talked to an AI version of Matt and have rehearsed an answer and been like, “Matt, does this align with the story that you’re trying to tell with the Madrona ecosystem?” And you might’ve said, “Well, maybe you think about pilots and scoping it out this way?” So that’s the experience that we now get to give our customers before they then practice for the podcast.

Matt: Anup, tell us maybe about one of those early customers who you won and now they’re pulling you into adjacent use cases.

Anup: I can think of a large telecom provider who was using us primarily for web authoring across their digital experiences. And there was a particular instance where they had their largest partner changed prices on them, and usually this very large partner doesn’t tell any of their partners about the price changes because they do it overnight. And when that happens, you have to change all of your digital experiences to reflect that new price or else you’re technically liable for that price increase. And in this case it was $7, so it was pretty substantial across their customer base. And so it usually takes three weeks to go through the digital experiences, sweep everything, actually replace the content with the right thing, go through brand compliance and legal and get it approved, it took them 30 minutes with prompting with Gradial. And the email team took note of that and they were like, “Can we use this?” And the creative studio team took note of that moment and said, “Hey, we have assets that we have to replace all the time. Can we use this technology to do that?”

So seeing the use cases in action and broadcasting that across the customer bases will naturally fire it up. And if they don’t have to go through procurement themselves and there’s already an MSA in place, they’re much more likely to be pretty gung ho about it, so…

Matt: Absolutely. So maybe let’s pick back up then because that was really helpful as well. But I want to pick back up on this notion, okay, I’m in this pilot, I’ve set the ground rules of the pilot, gotten alignment on the objectives. Is there anything else in your playbook that you think other people should understand about how do we get through that pilot successfully and then fully upsold into some kind of an annual commitment?

Esha: So my background was not in sales, starting Yoodli. I had a background in engineering and product. But of course, founders become the first salespeople of the company. And I remember with one of our very first deals, we found a champion, she was excited, saw the value, saw even the potential outcomes, and I was like, “Great, my job is done!” And then it turns out that you have to then enable this whole ecosystem. So champion is step one, then you have to convince IT, you’ve got to convince procurement. Now, with a lot of SaaS IT companies, there’s an AI governance council; you’ve got to convince the budget holders. So to answer your question, going from a pilot to a full 12-month contract, really anything with enterprise sales is you’re convincing the whole ecosystem. And every single person in that ecosystem has something that they care about, whether it’s a discount, risk mitigation, outcomes to the company, onboarding, whatever it might be. And I definitely did not think… It makes sense intuitively, but I did not think about that when I first started.

Matt: How about you, Anup?

Anup: Yeah, I think reading, listening to any founder talk about it could not possibly prepare you for all of the things that Esha’s basically talking about, of needing to appease the people who are higher up that the champion reports to, needing to also get the lower rung. So there’s a three-legged stool to enterprise sales. You need to get your champion, you need to get their VP or the CMO, and then you need to get the person who is the power user of the product. Right? And to get all of those, you need to work fast, right, during the pilot to actually get them all engaged. And so you start learning how to expedite and some tricks up your sleeve to accelerate parts of those.

And there’s this whole new concept in a lot of enterprises called the AI review board. Right? This wasn’t a thing two years ago. Many of the people are just procurement people that basically got a new job at their own company to lead this board. And so, anticipating some of those questions that just get asked because there will be added steps because you’re an AI vendor that will come up, but there are also other things because you’re an AI vendor that people want you to win. You try to make people win with you as well, and there are different tricks we can get into that we’ve thought of to try to accelerate that.

Matt: What strikes me here as interesting is outside of the AI review board dimension and probably some of the additional data security layers, because it used to be more software security validation, how similar this enterprise, understand the different players that are in the mix is to classic selling. What I’m interested though is it seems like both of your companies, you both had spectacular growth last year and on a great start again this year. It seems to be happening faster. And so what is it about how you run a sales program in an enterprise? Maybe a lot of the same techniques that are around, but how you’re doing it with more agility and with more speed?

Esha: Step one, we have more salespeople.

Matt: So that’s more capacity. That’s one.

Esha: Capacity. So we’re going with more speed.

Second, specifically for Yoodli, some of these big enterprise accounts, Google, Snowflake, SAP, we have initial success stories. And so we have little groups of champions across the organization and they’re speaking and we’re sort creating these serendipitous moments where we’re getting them to speak up by planting success stories and intel in their laps that they may not know from other parts of the other side of the organization. And to your point, we have an MSA and we’re through the security process. The AI governance process is from my experience, a royal pain in the butt. And a lot of these companies-

Anup: It always comes in at the last second, too.

Esha: Exactly. And a lot of these companies are like, “Oh man, I hope we don’t have to go through that again.” And so I think the fact that we’re already in through the organization and there’s some success stories, capitalize on that momentum and make them feel the pain to not move off and go somewhere else.

Matt: Anything on your end, Anup? I completely agree with Esha on adding more capacity, anticipating some of these nuances. On deal velocity, anything that you can offer there?

Anup: I think the concept of revenue operations or sales operations, I think everything has a KPI tied to it. And I think in enterprise selling, the concept of ops should be to reduce the sales cycle as much as possible because that’s where they can actually influence it, right, filling in all of the gaps.

I mean, I’m sure you and everyone went through the same thing; it’s all founder-led sales at the beginning. You do so many things naturally that you don’t really think through the gaps that may emerge for a person that’s newly coming into your company and needs to sell. I’m sure you guys have it on lock with using Yoodli to actually train your sellers. We don’t, I wish we had that, but basically giving them the runs, that’s having them shadow folks, you can actually start spreading that around a little bit more.

And we have a go-to-market hub at our company that basically has anything and everything you would need to know to sell Gradial, and we’re always adding to it. And it’s a collective knowledge base, so everyone can add to it. I’ve assigned owners to each section of it that gives them ownership over a particular section of enablement, so they really take pride in that. And one person owns workshops, one person owns different kinds of way of communicating pricing, one person owns the BVA deck, one person owns the pilot cadence. And you can have them be the centers of excellence and try to scale what used to come to us. All the questions that come to us, you try to scale that a little bit better. So that’s been much easier.

Esha: Good idea.

Matt: And I think I would describe that almost as AIOps. You’re using collaborative systems, but also AI to help make your own internal business better to help you accelerate sales cycles with your customers.

Anup: I mean, it was pretty unbelievable. With Opus 4.6, we plugged it into HubSpot, and I asked it exactly what I asked my VP of sales every week for a full Q1, Q2 forecast. It gave me every single detail I could possibly have wanted, and it made it an interactive app that I can just click through and see whatever I wanted. So it sent it immediately to me and the other three guys, and it was pretty awesome.

Matt: Now, in your customers, though, that kind of story might make them nervous. And I think one of the other aspects of this AI world is the customers you’re selling into worrying that by adopting your technology, you might be replacing their jobs or facilitating other people losing their jobs inside their company. How do you navigate that issue?

Anup: I think the most important thing is we talk about the three-legged stool and the champion, but realistically, everyone is a champion. And in this case, the power users need to be champions of the product. Right? And I think something that I model after is the way Clay built its ecosystem, Vercel built its ecosystem. People are so proud to talk about the way that they’re using these agents and this technology that they write pieces, they have training videos that they… And this is completely unprompted by the company themselves. Right? There’s whole communities that have risen up from the ground on being an expert in Clay, right?

Matt: “Claygents” I think they call them.

Anup: Yeah, exactly. And you need to build that camaraderie and basically that base up somehow because you’re not going to get replaced by AI, you’ll get better by AI. Right? And you’ll be one of the people, one of the few people in the world who can actually be a superhuman with AI. Right? So partner with them through the journey of actually having that super remote control, I think has been a key focus of ours this year and the end of last year, for sure.

Esha: I think Yoodli has given our enablement team at companies, at our customers, like superpowers that they never had before. So first, we’re able to now quantify the improvement of people in the organization in a way that we couldn’t. And so we’re making our champions, which are enablement leaders and enablement stakeholders, look like superheroes because they’re able to go to their CRO or go to their people leader or talent leader and talk about here’s what this tool system has done to enable folks, and here’s what this actually means from a revenue attainment standpoint. So they feel really great about that.

I think personally, and one of the reasons why I’m really excited about what we’re doing and why we started it is we have stories of people saying, “I was not replaced by Yoodli AI, I actually got promoted because of Yoodli. I practiced, and I became an AE or a senior AE, and before I was just an SDR or BDR, and it feels good.”

Matt: The adage that you’re not going to lose your job to AI, you’re going to lose your job to somebody who’s embracing AI.

Esha: Yes.

Matt: So you’re looking for those kinds of champions that are, “I’m going to embrace, I’m going to invest in this technology to help advance our company and inevitably help advance my career.” What happens though, because I’m sure whether it’s in the pilot stage, we are working with non-deterministic models and non-deterministic systems. There’s probably some good stories. I don’t know if you want to tell any of them of like, “Oh, this is when things didn’t go quite right with our system, with our models,” and how do you recover from that with a customer? And so because this trust element of I trust not only this company, but I trust their underlying application and the models they’re using, you have a good story or two to share there and how you navigated it? Do you want to start, Anup.

Anup: There are a number of examples of this. People claiming that the tech earlier on, especially not as much anymore, but that it didn’t work, that something that they ran just clunked out and just didn’t do it. And we ran the same thing and it did work and you had to try to show them, but they were like, “No, it didn’t work when I ran it.” And there’s no way to prove that other than go back. And so it was a little bit of trouble here and there of saying, “Okay, some people think that they see some writing on the wall and they’re reacting in a certain way.”

And we’ve learned from those lessons of how do we bring everyone together on a culminating success plan? And that’s what led to the mutual success criteria here of how do we actually beforehand establish if we completed these things, we would be successful with you guys. And get everyone’s agreement. The VP was in the room, the director was in the room, so now everyone feels like it’s an actual assignment that they need to do versus, “Hey, I tried this thing and I tried some things and it just didn’t work.” So, you’ve got to keep a close eye.

Matt: That’s great. Yeah, there’s always going to be bumps in the road along the early adoption phase, and so if people are focused on the end goals, then if somebody takes that mistake that happened and tries to use it against you, it’s like, “Well, let’s focus back on the goals we agreed to.”

Anup: Yeah, the pilot doesn’t stop either. I think we’ve learned that the hard way as well is just because you have a production contract doesn’t mean that they won’t still treat it as they want tangible goals, they want tangible ROI. And so we have a philosophy that the pilot never stops. Right? We want to keep having this mutual success plan keep — is the thing you want to accomplish at the end of the year and how do we help you get there?

Esha: Yeah, what you were saying, it’s either the pilot never stops or really the renewal is actually in the first three months. That’s the other one that I keep hearing. It’s rarely happening throughout the course, those first touch points.

But going back to your question around what’s a story that you can share. Early on in an early adoption phase, one of our customers was using Yoodli for a product certification and had used Yoodli. And this was a company that had several thousands of sellers using Yoodli. And one individual person reported that the product made a comment on their clothing, like the AI referenced the clothing.

Anup: Wow.

Matt: Yeah, a fashion review.

Esha: A fashion review. A fashion review. And I talked about that to my team and various members of the team were really upset. Other members were, when they heard it, they were like, “Well, I should go have AI comment on my clothing.” But long story short, it was obviously very disappointing. And so what we needed to do is, our engineering team was figuring out what kind of model magic could happen on the backend to serve that. But then we had to go talk to that individual user, hear them out, reassure them that they’d be good for the future. We had to go talk to our champion and the leader above and basically regain trust. It was a process, it was nerve-racking for a second, but I think we’re on a better path now.

Anup: That’s interesting. I guess sometimes in these situations, I mean, it’s better in a sales training and not a product sort of situation, but sometimes people on the other side are very unprofessional. Is that part of your guys’ work? How do you react to someone who’s being super callous? I mean, that happens all the time, especially in sales situations.

Esha: 100%.

Matt: Yeah. It’s maintaining your respect for them as a customer. And even if it’s like, wow, that seems to be a little bit out of bounds, but we’ll navigate it through and then solve those problems.

Anup: Yeah, that and a real-world scenario that you train for is someone being mean. It’s just like on the other side, they’re just like, “Yeah, I don’t like this.” Or they just came in with a bad mood. Is that something that you guys train against too?

Esha: 100%. We account for that. You never know how somebody else is going to react in a conversation. And there have been instances where I’ve left conversations or co-founders left conversations where that person was really mean, and you just deal with it. Not everybody has maybe as much thick skin. And so that’s actually what we say, we can build personas with different personalities ranging from really timid and meek to really aggressive and mean and rude.

Anup: That’s awesome.

Esha: And it’s not a bad thing, it’s just realistic. That’s the whole point of this.

Matt: Well, let’s switch from some of these product and selling cycle experiences to something closely related: pricing and packaging. And so maybe you could first tell us about how your product is priced and packaged today and how you’re thinking about continuing evolve. But let’s start with how it is today.

Esha: We’ve been experimenting with pricing for a very long time, so I would say our most frequently used pricing and packaging paradigm is seats-based model with a fee associated. Fee can be anything posed as a platform fee, which is the cost of AI features and functionality and also customer success support as part of that. We typically use that. We’ve had some pushback from customers, and then we can switch it up. Typically, the customers who push back or the ones who themselves don’t value and don’t sell to their customers seats and another fee, they’ll position it either as purely just seats with add-ons on top or usage-based pricing models. And so, we haven’t yet really seen an AI outcomes pricing model approach, which I know is causing a lot of excitement. But we’re keeping our eyes out.

Matt: But you don’t have much of a consumption component to your pricing system today. It is more on the seats-based and a platform fee type of a model.

Esha: Correct.

Matt: And then not an outcomes-based either.

Esha: Not an outcome-based model.

Matt: And, Anup, yours is a little bit different, so let’s hear about that.

Anup: We’re almost the opposite. But we do have a platform fee and a support fee. Those are two parts of it. The bulk of the fee from basically a volume-based metric because at the end of the day we are measured on output. So it’s like you’re more efficient with Gradial and you’re able to get more stuff out the door with Gradial and people have these output goals that they want to reach at the end of the year. Right? I want to have 200 more personalized blog pages, or I want to get 5,000 more emails out the door to targeted. So there’s I want to do this much more analysis. There’s an output measuring of it. And the good part is in marketing, there’s actually a very small finite list of those things, of those artifacts that people want to actually push out the door.

And so ours is a little bit, we map out a tentative plan of what they want to do and say, “Can you get more ambitious than that?” And they’ll be, “Amazing if we reach this.” And then we package that up into basically a volume-based metric of output. Right? So we don’t charge for consumption per se. If you’re asking it like a simple query, you can get that query out, but if it actually produces something of value, then that gets action.

Matt: And my sense is that that is generally working for you and that is what you plan to run with. Are there areas around the edges you’re exploring or experimenting like Esha was saying?

Anup: So, I think version one was seats-based module, use-case based. The module use case was how providers were doing it. And so we modeled after the best that there was, right?

And now we had the general pool of actions. And I think where we’re experimenting now is, we actually realize that when a customer tells us they want to get this host of artifacts out the door, why translate that list into some amorphous definition and then read translated when you’re trying to communicate the business value? Why don’t you just say, “We’re going to pledge to deliver that. Exactly what you want.” Right? And then find a way to, because we’re lucky that we are in a space where that list is very finite and it’s always going to be finite. It’s an artifact that you’re creating, whether it’s a page or an email or some kind of analysis report or a buyer report or an update to something. It’s always going to be like this finite list. And so if you can have this uniform list of categories, then why not just align it directly with what the customer wants to see at the end of the day? So that’s a little bit, it’s like it’s still volume, but it’s more tangible business value volume that is easier to explain to procurement.

Because the other thing we found is when we had the amorphous definition, oh, it was…

Matt: Tough for procurement.

Anup: Yeah. They’re like, “Okay, please walk through multiple types of prompts and use cases and how that would result in this actions, and do you guys have a dictionary?” And then, “Oh, well, how many actions would this be.” We are starting to get those questions a lot, so now we’re getting a little bit more simpler and saying, “Hey, it’s basically what an agency does, honestly, it’s how they price.”

Matt: Any other thoughts on this of where it’s evolving?

Esha: We’re seeing it evolve more to volume-based pricing because our tool is associated with behavior change and with educational learning. And because it’s behavior changing and practice and learning is hard, it’s less based on some folks, again, who sell this way to their customers want it based less on how many seats they’re buying, but the usage of those seats they’re getting. So different ways we’re doing that is of the actual seats that are being used, can we price on that, or of the unit itself, which is the role play practice in our case. Based on that you have a per-units cost.

Matt: Consumption type. Got it. Maybe let’s just touch quickly on two different things. One is an example of a key go-to-market partnership. You all are living in ecosystems and so maybe, Anup, you can give us an example on that front of how you’re partnering. I often talk about the AI ecosystems are really an important part of the equation. And then love to hear one from Esha as well. And then we’re going to go into a couple quick rapid fire questions.

Anup: We have a multi-pronged approach to partnerships. I mean, as we all know, enterprises are tough beasts and you need to attack them from multiple angles in order to conquer the mountain here. And so agencies was how we basically started off. So agencies dominated our space. Outsourced marketing spend is the biggest outsourced services spend category in the world, so about half a trillion dollars of spend in just the Fortune 500 per year. So agencies basically reign supreme there and they have all the relationships, but they’re also called upon to implement AI for their clients.

So we caught onto that train and started making relationships with a lot of agencies, the blue chip top agencies in the space. We purposefully chose, you could go in two directions there, you could go with blue chip top tier, you could go with the commodity volume players. And we went with the former just because we saw ourselves as sitting in that camp of people associate blue chip with similar things, and we wanted to continue to be in that. So agencies like Dentsu, Accenture Song, Deloitte Digital, folks like that that of top tier in that.

Esha: A lot of similarities in what Yoodli experiencing and what you’re saying. So the equivalent of agencies for you is coaching and training companies for us. So this can be anywhere on the spectrum of smaller Mom-and-Pop coaching companies. When I say coaching, executive communications, interview training, sales training. So anywhere that has 10 or fewer people these kinds of companies will use Yoodli to augment the coaching they do, to something like a Toastmasters, which is one of our first partnerships that has incredible distribution to like 4,000 Toastmasters, to anything now on the other end, like Sandler Sales or Corporate Visions, which are two of the top most well-known sales training companies. And in either case a lot of the bigger ones, like you said, have a mandate to use AI to help with their training and enablement and reinforcement of whatever they teach. And also there is additional exposure into their clients of Yoodli that’s sort of branded by them.

So I would say it’s been a really interesting process with the partners because you have to enable them so that they can go and sell your product, and also you have to work together symbiotically. So we’re learning a lot by that every day.

Matt: Do either of you use this concept of a forward deployed engineer or is your product so ready out of the box that that concept doesn’t apply?

Anup: We definitely have a forward deployed team. So every customer gets an account director or a deployed PM and a forward deployed engineering team. And there’s a lot of help that’s needed at the beginning to kind of introduce them into the agent world, because for some of our customers we have, I know we talked about some of the technology customers that we have, but most of our customers are actually legacy incumbents in financial services, healthcare, industrials, automotive, telecom. A lot of users in our platform oftentimes have never actually used ChatGPT before, which is kind of crazy in this year and day, but it’s the truth. And so how do you get people through a complete paradigm shift at the beginning has been interesting. Obviously, that gap is shortened, but it’s good to have. And everyone, all these enterprises like having a helping hand, they’ve always paid for carte blanche white glove service, and you need to give them that.

Esha: The equivalent for us is a sales engineer slash solutions engineer, also experimenting with professional services, to your point around white-glove service. Technical account manager for some of the really, really incumbent legacy IT that need that. And so, it’s all kind of a derivative of the same thing with slightly different skills. But yeah, certainly needed.

Matt: Well, we could talk all day. I’ve learned a ton just listening to both of you. So I think I’m going to end it there and hopefully in a couple of years when you all have continued to learn all these lessons of being truly agentic AI-native companies selling into the enterprise and making those customers successful, we’ll have another conversation. So thanks, Esha and Anup, very much for joining me.

The Infrastructure of Intelligence: Inside Crusoe’s AI Factory in Texas

 

In this episode of Founded & Funded, Ben Gilbert, co-host of the Acquired podcast, sits down with Chase Lochmiller, co-founder and CEO of Crusoe, the company building what it calls AI factories, including its massive campus in Abilene, Texas, which are designed to power this new era of intelligence.

In this conversation, Ben and Chase explore the physical reality behind today’s AI revolution. Why modern AI workloads demand entirely new infrastructure. How energy has become the primary bottleneck to scaling intelligence. What it takes to compress multi-year building timelines into months. And how Crusoe’s energy-first philosophy, from capturing flared methane to siting facilities near abundant wind power, shaped its path to building one of the world’s largest AI computing campuses.

This is a must-watch for anyone building in AI or rethinking infrastructure for the next era of intelligence.

Listen on Spotify, Apple, and Amazon | Watch on YouTube.


This transcript was automatically generated and edited for clarity.

Ben Gilbert: So, I hear you’re building a really, really big AI data center in Abilene, Texas.

Chase Lochmiller: This is accurate. We have a project in Abilene, Texas, that we’re doing with Oracle. This has been a very special project to work on because it’s really been reinventing a lot of the infrastructure layers of computing that are really the base substrate that enables all of this innovation and enables innovators to do their life’s work.

And we’ve really had to rethink the data… We don’t even really call them data centers; we call them AI factories. These are the factories of intelligence because we think about it as one coherent cluster of computing. When you look at it from an aerial view, it almost looks a bit like a motherboard when you’re sort of looking down on it, because there are these giant clusters that are all interconnected and designed in a way so that it can think as one coherent, giant brain.

But the scale of the infrastructure is just really dramatic, and it’s a huge shift from legacy web applications to the energy needs and the infrastructure needs to support intelligence at scale. Maybe to give you a couple quick tidbits about this site, so there’s about 1.2 gigawatts of power that powers the site.

Ben Gilbert: Can you help us understand that compared to historical norms?

Chase Lochmiller: Sure. So, I’m from Denver, so Denver runs on maybe a little bit less than 1.2 gigawatts. It’s about the power of Denver to power this data center.

If you look at Northern Virginia, where I think many people would consider the center of the world for data centers, this is where the bulk of the internet runs. At the end of ’24, JLL published a report, which states that there are about 4.5 gigawatts of total capacity in Northern Virginia. So, this one site, is a quarter of that.

Ben Gilbert: And that’s like all the hyperscaler clouds have very large presences in Northern Virginia.

Chase Lochmiller: Exactly. Yeah. And I think that’s sort of the evolution that we’re seeing. I mean, Sam recently just announced that his big KPI that he’s targeting his team for is basically one gigawatt per week. So, it’s basically one of these projects per week that would be delivered. It’s like total insanity. And 250 gigawatts by 2030.

And just maybe speaking about a bit of what it takes to make this happen, right? I think there’s a lot of talk about all the jobs that AI is going to take. We’re creating just an insane amount of jobs. It’s wild. We have 7,000 people on site working on this every single day. And this is in a city of… Abilene, Texas is city of 120,000 people. So, having 7,000 people working on this one project is a tremendous amount of labor, and it’s a lot of blue-collar trade work. It’s electricians, it’s plumbers, it’s construction workers that are making all of this infrastructure happen.

Ben Gilbert: So, what we saw here is two buildings. And John, could you loop this video again? What we saw here is two buildings. Across the street. There are six more of these, is that right?

Chase Lochmiller: Yeah, you can kind of see them in the background there. So, basically the first two buildings we started in June of 2024, and-

Ben Gilbert: That was dirt in June of ’24.

Chase Lochmiller: It was like dirt and mesquite trees. And then actually behind it, you see the six buildings going up behind it. Those started in February. So, it’s the whole project has moved very, very quickly. I think one of the reasons we sort of had the opportunity to do this is that we kind of came over with a bunch of creative ways to really accelerate this. So, there was actually an RFP that went out for the initial first 100 megawatts, which is going to be one of these buildings. And the next, fastest bid to basically make this happen was two and a half years. And they called me, they were like… I was the 35th person. I was the last person that anybody called to do this.

Ben Gilbert: Crusoe was a seven-year-old company. You weren’t a startup.

Chase Lochmiller: Yes, totally. And they were like, “Do you think you could do this in a year?” I was like, “Yeah, for sure. Absolutely.” But we had to do a lot of creative things. And I think part of it was the fact that I had a lot of great data center insiders on my team, but I was very much an outsider. I had never built anything of this sort of scale. And I was really coming at it from building out large GPU clusters that we’d been doing for our Crusoe cloud business, and really having this framework of thinking through, “Okay, what does it take to build out a large-scale coherent cluster? And what is the actual building design, the cooling design, the power design that you would want to have to support that at a much bigger scale?”-

And that led us to all these different modular optimizations in terms of how we brought this infrastructure to life. We actually, Crusoe, we stood up a whole… It’s kind of the journey of an entrepreneur is so funny because you end up doing all these things that you wouldn’t have expected yourself to do when you started.

We have a whole manufacturing arm at Crusoe called Crusoe Industries, where we have a factory that makes electrical equipment for data centers. And the reason we did it was so that we could control our own supply chain, and we could control time-to-market for a lot of these things. So, to give you an example, there was a critical component called a power distribution center that distributes the medium-voltage electricity, the medium-voltage power, to get to the data center. It distributes to all these low-voltage transformers. When we went out, and we were requesting bids from all the different suppliers in the industry, the fastest offer we had was a hundred weeks. And we were like, “I committed to 12 months, so a hundred weeks isn’t going to work.”

Ben Gilbert: And a hundred weeks, that’s like three or four big foundational model releases for Oracle’s customers.

Chase Lochmiller: That’s not going to work. So, we figured out how to make it ourselves, and we can make it in 20 weeks. And so it was a lot of different modular components that we were basically manufacturing off-site.

What you see is you see these big buildings, but a lot of the guts of the data center, the electrical components, the switch gear, the low-voltage UPS, the RPPs, the hot aisle containment systems, these are all modularized into these data center Lego blocks. And what we do is we actually do that off-site in a controlled manufacturing environment, and then we bring it to the site inside this big building, and then we assemble it. And it’s kind of a lot faster to do that on-site than having to build everything on-site, essentially.

Ben Gilbert: So, just to make everyone in the room aware, a lot of the AI applications you are using when you go kick off some interaction with image generation or a chatbot is happening right here in this building. And it’s pretty recent that this has been full of GPUs and actually operating.

Can you take us through what it actually takes, the inputs to a site like this, and how they’re different than the old world of building sort of classic data centers versus the new world of AI factories?

Chase Lochmiller: Yeah. I kind of touched on this, but I think just from a very first principle basis, the number one thing is we think about this as one giant, coherent cluster, and then you sort of end up building and designing around that.

Ben Gilbert: The data center is the computer as Jensen would say.

Chase Lochmiller: Exactly, the data center is the computer. And because of that, you see that central core, that T in the center of the four wings, that’s where all of the network and storage lives. And then it distributes out to each of those four wings, where you have these giant clusters of liquid-cooled Blackwell GB200 NVL72 racks. And on the perimeter, you see this stream, it kind of looks like we were sort of talking about-

Ben Gilbert: It looks like RAM.

Chase Lochmiller: It looks like memory. Yeah, yeah. But it’s actually, they’re chillers, it’s air-cooled chillers that are cooling the water. So, we have a giant water loop. There’s a million gallons of water per building that are cooling these high-density GPU racks.

And while these buildings look quite large, it’s a way denser configuration than a traditional web cloud data center. For this amount of capacity for an AWS data center to serve EC-II or something like that, it would be probably three to four times bigger in terms of square footage. So, each of those buildings is about 500,000 square feet, probably 1.5 to 2 million square feet, if that were a traditional web data center.

Ben Gilbert: And so obviously, this increased power need has to come from somewhere. Where do you source power from?

Chase Lochmiller: Power is definitely the key bottleneck in a lot of this. And I think we’ve sort of seen this evolution of scaling laws and scaling AI infrastructure, and it’s very rapidly saturated, the infrastructure we have to support computing and the energy we have to support computing. So, that’s led us to a place where we fundamentally just need a lot more new energy generation, and we need a lot more new data center development.

Crusoe’s always really taken this very energy-first approach to computing. It’s what led us to build in Abilene, Texas. Abilene was not a data center market before we put a shovel in the ground in Abilene, and now the world knows about Abilene. But what brought us there is actually, it’s an area of Texas in West Texas where there’s very abundant wind energy. It’s one of the most consistently windy areas of the country. And what had happened was a lot of wind developers had built out these large-scale wind farms, and they were having to curtail because power prices were going negative, they were getting paid these production tax credits, which only last for 10 years, and at the end of the 10 years, you’re subject to economic curtailment or face negatively priced power. So, it wasn’t like a great outcome for these renewable energy developers.

Ben Gilbert: They have wind turbines that are sitting there that they’re intentionally not letting run?

Chase Lochmiller: Yeah. You’ll drive by it on a windy day, and you’ll see this wind turbine not spinning. You’re like, “What’s going on? Is this thing broken?” But it’s actually that they’ve turned it off because of economic curtailment. So, there’s no marginal bid for that power. And so this is obviously… AI needs a lot of energy. They sort of had a lot of energy. And it made sense for us to basically, instead of trying to build the next data center in Northern Virginia, our focus was bring the demand for compute to areas where we can access low-cost, abundant energy. So, that was one of the big drivers of us coming to Abilene. We’ve done other stuff across Texas-

Ben Gilbert: That happens to work well in AI in a way that it wouldn’t have worked well in the previous internet era, right?

Chase Lochmiller: Yeah, that’s right. I think certainly for these mega clusters, latency, it’s not super sensitive from a latency perspective. Certainly, for training, if you’re training a new model or fine-tuning something or post-training, you’re far less concerned about adding 10, 20, 30 milliseconds of latency to get to the data center. And then even for most inference applications, frankly, you don’t really care about that latency.

So, one of the beauties of AI from an infrastructure standpoint is it’s far more agnostic as far as where the workload’s actually running. There’s certainly applications that you want to be very low latency. I don’t think anybody wants their self-driving car to be running on a cloud data center network-

Ben Gilbert: Right. But if I’m going to go train GPT-6 for six months…

Chase Lochmiller: Exactly.

Ben Gilbert: I don’t care where it trains.

Chase Lochmiller: You don’t care. And even for all these chain-of-thought reasoning models, where you’re doing test-time compute scaling, and the model’s really thinking about the answer before it gets back to you. The response time is order of seconds, minutes, days, weeks. It could be a very, very long period of time. Adding 30 milliseconds, it’s irrelevant. It just doesn’t matter at all.

That’s set up this entirely new framework for how we think about the compute infrastructure to support AI in the future. And it’s very much in line with Crusoe’s philosophy of energy-first. There’s going to be large AI factories built in areas where we can access abundant energy resources. Abilene’s a great initial application of this.

We’ve announced a project that we’re doing in Wyoming that has initially 1.8 gigawatts of power. It will scale to 10 gigawatts of power. So, again, this is two New York Cities or something like that. It’s a ton of power, and I think there’s going to be many of these sorts of facilities that are in these naturally energy-rich areas.

Ben Gilbert: I’d love to talk about your entrepreneurial journey a little bit because I think everyone in the room who doesn’t know much about Crusoe is probably trying to figure out how you, a seven-year-old company, are building what I think is currently operating the world’s largest and most power-intensive AI data center. Is that fair to say in the current world?

Chase Lochmiller: I think that’s right. I don’t want to go on record, and Elon get mad at me or something, but…

Ben Gilbert: A very large-

Chase Lochmiller: It’s big.

Ben Gilbert: Before you were doing this data center business that you have, you were doing Crusoe Cloud, and you currently do both of those. You had several other iterations of the business, too. Can you take us to… If I’m a founder sitting in the room starting a business, what unknown steps may be on the journey ahead of me?

Chase Lochmiller: Yeah. So, my background — I was working in AI research for the financial industry, where I was actually a quant portfolio manager for about a decade, building AI and machine learning models to forecast security returns. We were historically using a lot of classical machine learning techniques, and then doing a lot of feature engineering on these economic relationships between stocks or commodities or whatnot. And then we’d sort of train these models.

And then deep learning sort of appeared, and AlexNet was published, and we shifted a lot of our workloads from classical machine learning with heavily engineered features to deep learning models, where the model was actually discovering the feature for us. And that shift led us to shifting from training on CPUs to training on GPUs and consuming a lot of computing power. And I sort of had this firm belief at that point that intelligence was going to be embedded in every aspect of the economy. Everything could be made better by having Silicon-based intelligence, making it more optimized, more accelerated, more improved, compared to just humans doing it.

So, when I set out to build Crusoe, my big ambition and goal was to really build this AI cloud platform, and I recognized early on how important energy was and the scaling of that. And so when we first started, we were actually capturing this waste methane that was being flared in the oil field, and it was basically a free energy resource that was being wasted by another industry, and we were capturing it, we’re utilizing it to generate our own power to power these mobile and modular data centers. Our first off-take to monetize it was actually Bitcoin money.

Ben Gilbert: It’s the craziest thing. You had these effectively shipping containers that they would fill with GPUs, and you’d drop it in an oil field and power it with flared methane.

Chase Lochmiller: Yeah, it was like…

Ben Gilbert: But you learn a lot doing it.

Chase Lochmiller: But I learned a lot. Exactly. My co-founder came from this very energy background, and I’d never even been to an oil field, and I was like, “Wow, they’re just burning this stuff all day, every day. It’s just being lit on fire.” And so we sort of found this unique and interesting opportunity to basically turn otherwise wasted, stranded energy into money via computing.

And I think that energy-first mentality has continued to persist at Crusoe. I think what led us to getting into the data center design engineering and development business was that as we were growing and building and scaling Crusoe Cloud, I spent a lot of time with the Nvidia team going through the roadmaps, looking at the architecture of chips, looking at B100s, A100s, and looking at future generations of chips saying like, Okay, this generation’s 200 watts per chip. The next one’s 300 watts per chip. Okay, then you think you’re going to do something that’s like 600 or 700 watts per chip, like, wow. This is really accelerating in terms of the power density being consumed by these chips, and it got more competitive.

And a similar thing had actually played out in the Bitcoin space, where people were initially mining Bitcoin on their CPUs when Bitcoin was like this new network, you could do mining on your laptop, you can… 50 Bitcoin from mining a block on your laptop. And then people shifted to GPUs, and it got more competitive, and difficulty went up. And then people were actually doing this in tier four data centers that have 99.999% reliability.

Turns out you don’t need that for mining Bitcoin. All of the added CapEx to build a “five nines” of reliability data center was completely unnecessary for Bitcoin. And so people started building ASICs, Application-Specific Integrated Circuits just to do this SHA-256D hashing function, and they started moving them into these ultra low-cost data center infrastructure solutions where a tier five data center or tier four data center may cost you $10, $15 million a megawatt, a Bitcoin data center that’s essentially like a chicken coop or it’s like a power plug. That’s where it’s passive air-cooled. You don’t care about cleanliness. You try to get that down to $200,000 per megawatt, maybe even cheaper. So, you’re talking about a massive 98%, 99% reduction in overall CapEx per megawatt for this specific use case.

Well, I sort of felt like that thing was going to happen in AI as well, just watching the chips evolve. Cooling architectures are different. Reliability concerns are different. We’re seeing this massive shift in the industry away from five nines of reliability. Because you don’t need it, right? It’s like if you can have some amount of reliability, call it three nines of reliability, that’s plenty to support a training workload. It’s probably somewhere in between where Bitcoin is and where serving webpages is. But it’s definitely a new and unique application that requires new and unique infrastructure to support it, ranging from everything from the building design, the cooling design, the mechanical, electrical, the whole thing. It’s just fundamentally a new thing.

And I think the more people get in their heads that this is a new thing that requires new infrastructure, I think that’s the massive opportunity that we’re focused on going after.

Ben Gilbert: Awesome. Well, I always like talking to Chase because, to me, AI is software. It’s me interacting with Chat, me interacting with Cloud Code, me interacting with image generation, video generation, and my head never goes to, “AI is a completely rethought building with 7,000 people and moving dirt and drawing on new power sources.” And my eyes are always open when I talk to you with just how much physicality there is to AI.

Chase Lochmiller: Yeah, I mean, it’s a cool thing for people to be exposed to, but our goal with our cloud platform is really to abstract away all of that complexity. It’s all behind the scenes. So, our goal is to build these factories that can convert electrons into tokens, and so people can just interact directly with our core services, whether it’s managed high-performance virtualization of GPUs, managed Kubernetes, whatever core primitives you’re looking for, managed infrastructure, we deliver to you, and you’re abstracted away from all of this complexity.

If you want to host a model and run a high-performance managed inference service, you can do that on Crusoe Cloud. So, there’s a lot of cool ways you can interact with us, and we’ve tried to create different solutions for different participants across the AI industry.

Ben Gilbert: Awesome. All right. Give it up for Chase.

 

Can We Trust AI? The Future of Verified Reasoning in High-Stakes Systems

 

Today we’re bringing you a special live episode of Founded & Funded hosted by Madrona Partner Jon Turow, featuring Carina Hong, Founder & CEO of Axiom, and Byron Cook, Vice President & Distinguished Scientist at AWS.

Carina is building foundation models trained on verified proofs rather than human-written reasoning. Byron leads the automated reasoning group that secures AWS’s massive infrastructure by teaching machines to reason about real-world systems.

In this conversation, Jon, Carina, and Byron explore what it really means to move from models that appear to reason to systems that can prove they are right. They dig into why verification becomes essential as AI moves into high-consequence domains, how formal methods and machine learning are converging for the first time, and what happens when reasoning shifts from a scarce human resource to something that can scale with machines.

Listen on Spotify, Apple, and Amazon | Watch on YouTube.

You can also read Jon’s takeaways for founders here.


This transcript was automatically generated and edited for clarity.

Jon: So help me orient, we’re in a world where we’ve been talking about reasoning machines all day. And actually the models that we have that are sentence completion, they seem to reason, they seem to think.

So why do we need to be talking about next generations of reasoning when we have pretty amazing models already. Why do we need to be doing advanced ambitious new moonshots?

Carina: Yeah, for sure. I think models are getting really good at pattern matching and matching specifically for human preference. They are trained on chain of thought data, which is made by expert, lines and lines of chain of thought to approach a specific problem.

However, there are certain areas where human preferences don’t cut it. Those areas require objective truth. Absolute right or wrong answers that you cannot get right or wrong because some humans prefer it more than others. And one example is mathematics. The other examples are engineering, coding, quantitative finance sector, those more quantitative scientific areas. And recent models fall short as that.

For example, if you ask a model to approach a really difficult high school level mathematical proof, it actually cannot generate the correct intermediate steps. And that’s a really sharp kind of contrast with some other domains that it achieves expert level performance.

And there’s also a question about scaling, which is currently, we see them doing okay, but as you try to heal climb in these domains, it gets harder and harder to make recent models continue to work because currently, the way to give them some signals about whether they’re getting it right or wrong is by 10,000, 100,000 large language models, agents voting yes or no.

And that gets extremely expensive as you try to scale up to harder and harder difficulties in those domains. So verification must come in, which is objective shoes, right or wrong. I’ll leave it at that.

Jon: Byron, what do you think?

Byron: So I’m coming in, I’ve been at Amazon for 10 years applying proof tools in AWS infrastructure and also for customers. And so I come at it from the other side where an incorrect answer is a crime, so you can’t give the wrong answer. So we apply mathematical logic to-

Jon: $100 billion of billing in a year. We got to bill it right.

Byron: Exactly, yeah. The cryptography needs to be correct. The virtualization needs to be correct. The storage needs to be correct. The interpretation of the policies needs to be correct. So that’s where I’m coming at it from.

But the sort of challenge for us has been that the tools that you use to achieve that 100% correctness were harder to use and required PhDs. So for me, what’s quite exciting about the recent trends is that suddenly these tools in combination with machine learning become actually potentially much easier to use.

Jon: Can you say a little more about that? One thing that’s amazing about this field, applying reasoning and formal verification in technology has been around for a very long time. I think even before you got started with it, which was some time ago, but it has been an artisan craft and it required experts to verify line by line that this airplane autopilot system is going to work as designed.

So we couldn’t use it everywhere because we had to use it only where we could find the resources to do it, justify the resources, find the people. So what changed, Byron?

Byron: First of all, as computers were networked and as society began to take more dependence on computers, the need for that work became more and more important. So as people move to the cloud, for example, proving the correctness of the cryptography became the reason the customers actually came to AWS, proving the correctness of the policy interpreter became the reason that people came to AWS. And so it helped AWS grow. It helped customers move and that became an area of importance.

And then we and AWS have been able to make the tools more and more usable, but also there’s been other advances like, for example, the rest type system is essentially separation logic combined with really good error messages. So there’s been effort there. And then more recently with machine learning tools, that’s really been a breakout. The other thing to keep in mind is that the models are often trained with data from proofs, either synthesized or discovered.

And so the models are actually often quite good at using, for example, the Lean theorem prover. And so that combination of the human and the AI tools, understanding Lean or Isabelle makes the tools much more approachable.

Jon: I want to put a fine point on this because it’s something that took me a while to get my head around, chasing bugs, chasing incorrectness in something like a billing system or a security system. For a long time, or even like an autopilot system, was a bit like being a goalie, where you’re just trying to stop every vector of something that could go wrong and trying to think of all the things that could possibly go wrong and knock them down one by one.

And you’re using a word proof, which is a big word, which says, “It’s not just everything that I’ve thought of that is going to break, but this will not break. This works and I can prove it.” Can you say a little more about that?

Byron: So you can identify properties that you want to hold in a program. So, for example, all data at rest must be encrypted or there’s no memory corruption errors or that strong consistency is actually guaranteed as opposed to eventual consistency.

And then up to that property, now you can prove using techniques like mathematical induction, which is the same technique we use to prove the Pythagorean theorem or the Four Color theorem. And so it’s the sort of foundation of how we know what’s true of the infinites, how we know there’s an infinite number of primes. It’s the foundation of truth in our society.

Jon: And so when we say proof, we mean it in a mathematical sense that yes, your AWS bill’s going to be right?

Byron: Yeah. What’s amazing now is with the world and society, the usability of generative AI means that people have discovered this problem too. And so they are trying to remove the sociotechnical mechanisms with agentic tools or chatbots and that notion of, “Oh, the answer is wrong. How do I know the answer is right? Can you give me an argument as to why the answer is right?” Is something that the world is waking up to the importance of that, whereas before it was quite an obscure area.

Jon: Yeah. So, Carina, there’s something we’ve spoken about a bunch of times, and I want to sort of share this with the group, that there is a market for reasoning on this earth in this world, and there is not enough really sophisticated reasoning to go around to solve the important problems that we face.

How in the world is that true? What does that mean? And what does it look like when we go from a world of scarcity to abundance in reasoning?

Carina: Yeah, 100%. Double touching the point Byron was making, which is that for a lot of the programs we currently have an output and a property, large language models are good at generating the output. They’re unable to prove that this output satisfied that property. So that critical link is missing, and that’s why proofs, in addition to computation, becomes extremely important.

Only when a model is trained on proofs and similarly executable programs, you can know for sure it’s right or wrong. The really technical breakthrough happened in formal language called Lean and similarly like Isabelle and there are a few variants of theorem-proving languages, they term mathematical proofs into computer programs by Curry Howard correspondence, which is a scientific term. So that means for every single math problem and the proof associated with it, I can turn it into a computer program, which I can then click it and wait for it to run.

After it finish running, I can see either a check mark, which tells me congratulations, it’s good, proof is good to go, or an error message, which I can then feed it back to the model to help it to revise and adjust for the correct answer next time. And then it keeps doing so. And eventually, I will get a check mark, which is, wow, congratulations. This proof is actually logically sound.

At Axiom, we are trying to solve the problem of the scarcity of outlier human reasoners. If you think about the history of scientific discovery, and I keep thinking back to some mathematicians lives and their legacy. Ramanujan died, I think at the age of 32 and before he died because of starvation and ill health, he had tons of notebooks. And those notebooks have formulas in it that took subsequent mathematicians, I think, decades, if not centuries, to prove.

And I also think about Galois, the guy who really pioneered group theory. He died incredibly young, I think 22 because of a duel for a girl that he loved. These are the kind of human outlier reasoners that the society lacks for true scientific breakthroughs accelerating to market applications to happen. If we think about Abacus giving rise to accounting, integral calculus giving rise to thermodynamics, mechanics, and then industrial revolution, we think about babbage engine, which is a mathematical tool to make lock tables faster, that’s the prototype of a computer. But there’s a significant time lag of 200, 300 years.

AI compresses everything. You can consider what happens if billions of AI mathematician agents go to every single complex system that is currently not understood and work with applied scientists, domain experts in this field in chips, for example, have a theorem prover, help formally verify chips to work with people who are the best software engineers, have a theorem prover, help formally verify code, work with people who really understand the economics. I mean, quant trading has previously been really restricted to the domain of high-frequency or some mid-frequency quantitative trading. I was previously working at XTX Markets, a leading hedge fund, but really why has anyone not thought about combining those really outlier quantitative methods to some more traditional, say, day trading? What happens if you have an AI Terry Tao, AI math wiz work with every user on Robinhood? What happens? And I think that’s the question that we need to ask ourselves about. When we enter the era of math intelligence and math outlier reasoner grounded by verification, because we couldn’t afford to have really these billions of agents that we don’t really know whether the answer is right or wrong.

Okay, I see a 1,000-line math proof and you tell me that I have no way to check it. Surely, I cannot have humans check it. That’s extremely expensive when it comes to scaling, but once you have these thousands of math proofs and you know that there is no chance it’s incorrect, it’s good for use, that’s incredibly impactful. What’s even better is, in an era of agents, I can make sure that the stuff I pass from agent A to agent B is correct. I can’t really afford to pass incorrect and lousy work to my teammate during my legal training at Stanford Law. But what happens if you have verification, if you have agentic workflow, if you have the ability to conjecture and generate new problems to think about and then have the verifier to try ground it and prove it, this sort of self-improving AI, recursive, self-improving loop, I think that is the amazing future that building a reasoning engine as a model layer helps unlock.

Jon: So let me try and make this really tangible. Take me back to your work at this quant fund, XTX, which hires people who are really, really good at math and basically hires every person who is good enough at math that they can find on earth. Not enough people that they can find.

And there are other fields that do this too, by the way, but let’s focus on one field and one firm. They’re in a world of scarcity of reasoning. What happens to a place like that when they can suddenly turn on more reasoning like a tap of water, like a tap of electricity?

Carina: I think the future in these players is quite interesting. There are different dynamics with each of these players. So imagine you are a firm that actually doesn’t have access to the top MIT Princeton math PhD students. You just do not have the talent pipeline to employ these extraordinarily bright young minds to be quantraders.

Suppose you are a hedge fund in a market that’s total trading volume is quite low, which means low-hanging fruit is everywhere, like really not so hard to derive alphas or everywhere. You just don’t have the talent to help unlock it. I think that’s really meaningful for these players, first of all.

Second of all, for these players that have already amazing talents, then they have these AI mathematicians to collaborate with. They don’t have the sort of problem where I think as some other hedge funds, which is, I have never have worked in, but I just read it from all sorts of books where different desks have different trading strategies.

For example, you can be quite isolated to develop one thing that you own full end to end. And you have these AI mathematicians that’s bound to be correct to bounce ideas over. And you can think of talking to the AI at your tea time generates really interesting quantitative methods with the AI collaborators.

So having this sort of strong supervision from the existing amazing talent outlier, human reasoner is going to only accelerate the discoveries and breakthroughs in the algorithmic, in the scientific sense.

Jon: Yeah. One thing that I think we’ve seen with the first waves of GenAI is that the marginal cost of creativity and experiments is dropping. We saw this in the cloud, Byron at AWS when the marginal cost of experimentation for a cloud developer dropped and I could run an experiment much faster with it ordering racks.

And now I can do that with all kinds of creative fields. And I think to your point, Carina, if I’m now doing it in fields where being right super matters, be it because I’m betting billions of dollars or people are going to get hurt if I’m wrong or any number of reasons, you start to get to that next level if you can move from a world of reasoning scarcity or reasoning abundance.

Carina: For sure. Jovan’s paradox tells us that when the price, when the cost of the tool becomes lower, there will be new market cases, use cases to be unlocked. And I think really we are in an era where Jon’s paradox is our opportunity. What happens if every cloud compute provider can afford, say, Facebook engineers that currently pay two millions each year to work on the routine Mundane IT optimization ticket task?

Just think about what that means. And what if a hedge fund in a trading volume of only say eight millions each day can now afford a trader, an AI trader that’s as good as those being paid currently 20 millions each year as a starting salary in certain quantum places as this does happen to afford these AI mathematicians, but only pay them $5 each hour.

Jon: So help me, I’m laughing with maybe some other people at the term only eight million. Help me understand how broad is the application of reasoning with math? Is this about creating more math for its own sake or there was huge generalization here and I think that’s counterintuitive at least it was for me until you guys explained it to me.

Byron: So there’s a little bit of survivor bias going on because I’m pulled into the conversations where it’s needed. But from what I see, like I’ve worked in reasoning about genetic regulatory pathways, railway switching, safety, microprocessor design, operating systems, virtualization, storage systems, networking.

So any question where there’s a system and the system is evolving, there are questions around termination, around reachability, and most of those turned into questions about unbounded, infinite or intractably large problems. I haven’t really talked more about Lean, but there’s also the propositional satisfiability solvers and SMT solvers, which handle combinatorial reasoning.So you’ve got the problems are typically MP complete or they’re undecideable, and these tools eat those problems up for lunch. So many, many problems are translatable to state space reachability or termination.

Jon: Okay. Yeah. So if I follow on this line and I have tools and techniques that are broadly generalizable to be really, really powerful, super intelligent brains that we can turn on like a tap. Each of you are leaders who have to straddle two lines because you’re doing advanced research to push the frontier.

And I know because each of you have told me that you’re also pushing really hard to apply this and actually put it in front of customers and make it land as opposed to stay back in the lab. And so maybe you can talk about the leadership experience of that and how you balance your resources and your team and even your culture about the relative importance of awesome math and code for its own sake versus the impact it’s going to have for customers.

Byron: So, Strachey, who is a contemporary of Turing and founded the Oxford Computer Science Department, says that the division between theory and practicing computing is injurious. That the theoreticians essentially don’t know what problems to work on and the practitioners don’t have a grounding in what they’re doing. So what I have found is that even when I was in Ivory Tower, Blue Skies research labs, that those who ground themselves and customer problems were the most successful because they were able to extract out theoretical challenges.

So for example, I worked on termination proving, driven by the observation that device drivers need to terminate and also genetic regulatory pathways need homeostasis, and that is the same problem. So what I’ve done with the scientists that work at Amazon and in my career is to help people see it’s an and not an or to work both in the theoretical and be applied.

The other thing I’ll say is that one of the big challenges for the space is figuring out what it is you want, what should be true. So I can systematically remove incorrect statements from a chatbot about the Family Medical Leave Act, but first I need to encode what are the rules of the Family Medical Leave Act and to do that has typically been a challenge. So you end up spending a lot of time identifying what is the correctness of the AWS IAM policy language, what are the semantics of AWS VPCs, what is the semantics of Rust precisely? And that actually turns out to be surprisingly hard to do.

And so to make practical progress, you have to drive on that and then sometimes make some compromises. One of the very nice things about the rise of generative AI is that these tools now can help you do that. So you can get them to help translate the natural language into logic and help you figure out what the logic means and figure out what it is you want to try and improve and then maintain that over time.

So you might maintain your formalization of the Family Medical Leave Act and then as you deploy your agentic system or your chatbot based on that formalization, over time you might realize you got it slightly wrong and fix the formalization.

Jon: Got it. Let me turn this upside down before I bring it to you, Carina. Byron, explain it not to your team who are going to want to know how much should we be doing awesome research versus applying it? How do you explain this to Andy and Matt Garmon about why in the world we should be investing this level because Amazon has been investing bigger and heavier and longer in formal methods and automated reasons than anybody. And how do you explain why that matters? Why that’s so important?

Byron: The customers did it for us. So I was in the Blue Skies research lab explaining things on a technical level in the past. And then when I joined Amazon, because I’m in New York City, I was pulled into conversations with financial services.

When they discovered the work we were doing, they went and drove that discussion. Actually, they went to other customers and helped them appreciate the importance of the work. And they drove those discussions with Andy and so on.

Jon: Because they said,

Byron: This is why we’re going to move to AWS.

Jon: This is why we’re going to move to AWS?

Byron: Yeah. Orders of magnitude workload because now we understand that the cryptography has been proved correct. We understand the story around the virtualization, we understand the durability story, and this is why we’re now comfortable moving our workloads over, and this is something we would never be able to do ourselves.

Jon: Because previously it was a list of things that might go wrong crossed off the list. And no, Byron comes with a proof that says, no, seriously we mean it.

Byron: Yeah. And you can also, it’s a transparent way of describing to your auditors or to countries why certain things are impossible. You can provide a proof and mathematical logic without showing the data why data can’t flow to certain places. And that’s important to certain governments and customers.

Jon: Got it. So, Carina, if I bring it to you, you’ve assembled a team of mathematicians and developers and researchers who like things that are awesome and they want to do advanced cutting edge work and push the boundaries of reasoning. How do you set the point with them that we’re going to apply this? And how do you set that balance and how do you calibrate what inputs they’re getting?

Carina: For sure. I think this is actually a contrarian panel in a way. The call to action of formal methods applied to reasoning is not something I think that has been talked about and appreciated enough. And I have this challenges when I’m talking to investors, when I’m talking to talents, when I’m talking to my mom. So what exactly are we doing? And so we launched yesterday on morning Axiom came out of Stealth.

There’s a Forbes article, can read about it. We’re building a self-improving, super-intelligent reasoner, starting with an AI mathematician. So all these terms are contrarians like VCs don’t know why they should invest in an AI mathematician. People don’t know what does self-improving AI mean besides that being in the first paragraph of Duckerberg’s personal super intelligence memo. And what is super intelligence anyways? I think we know for sure this technology is likely to work more likely than not.

And this is the reason why Axiom has been a hiring magnet. I think we are able to attract talent at the caliber of Facebook, AI’s directors and other companies, senior tactic managers level, because they believe in the mission. They believe in working for a company that is developing the next step function change of technology.

That is, if you compound formal verification like we talked about with better and better learning and search, then you really have a decent shot toward super intelligence that is defined as a model that can generate new problems for itself to work on with 100% correctness and then generate more interesting, harder problems based on how it did last time. So I think the technical vision being sound and that really helped with hiring.

And in terms of say justifying to investors, customers, I think people are seeing three things coming together right now as consensus in the AI field. One is the reasoning models on the informal reasoning side have gotten stronger and stronger as a good backbone of building things on top of it.

Second is the formal verification tools like Leanford have really been widely accessible and so many Lean developers in the world are trying to build tools within Lean or trying to expand the mathematical library and the CS library in Lean. And that has really become a phenomenon from a subculture. Back then, I think two years ago, people who are Lean developers talk to each other.

It’s almost like, “Wow, I can’t believe you are also doing this. This is a language that no one has heard of.” But now it’s while we have a vibrant community of Lean users. So that’s the second part. And the third part being that co-generation have got extremely, extremely strong. Reinforcement learning applied to a verifiable domain is going to have performance gain that has shown encoding.

And now by turning math into programs the first time, we can see reinforcement learning being good at math, logic and reasoning. So these three trends powerfully come together, which is also the hiring thesis at Axiom, which is we need experts from three fields, from applied AI deep learning, from mathematics, and from programming languages.

Really, the combination of these three fields and math obviously include both mathematicians and people who are Lean experts who usually are mathematicians. So these three pillars is what it takes to build an AI mathematician as the first step to a super intelligent reasoning engine.

Jon: So we’ve got time for I think one question from the group. Who would like to hear from Carina and Byron? Yes, sir.

Question: It’s fascinating to hear from you. So, Carina and Byron, so with your outlier intelligence or super intelligence, what problem domains are you most excited about going into or solving problems? So we’ll be good to hear some examples of that.

Byron: I’ll characterize the solutions where the space of satisfying assignments is very sparse. I appreciate that’s probably a mathematician’s way of describing the space, but when you have in biological systems and physics and in programs, security, there are many, many programs that are incorrect and buggy and insecure and very few that are secure.

And so to balance the different dimensions of privacy, sovereignty, security, availability, durability, and to meet all of those and to write down those specifications and then to find the programs that meet the specifications. And especially in the enterprise, that becomes increasingly, increasingly sparse. And so I think the generative AI combined with the formal reasoning, neurosymbolic AI, I think that is really great. An example which would not be great is poetry.

So I don’t think that formal verification is going to be especially good at proving the correctness of a song or a poem or something like that. So there’s a spectrum for sure, but those areas where very large amounts of money is on the line, people’s safety’s on the line, security, privacy, sovereignty, and where the consequences of getting it wrong are pretty bad is I think the area where the methods are a best fit.

Carina: Yes. I’m most excited about, it’s really two parts when it comes to go-to-market, which is scaling formal reasoning in all existing market, like really looking around the corners because we’re talking about things that previously, people just accept that the verification of a protocol could take three years. And now you’re telling me, wow, it takes two weeks. What happens then?

So I think really scaling it up by looking around the corner and really that depends on collaboration between various industry experts. And we really appreciate customer inbound, which is, “Hey, you have this problem that you tried to solve. And I think maybe formal reasoning is helpful. I don’t really know, Carina.” And then we can talk. And these conversations starting are really just a generally promising signs that perhaps we can find a corner that’s immensely valuable.

And then the other part is once this math intelligence era is unlocked, what does it mean for unlocking new markets and use cases? That part requires some part of dreaming about the future. I read a lot of sci-fi novels and actually I was asking sci-fi authors this question is, what does it mean? And no one really could quite figure out exactly what that is going to mean, but that part is definitely something I’m extremely exciting about.

On the technology side, I want to emphasize that technology is extremely hard. We have a data scarcity problem. If we think about how many tokens of Python code there is on the internet, it’s probably about one trillion. I think that’s a safe estimate. And how many tokens of Lean code, the programming language for proofs, probably 10 million. So that’s a 100,000 times data gap. And we know that scaling works really well in AI.

So when your data is scarce, that generally means the technology is going to have a co-start problem, which is why we are taking bold data bets by various methods of automatedly not relying on human experts to write the link code, but to automatically synthetically generate high quality, extremely interesting in terms of helping model training data. I will highlight that as one pretty difficult technological problem.

Jon: Cool. I think we can keep this going for a long time. I’m going to be thinking about what is the most true and correct song that I know. We can leave that-

Carina: The song of pie perhaps?

Jon: Good, good. So anyway, that’s the discussion for later. But thank you so much to Byron Cook and Carina Hong. It’s been a fascinating conversations. Thanks everybody.

Microsoft’s Agent Factory: The Future of AI Software with EVP of Core AI Jay Parikh

 

Today we’re bringing you a special live episode of Founded & Funded, featuring Madrona Managing Director Soma Somasegar and Jay Parikh, Executive Vice President of Core AI at Microsoft.

Jay leads the team responsible for Microsoft’s core AI stack, the systems that power Copilot, the tools developers rely on, like GitHub, and the infrastructure that makes large-scale AI possible. In short, his group builds the underlying tech that Microsoft and thousands of companies use to create AI-powered applications and agents.

In this conversation, Soma and Jay dive into what Jay calls the Agent Factory — a new paradigm reshaping how software gets built in the reasoning era. They explore how AI changes the development lifecycle, why observability and evals are becoming mission-critical for enterprises, what it means to collapse traditional engineering functions, and how organizations should prepare for a world where models, agents, and human builders all collaborate in real time.

Listen on Spotify, Apple, and Amazon | Watch on YouTube.

You can also read Soma’s takeaways for founders here.


This transcript was automatically generated and edited for clarity.

Soma: Thank you for being here today, Jay.

Jay: It’s a pleasure. It’s a pleasure.

Soma: It’s really a pleasure to have you here. We’ve been sort of spending the whole day here talking about how we are in the beginning or the early days of a reasoning revolution and we sort of think about the industrial revolution and fast-forward to today, to me the big difference between when we say Industrial Revolution and when we say reasoning revolution is the machines that were part of the industrial revolution where right from day one, I should say, they were a depreciating asset, whereas reasoning machines, I think from day one, it’s an appreciating asset. And it’s sort of the flywheel that happens with data that makes it more powerful and appreciating in value. And so I look at the opportunities ahead, it is phenomenal.

As I was at Microsoft for many years before coming over to the side of the house, one of the things that I remember vividly is what Bill used to talk about in the early days, in the early days of Microsoft, it was like, “Hey, we want to build a software factory.” And then I heard you talk about this a couple different times about, “Hey, this is now the time for us to go build an agent factory. If you think about it in today’s day and world with agentic AI, can you tell us a little bit about what you mean by Agent Factory and how that’s going to change the world that we live in?

Jay: Yeah, so fun story, I didn’t actually know that Bill used to talk about Microsoft as the software factory. So my one-on-one with Bill, my first one was I was telling him about this concept of the Agent Factory, which I’ll answer your question, he’s like, and he stopped me. He’s like, “Jay, you know when we started a company back in 1975, I always had this,” and then he went on to explain the software factory thing. And he was very exuberant and excited, but I was like, “Wow, 50 years later, here we are, and we’re talking about the Agent Factory. So I had no idea about the software factory, so I kind of sketched this out in a completely different kind of time. And it was just fun to hear that original inspiration and vision for the company.

So the idea behind the Agent Factory, I think originally I thought about this internally in terms of how we need to evolve our infrastructure, our platform, our tools, and effectively how we make products in the next era of Microsoft. And then I realized, as we were cooking this and thinking through, it’s a cultural thing, it’s a technology thing, it’s an incentive thing, it’s a systems thing, that I actually realized that every company that isn’t maybe founded today is going to have to actually create a similar system.

And the idea is, I mean, honestly, I drew it out, and it’s literally like, I did it in ASCII art for the team in terms of how this looks. And there’s these things that I envisioned showing up on the loading dock, and they’re like models, they’re different things, there’s technology from MSR that shows up here. And then we kind of need to put those together into a production line. Then at the end it produces some AI capability, which today likely is some type of agent, whether it be a coding agent, a code review agent, a biology agent, a fraud, something, something agent. And we’re building those agents internally for our core products, but as more and more customers adopt different parts of our tools, our products or platforms, they are enabling their business or transforming or changing their businesses to be building their own agents to transform their organizations, their business.

Some of it is about time savings, but I always push customers to really raise their level of ambition in this reasoning era, to use your talk track here, to raise their level of ambition of the types of things they’re building because we’re on an exponential and we’re all terrible about understanding what’s actually possible with the collective set of intelligence that we have in all of these models today.

Soma: As Microsoft, as a company that is sort of in the forefront of AI in a variety of ways, how do you think this concept of Agent Factory is going to change how Microsoft operates and how Microsoft is going to be building software?

Jay: I mean, the way I think about it is what we’re doing in our team with several others is building the future of Microsoft. I think it is going to change everything we do.

And yes, we are a 50-year-old company and there’s lots of technology that’s there classic and is going to be running for a long time and we’ll continue to optimize and support those products, but everything that you see going forward from the company, whether it be knowledge worker in M365 products to security, to platform, to tools, to the health and life sciences, to the focus of research, to just keep going, I think is all going to be around this concept of really unlocking, what can you do to drive productivity, drive creativity, drive collaboration, solve these technical or science problems that have been eluding us for decades faster now because we have these frameworks, we have these tools, we have these models, we have the ability to get data, synthesize data, do this flywheel of reinforcement learning, and I think we’re still early in terms of how that is going to manifest itself in both Microsoft’s products, but then how do we enable those capabilities in our platform, our tools, our data, our systems, so that every organization out there can use these to really unlock that in their own enterprises.

Now, the fact is, there is inertia. In organizations, it takes them time to adopt this stuff. And so it’s not happening overnight. The technology is way ahead in terms of capability than most cultural transformations will take in an enterprise. So there is that, we don’t have time, and everything is moving fast, but these organizations, the larger enterprise organizations, global 2000, they are slower to adopt all of this and change.

Soma: Great. All kinds of numbers get bandied around that, Jay. The most recent thing I heard that was inspirational, and it feels like, as much as I want to see that happen today, it’s probably a little bit in the future, was where Dario at Anthropics said, “Hey, by the end of this year, 90% of all code is going to be generated using AI.” So if you have to pick a number for where Microsoft is in the journey, just internally from a development perspective, how would you characterize that?

Jay: Yeah, so I’m going to give you a non-answer, but in answer, and the way I think about this actually is these numbers around how much code is written by AI is largely uninteresting to me having run and scaled some of the largest engineering teams because I really think it’s more about the capabilities it’s creating or giving to builders in the company. So to me it’s less about lines of code that AI is generating, it’s really a few things.

One is of our run the business stuff, of our technical debt, of our operational stuff, things that we’re optimizing or we need to, technical debt, upgrading frameworks of stuff, fixing security, improving performance, there’s a lot of toil that engineers spend in organizations dealing with that. And I can say, “Hey, we have X percent of our code being generated by AI,” but who cares because I still have engineers toiling away at all of this tech debt and this other stuff and bug submissions from customers.

So what I want is to shift that to, hey, this run-the-business stuff, this technical debt stuff, this bug fixing, this… In some ways, there’s true toil there, but there’s stuff there that is just holding these very talented people back from achieving higher levels of creativity and collaboration. I want to shift that dramatically and I want to shift the amount of time that we spend, classic engineering time, whether it be in meetings or upgrading stuff, fixing vulnerabilities, pushing stuff to production, optimizing some performance bug, I want to squeeze that with AI, I want to get that down so then we open up and we give back time for the creativity part.

So we actually have an initiative inside our team that we run, and we’re scaling across the company called Engineering Thrive. And we actually do measure the time when engineers are stuck. So largely there’s unfocused time, there’s focused time, and some other categorizations, and we’re watching longitudinally as we make cultural shifts, as we improve the tools, the technology, and even just the know-how, the skilling of our technical organization, how that percentage is, how they shift. Because, ultimately, if I can get more builder time back, more creative time, or focus time, then we’ll just accelerate our ability to build great products and to bring more value to customers and all of that good stuff.

Soma: That’s great.

Jay: There’s one other thing I would just say from a macro perspective, I would say we’re at a point today where if you think about it, if you add up all the software that’s ever been written in humankind, we are probably sub-one percent of what’s going to be written in the next 10 years. And I think that’s the more interesting thing, of how much more we’re going to be able to do, solve, create, collaborate, because now we have this superpower that’s getting better and better every day in terms of being able to build, prototype, solve these types of problems.

Soma: I should tell you this little anecdote thing because it involves GitHub Copilot. This was probably a year ago, or so, I was talking to Marco Argenti, who was here earlier today, the CIO of Goldman Sachs.

Jay: I know Marco.

Soma: And he was telling me that, “Hey,” because he had deployed GitHub Copilot across his organization, I asked him, “So what are you seeing in terms of productivity benefits?” And this is within six months of him having deployed GitHub Copilot, and he said, “Hey, I can, without thinking too much, I can tell you that my workforce is 20-25% more productive.” And I asked him, “So that means you’re going to fire 25% of your organization tomorrow?” He said, “No, no, no, we don’t think that way. The interesting thing is that it has made my job easier in that I go into a meeting now knowing that I’m not waiting for a shoe to drop at the end, where people come and tell me exciting things and then say, ‘Here is a headcount bill.’

Now, instead, it’s like, hey, that person knows there is headcount available, meaning resources available in the team, because of the productivity benefits, I know that, so there is no headcount corner. It’s all about what do we want to prioritize? What do we want to do? And so to me, I think that’s an exciting part of what AI is able to do for all of us.”

Jay: I agree with Marco on that.

Soma: Microsoft is one of the companies, Jay, that is spending on AI infrastructure as much as anybody else, I would say, if not, maybe more.

Jay: Maybe more.

Soma: This year, the stated dollar amount from Microsoft was $80 billion or whatever. But even apart from the dollar amount, the Azure platform, the Azure AI platform, the Azure AI Foundry, all the tools, all the developer stuff that you guys do, you have a fantastic platform for AI builders or AI developers. But having said that, and having made all the progress that you’ve made, what are the two or three things that you think Microsoft ought to be doing more, whether it is from an operational perspective or a technology/product perspective or even a cultural perspective, to position Microsoft in the best possible place for an agentic AI world?

Jay: The thing that I would say is with our team and our focus here, and some of you were, I think, supported many of these teams when you were at Microsoft, so you know them well, is really putting together this full-stack approach to how the future of software development is going to look. And it’s starting at the core, at the infrastructure handling and evolving how the workloads are going to change and scale. Then it’s the platform in terms of these AI applications, the agents, remember for decades we’ve been building software where we go interview our customers, we come back, we design some schema, we write some business logic crap in front of it, and then we put some UI around it. And that’s what we’ve done in terms of software. And we’ve built a lot of incredible stuff over the decades.

Now though the entire, it’s not even that the paradigm goes on its head, it’s like a completely different alternate universe because now you have these models that can think, they can plan, they can reason, they can call different tools, they can collaborate with each other, they can do things faster, better, sometimes slower and worse as well. And then you have to think about the scaffolding that sits on top, below left and right of these models, from a, what is software, and, what is this new era of applications and software and agents, whatever comes after agents is going to look like. So it really is actually a completely different stack, and that’s why we put this team together, it’s really, first principles, think about the [inaudible 00:16:35]], the platform, the tools, security and trust and everything.

Now to answer your question, it’s really important, especially in the tools and the platform part of this where… and my philosophy is that we have to do this… Sure, we’re going to build a glorious platform and it’s going to be delightful, but I actually believe for builders, for developers, for scientists, whoever is building on this platform, you also want choice. So having a vibrant ecosystem of different partners is incredibly important to our strategy.

Now, the things where I think the world and we need help with is, some of this was covered, I’m sure today is, I’m getting into a very specific example here, I’ve been saying this for a while, but I think it’s finally starting to be more of the conversation, I think evals are going to be more and more invoked in terms of what everybody needs to go solve. And I think there’s some clever startups out there tackling this. We’re all trying to tackle it. Everybody’s struggling with this. Marco is struggling with it. I struggle with it. Because there are benchmarks and there’s, what I say is one-dimensional evals, and we’re pretty good at that.

But humans like the lived experience, they’re so personal, they’re so different. And then you can have a set of evals, build your application, evaluate some models, be like, “Hey, we got these scores,” et cetera, et cetera. Then you put it in the hands of your customers and they vomit all over the experience and you’re like, “Wait, my evals are great, look at my scores.” And they’re like, “Yeah, this sucks. I don’t want to use this thing.” So that there’s a human touch or a three-dimension, four dimensionality to the lived experience of these products, and I don’t think we’re still really, really, I think basic and not great there.

The other area that I think is, and it comes up in every conversation when I talk to the enterprises, nothing in AI is going to work in the enterprise without observability, and I know there’s some great startups out there doing cool stuff out there. We invest in this a ton, there’s a bunch of stuff that we’ll be announcing through the end of the year on observability orchestration, but I think this is also one where we want, and I expect a vibrant ecosystem of startups out there to be building these really sophisticated and awesome observability features, products, whatever it might be, because that’s going to help really drive the diffusion of this technology in the enterprises.

Soma: When you talked about-

Jay: I have a long list today. Is this my shopping list?

Soma: When you talked about a vibrant ecosystem, one of the things that I think the world knows this now, that Microsoft gets a lot of credit for the partnership that you guys struck with OpenAI a handful of years ago. I think it has served Microsoft really well. It has served OpenAI really well. It has served the world really well.

But having said that, I would say in the last few months, at least in the last few months, maybe even a little longer than that, we’ve started seeing Microsoft come out and say, “Hey, for this part of Microsoft,” meaning say for office Copilot or for some other thing, “Hey, we might take advantage of different models, whether it’s Anthropic or what have you,” which is fantastic. And Microsoft was one of the first companies that came and said, “Hey, we really want to build a platform that delivers model as a service,” which is great because it’s all about a vibrant ecosystem and embracing the ecosystem. But having said all this, what do you think Microsoft should be doing? Maybe it’s already doing some of it in terms of having investments and building first-party large language or frontier models. Is that important or not important?

Jay: So there are three things. One is our partnership, and I think history with OpenAI is still incredible for both companies and we invest both sides a lot in lots of different things, because the world is changing, the products are changing, the risks out there are changing, the business models are changing. So there’s a lot of collaboration across every level of both companies, and that continues to be a huge focus, and I think we’re both very happy and there’s lots of stuff to do, and we can’t get to everything, so I think it’s, one, it’s a very valuable relationship. It’s a very, I think, positive relationship.

The second thing that I would say, and you kind of gave the answer here, which is building these platforms, building these tools, it’s really important that we also meet developers and builders and knowledge workers where they are and what they want. So being dogmatic about, “Hey, we’re just going to have only one choice, or one model,” that’s not what people want. And so our strategy I think has always been, but I think it is way more noticeable now where there is that choice. I mean, GitHub Copilot has always offered the best models, irrespective of who the provider is. We offer Gemini, Grok, OpenAI, Anthropic, our own models, etc., in the model picker. And then in Foundry, we have thousands of models that somebody can pick up and use to build what they want, open source, closed source, it’s all there. And we’ll continue to invest in that breadth of choice.

And then I’d say the third part of the strategy is our own models. And we released, I think probably three or four weeks ago, our MAI models. There’s an initial set of models that have been built, trained all inside of Microsoft that we’ve released. There is more coming along that frontier, investment there. We have lots of other models that have come out of Microsoft research. There’s models actually that are Microsoft-trained models that sit inside of GitHub Copilot that most people don’t know today. And then there’s stuff that’s in M365 that’s also not one of the companies we’ve already talked about.

So there is all of that going on. And it is hard, honestly, it’s sometimes confusing that there is so much choice, but we’re all figuring this out and we do have to listen to our customers in terms of what they want, and we have to provide the best because everybody wants to tune and I think optimize for different things, whether it be quality or cost or performance or this thing that works in a specific region or in a certain language or a certain discipline better than another model.

Soma: Why don’t we turn it around and see if there are any questions from the audience for you, Jay?

Question: So I really just, as a point of clarity, when you say observability, is that also auditability, or are they merely observable in transaction, but don’t necessarily lead to an audit trail?

Jay: I’d say observability in the royal observability, I think it’s a, depending on the enterprise, observability is the bigger category. Inside of it, you can say there is monitoring stuff, there’s traceability stuff, there’s compliance, there’s audit, there’s even some security tie-in right there. So it really is I think, an expansive term and I think all of that falls within the comment I made earlier of where enterprises need help and they are evolving, but this is an area where whether it be the monitoring platform for monitoring usage and cost and performance and that stuff to having something that will do the audit so that you can hand this stuff over to compliance folks, et cetera.

So they’re somewhat obviously different functions in a company, but from an observability perspective, you need that data, you need that correctness, you need that evidence. And as the regulatory environment keeps changing on the compliance and audit side too, there is going to be a whole set of shifts that are happening there, but the data that we get has to be made into these different outcomes. But I think we’re really early on all of that, super, super early.

Soma: Great. I think there’s a question here.

Question: I’m curious to hear your thoughts on, you talked about the model picker. So with GitHub, with Cursor, with a lot of different tools today, there’s lots of different model options for the end user. Do you think the fascination with choosing your own model dies down as AI becomes more commonplace? I think of this as the electricity example. Edison built the electric bulb, everybody wanted to know how it worked. Today, when you flip on your light bulb, you don’t even think about it. I’m curious to hear your thoughts, do you think as models or will models become commoditized to the point where the end user doesn’t have to have the burden of thinking about what is the best model to use for what task and applying that also to the developers?

Jay: I think there will always be a need for developers to kind of go into manual mode and pick something to solve a particular task or something that may appeal to their sense of craft, so I think that’s always going to be there at least for a long time. I don’t think it’s just going to be automatic or hidden for everybody. So I think we have to make that option there, or there’s some personalization aspect to it.

I’d say the cognitive overhead that we’ve put on developers today in terms of picking models, testing them, trying them, I don’t think it’s as necessary in the long term, especially if you think, “Hey, we have 10 models in the model picker, but imagine a world where there’s 50.” And nobody’s going to know one from the other at that point, so I tend to agree with you on that.

But where the product then needs to evolve is how do we deduce and learn what that individual’s preferences, style that lived experience is and go into auto mode and that we have that long-running context, that memory, and then we can more, I think, abstract away what these different models do, but it’s going to be hyper-personalized to that repo, to that developer or to that team. And this is where part of the vision that we talk more about, but just think about, okay, hey, there’s end users, but there’s the repo in GitHub, and that actually has a ton of context today. Issues, discussions, the security stuff, documentation, et cetera, team dynamics, et cetera.

Then you have what’s happening to me or you as an end developer and how you are evolving, how you’re working, that memory, that context that you carry with you, and how do you put those two together and then really make a lot of this much more streamlined, but I think it’ll move even beyond just picking the right model for you. I think it’s going to do a lot more than just auto-picking models.

Soma: Why don’t we wrap it up with one last question?

Question: If you had to fast-forward five years into the future, what does an enterprise engineering team look like?

Jay: Five years in the future? That seems like a really long time. For me, I think there are a few things. So one is I think in two years the companies that have figured this out will look entirely different. I think functions will collapse. I think applications will collapse, and I mean merge, not collapse, and go to zero, but I do think functions will collapse, and I think that’s going to be a big reckoning, so to speak. But I think companies that embrace and understand how to flatten structures, how to combine functions, and really use this as a way to pick up the pace, and also prioritizing and understanding how to shift that time from this run the business time to this creative time, those are going to be the ones that start to really run lead. I think there’s going to be a lot of enterprise teams that look like they look today, but where those companies are going, probably nowhere.

I’d say the other thing that’s going to happen here is I think the enterprises are going to have to work through the system because there’s just a lot of, today, in some ways, we talk about waterfall, and we’re like, “No, we’re agile,” but honestly, humans still work in a waterfall way. We work on a project, then we move to the next project, then you have to go through some set of reviews, somebody has to do 45 approvals before it gets into production at a big bank. Those systems, those things have to be automated, those functions inside of the company have to be, I think… We have to rethink them from the ground up. So I’m talking about the companies that are going to figure it out, the the companies that are still going to try to operate the way they are today, I think are going to just see the, “Hey, I’ve got 15% productivity gain,” but the rest of your competition is away in a speedboat and you’re like, “Hey, I’m paddling faster.”

Soma: Great. Okay, thank you so much, Jay, for this wonderful conversation.

Jay: Thank you.

Soma: Thank you.

How — and What — to Build in the Age of OpenAI

 

In this special live episode of Founded & Funded, Madrona Partner Vivek Ramaswami sits down with Jason Kwon, Chief Strategy Officer at OpenAI — a 2025 IA40 winner — just days after the company’s Sora and agentic commerce announcements.

They dive into OpenAI’s thoughts on which markets they enter directly and which ones they support as a platform, their perspective on AGI, and their overall perspective on where the space is headed, from compute and data to reasoning and agents.

Anyone building in AI will benefit from this candid and in-depth discussion recorded during our 2025 IA Summit.

Listen on Spotify, Apple, and Amazon | Watch on YouTube.

You can also read Vivek’s takeaways for founders here.


This transcript was automatically generated and edited for clarity.

Vivek: Both of us came up from the Bay Area, me two days ago and Jason this morning, so we’re glad to have him here. Jason’s been part of the OpenAI leadership for almost five years now. He’s seen a lot there. There’s no shortage of incredible and interesting things coming out of OpenAI. Let’s go ahead and get started with this: What does being chief strategy officer at arguably the most strategic company in AI and tech mean on a day-to-day basis?

Jason: Yeah. The way I’m mostly thinking about my job is thinking about the external and bringing it in inside the company, so nominally, the functions I work with a lot are: legal policy, global affairs, parts of trust and safety. The legal part also brings me a lot into deal making, but really it’s the external world that’s reacting to the technology, and there’s a whole process of actually taking in these sets of considerations, especially when the technology is very disruptive, and thinking about how that should impact things like how we approach ecosystems through deal making or deal structures, how we should think about various policy measures and how we should think about building solutions, working also with startups as well as big companies to address various concerns, and so I think about that, really, on a day-to-day basis.

Vivek: One of the interesting things we’ve been talking about is that OpenAI is best known for frontier models, everything from when GPT-3 was exploding into our world and then GPT-4 and beyond. Obviously, the ecosystem around those models is really exploding, and we often have this term: full stack of AI. What does a full stack of AI look like to you? How do you all think about that in OpenAI lens?

Jason Kwon on AGI, agents & how startups can build in the OpenAI era. Recorded live at Madrona’s 2025 IA Summit.

Jason: Yeah, so I think it starts with compute and data, and you could see us becoming a lot more active in this space in recent history, much more so than we were in years prior. Then, certainly, there’s the foundation model layer, which we’ve been active in for a long time. Then, there’s the application layer, ChatGPT, and then the applications you get built on the API. Those are essentially the basic, several layers, and then sooner or later, you’ll have devices that are kind of custom-built or purpose-built for this kind of technology. Then, you can even go more specifically within each of these layers.

In the future, you’ll probably see additional stratification, like at the application layer or in the model layer, you’ll find infrastructure in terms of tooling, eval suites, things that have to do with cleaning or structuring data when it comes to task-relevant data that you need to collect for post-training. I think all of those elements, you’ll see further deepening and development, and I think that’s actually a very interesting way to think about what you should build as a founder going forward, because a lot of this will become very relevant, not just to us, but various other foundation model labs. The pace at which all these labs can actually build all these tools and technologies at the various layers is not necessarily going to keep up with the pace at which the ecosystem can build all these things, and so there’s a lot of opportunity there.

Vivek: Yeah. We’re definitely going to get into the meat of how founders should be thinking about where OpenAI plays and where it doesn’t, because I think it’s so important and something we think about all the time, but going back to the point you were making around compute, infrastructure, the data, we’re seeing all this news about hundreds of billions of dollars being poured into or about to be poured into the compute side and the infra side, and OpenAI has had amazing partnerships with Nvidia, Oracle, and so many others, and I think one of the cool things for us as early stage investors and founders is this CapEx build ends up being really, really great for building applications, but maybe just take us through some of those partnerships and maybe just a little bit beyond the headlines, like why is it so important for OpenAI to be spending significant resources here around compute infra and with these partnerships involved?

Jason: Yeah. Well, I think just the simplest answer is, just because the scale models continue to work and we believe that they will continue to work, and so that just implies that the amount of compute that we’re going to need, both for training as well as inference. If you want to continue to push ahead to capabilities, you’re going to need increasing amounts of compute, and if you also believe that the demand for this technology is still in early innings, which would be the case if you kind of look at enterprise adoption. A lot of them are still in proof of concepts or pilot stages, and there’s a lot more token conception to come. If you are able to deliver on deep use cases that get repeat usage, there’s a lot more inference compute that you need that’s implied in the future, and so that’s really it, in terms of what is the actual impetus or driver.

Vivek: Right. And so, one of the major themes that has kind of been running throughout the day and on the previous panels as well, is what we talk about this reasoning revolution, right? And I think on one of the previous podcasts you were on, you were also talking about reasoning, and the name of the Fireside chat today is “The Collaborative Reasoning Full Stack.” So curious, what does that mean to you? What does reasoning mean in the world of OpenAI, and how you all think about the future of the reasoning revolution?

Jason: Yeah. I think there are at least two elements to it. There’s probably more than that, but two that immediately come to mind. So, one is, just going back to the prior topic we were touching on, which is that the more compute that you apply at inference time, when you’re using reasoning capabilities, the better the answers actually become or the better the quality of the response. And so, if you’re in that kind of paradigm, you are now seeing this kind of relationship between compute and execution, essentially. The other thing that is very much, you can kind of start to see develop right now, and this has, I think, been a topic today, as well as possibly yesterday, is agentic capabilities, which is really enabled by the reasoning stack, right? Because you can get to task situations, and it’s actually the reasoning capabilities that let the model understand what it’s supposed to do, especially in situations that would normally confound, like sort of normal, programmatic software.

You can also get into situations where systems can start to interact with each other, because they both have reasoning capabilities. You see elements of this previously with some of the tools that we put out, like deep research or our search capabilities, which use elements of reasoning to unlock agentic capabilities, such as searching on your behalf, understanding what it’s reading, doing analysis, and then writing the content and the report for you, which is agentic capabilities, but it’s able to accomplish those things, because it’s reasoning through the task that you assign to it. Then, up to one of our latest releases, which was on Monday, which had to do with agentic eCommerce with Stripe, which is now the ability for ChatGPT or other applications if they want to make use of the protocol, which we open source, to have the ability to leverage API and the models to have transactions occur based on commerce stacks and reasoning logic interacting with them. I think this is actually probably going to be one of the primitives, in terms of capabilities or building blocks, when it comes to AI, that sort of unlocks a lot more building onto it.

Vivek: Yeah. No, that partnership and sort of announcement you had with Stripe on Monday was really interesting, because I think one of the things that we think about is, “Okay. In this future of agents, how do they interact with each other, and how do payments work?” And there are so many parts down the line. Is your sense that that’s something OpenAI will partner with a lot of the next generation set of companies, or in the case of Stripe, one of the best payment companies out there? How do you think about what makes sense to partner, versus what you’re building yourself at OpenAI?

Jason: Yeah. I think our core activity is focusing on general intelligence, and so everything around that needs to be accessed or to work with that is much more something that we’re inclined to partner at, and so focusing on the foundation models and then the surrounding stack around it that enables us to deliver that efficiently and at scale, that’s what we’re focused on, but then taking that and interacting with commerce, taking that and interacting with content, taking that and interacting with other capabilities like research or whatnot, enterprise capabilities, those are all things that other people have been doing for decades and have a lot of expertise in. This kind of ability to use reasoning to interface with an entire ecosystem, it’s also very consistent with the nature of intelligence to be able to work with lots of different other systems. So I think partnership is really, very consistent with that part of our mission.

Vivek: Right. The interesting thing here is that we always hear about OpenAI talking about its mission of finding and fulfilling our journey to AGI and doing that in a way that’s going to be beneficial to all of us. What does that mean on a day-to-day basis? So, a lot of us hear about AGI from the outside. You guys are around that mission every single day. What does it mean, inside of OpenAI, to be working towards AGI?

Jason: Really, I think it’s about staying research-focused, so the company is still very much a research-centric, research-driven company, despite the fact that most of the attention is still on the products. That is mostly what gets popularized, but inside the company, if you were to just kind of measure the amount of content at all hands, for example, 90% of it is still talking about the research. One of the things, for example, I like to talk about with the teams that I work with is, when you apply a “5 whys” framework to something while you’re doing something, it’s like somewhere along the way to the fifth why, it should be because that advances, helps, or supports some core research activity. I think that that’s the one thing that’s very consistent about the inside of the company.

Vivek: Right, so research at its core, but product working in tandem with research.

Jason: Yeah, yeah.

Vivek: Yeah. And I think one of the things that is always top of mind for founders is, if I build something, is OpenAI just going to build it, or if I’m building something, what happens if OpenAI builds it? I think the old question would be, what would happen if Google or Microsoft built it? Now I think it’s very reasonable to ask, what happens if OpenAI builds it? So, we see so many interesting applications/tools coming out of OpenAI, so how do you think about what you are building and what makes sense on a 6, 12, 18-month timeframe? And to flip it on its head, where is OpenAI not building?

Jason: I think that this is an interesting question, and we can talk about this a little bit, too, in relation to stuff that we just did this week. So, a very general answer would just be, if you look at our mission, it’s about AGI. You could take issue with the fact that, whether you believe AGI is a useful concept or not, but the crux of it is that if it’s on the critical path to this general intelligence capability, that’s something that we’re going to be interested in. If it’s not so much on a critical path to that, that, by definition, is something that we’re going to be less interested in. That seems ostensibly simple enough, but every so often things will happen, and maybe it’s not so simple, or people get surprised.

And so, I think a couple of days ago, we released Sora, and I think some people wonder, “Is this actually on the path to general intelligence or not, and how does this actually fit?” I think that it’s very easy to look at the form factor of the product and think, “Oh. That doesn’t necessarily make sense,” but I would just say, we’ll look through to the actual capability, which is to understand the physical world as a world model, as a simulation that’s represented through video, and there’s a live question as to whether you can actually fully capture what you need to in order to get to general intelligence just through text. And so, if it’s true that actually you need sort of representation of the world and movie pictures, then having this kind of advanced video capability is possibly quite important to having general intelligence of some sort, and that is why it is important.

Vivek: It’s interesting, because I think that’s true. I think a lot of people would say, “Hey. You talk about AGI, but then you’re coming out with really, really good video models, and there are a lot of companies going after building video models and applications.” So, your sense is, these are things that are the outgrowth of needing the data and needing these modalities and the information that comes to fulfill our ultimate goal.

Jason: Which is general intelligence.

Vivek: Which is AGI. So, you’ve got a bunch of founders in the audience here who are building various things and thinking about building various things. What do you think are the open areas that OpenAI is probably not going to end up building against, that you would suggest founders should focus on?

Jason: I think it’s just posing that question, and I think that there are two ways to think about this. One is, “Hey, what is not necessarily going to be straight down, the middle of the fairway, when it comes to that general intelligence question?” And so, if you’re going to do some kind of product that is very focused on figuring out how to apply the AI models to a manufacturing process, that’s very specific. That’s probably not an area that we’re going to go super deep in, in terms of developing a whole product suite, right? And that’s just an example of how to think through it. There’s another way to think through this, too, which is that, and we see this happen all the time, and maybe founders in the audience actually have experience with this, which is you might be in a particular area and the current class of models are just a little bit short of being good enough to address the use case that you have. Maybe they’re slightly too expensive, or maybe they’re slightly not reliable enough, or maybe they just don’t quite hit the use case to the level that you want.

The interesting thing to do, then, is actually to bet on the increase in quality of the general capability of the models, as just sort of like a macro force, and that is an interesting way I think to also build a startup, because we’ve definitely seen a few where it’s just, that is what they’ve done, rather than try to over-optimize on this particular set of capabilities that you see on the models today, but actually, just bet that the general capability will continue to improve. And I think that that would be perhaps the wrong way to build a startup, is to actually say, “Okay. The current class of model capabilities is here, and now I’m going to spend a lot of effort in terms of fine-tuning, customization, data collection to then fix this last-mile issue with the fundamental capability,” rather than actually building productization around delivering the intelligence and the actual utility, which is about probably more than intelligence, that is then going to be further enabled by actually the growth of capabilities, just natively of the model itself.

Vivek: Right, yeah. So, don’t assume that the models are not going to get better, basically, and build around that. Maybe I wanted to just make sure we have enough time for folks to ask questions, and so I’ll open it up to the audience here in case anybody has anything they’d like to ask. I see a hand in the back there?

Audience Question: From a profitability standpoint, with the cost of inference still being a non-trivial number, how do you think about the business model from a profitability standpoint, and at what point do you see an inflection where the gross margins, from a user perspective, become positive and compelling?

Jason: Yeah, so for ChatGPT itself, we’re actually profitable in most markets, if you look based on compute margin. It’s actually the overall company may not be profitable because we continue to invest in actually scaling the compute and also scaling research, or maybe scaling is not the right word, but continuing to reinvest in research. And so, maybe the question behind the question is, what is the lesson for other companies maybe, in terms of how you should think about the business model, if OpenAI is kind of in this kind of financial situation?

I think it’s probably that you should be thinking about the margins relative to serving compute, really. Then, what other costs actually go into your variable input? But I think if you are doing well on the basis of, here’s how much you pay for compute, and then here’s how much you actually get per unit of whatever delivery and you’re positive on that, then actually, in a very plain English sense, the core input, which is compute, you’re actually deriving more value out of than you’re paying for. Then, you can probably also have a theory that the cost of compute itself is, over time, going to decline.

Vivek: Actually, picking up on that comment about the, we didn’t talk too much about financials and the revenue explosion of OpenAI, which is unlike anything we’ve ever seen, and more on a personal note for you, you joined OpenAI in early ’21 as GEC, so months before ChatGPT really exploded on the public consciousness. What surprised you the most about that launch and what you’ve seen since then? Has that really transformed the company and from what you remember pre-ChatGPT?

Jason: Yeah. This goes back to what we were talking about earlier, which is trying to stay research-focused, which is actually a struggle, right? So, it’s something you have to work on, and it was much easier to be a research-focused company naturally when the rate of change inside the company itself was not ChatGPT-like.

Prior to ChatGPT, I think we were about 200 people, and we were not adding that many people per month, and then after ChatGPT, we 3Xed every year in terms of headcount since then, and that’s a painful amount of headcount growth. And so, we’re still relatively small compared to lots of companies, but in terms of experiencing that amount of change, it takes a lot of effort, I think, to sort of retain certain types of working habits, styles, prioritizations, and an idea of what’s important and what are we trying to do. It’s a lot of work.

Vivek: Well, maybe on that, how do you manage that? Because it’s somewhat unprecedented, right? You can’t really look at any other comps. There’s not a lot of other companies that have grown as fast and as quickly, and so you say the whole organization has to change. There’s no blueprint for it, so is there one or two things that you’ve thought about or you’ve done tactically that has helped in this journey?

Jason: I think, really, I would just kind of keep it simple and just pick one, which is it comes down to leadership and focus, and then what does the leadership focus talk about all the time? What’s been really interesting is, so Sam, when he does all hands, and he always speaks at all hands, he always has a few minutes, right? He doesn’t speak to all hands, but he always has a few minutes, but when he does, he chooses to focus on research or compute, and those are the two things he just always talks about. And so, that center is the company, because it’s just the CEO sets the tone.

It’s funny. There have been times when I or other execs have been like, “Well, no. We need to talk about a bunch of these other things,” and he will nix whole parts of presentations or take what should have been a 20-minute presentation on what a normal company would actually spend 20 minutes on and on all hands, and he’ll be like, “Two minutes,” right? And I think that, during the course of all this growth, you’d be confused, because no, these are important normal company things to talk about, but in retrospect, it’s very clear what he was doing, right? It’s like, through all of this chaos and change, it’s like he was trying to make very clear what the main thing was and just kind of have that be a ballast.

Vivek: Ruthless prioritization, right?

Jason: Yeah, yeah.

Vivek: Super interesting.

Jason: And sending a very clear message.

Vivek: I think we have room for one more?

Audience Question:

Thanks. Thanks very much for being here with us. What can you share about how OpenAI is thinking about its role to play in commerce? Obviously, you had the announcement earlier this week. People ask for recommendations of products, look for specific queries, comparisons, things like that. Sam Altman’s been on record, very publicly, as not wanting to go down the advertising route. What can you share about how you think about that huge opportunity of what people are asking for help with and what your role might be in delivering that, and then how you might monetize that? Thank you.

Jason: Yeah, so I think we want to kind of approach this like an ecosystem player. I think that’s the first thing, and so it’s kind of going back to what that Bill Gate’s quote, which is like, “You’re not really a platform, unless you create a lot more value for everybody else than you create for yourself,” and I think that’s probably one, a very good principle to think about how we approach this space. I think the other part of this is, even when you think about it from a technical standpoint, which is the real sort of beauty of the reasoning capabilities and the agentic capabilities expressing in this way, as an interesting engineering and science problem, is not if it does everything inside of ChatGPT itself, but if it’s able to actually execute an indeterminate number of transactions with an indeterminate number of commercial players, right? And that is actually a much more interesting, technical achievement. And so, I think that people are excited to build on that, and that is partly what the agentic commerce protocol is about, and so I think those are probably two strong indicators of how we think about this.

Vivek: One more?

Audience Question: I just wanted to get your thoughts on the application versus the model layer, Cursor versus Claud, where does the value come from, and do you see OpenAI coming into moving up the stack?

Jason: I mean, I think you can just look at our actions here with Cursor, which is that we’ve decided to partner with them, because they’ve just done a very good job executing when it comes to actually the application space. Adropic’s probably taking a slightly different view, in terms of how they want to go further into the application space. It might just be that it’s not necessarily always a matter of, you have decided permanently that one thing is a particular type of approach.

It’s just that, here you have this company with an amazing founder and Michael, and they’ve done an excellent job executing, and they’ve cracked some way in bringing the light to engineers, and so it’s a good partnership and a strong one that we are pretty happy with. I think when it comes to the coding capabilities itself, for our own research, what’s interesting to us, really, is thinking about that as how to advance our own research, because the ability to actually have automated software engineering is going to also increase the rate at which you can run a bunch of experiments and do the research itself for AI, and that is the core of our interest.

Vivek: That’s great. Maybe just to wrap everything up, if you’re sitting here one year from now at our next IA summit in 2026, what’s one thing that you’d be excited about that OpenAI would’ve accomplished over the last 12 months over everything that OpenAI is doing right now?

Jason: Man, so I think that’s such an interesting question, but there’s a bunch of things on the model side that I could pick, but I think it’s like you could just look at our current state of activities, just extrapolate that out over 12 months, and say, “We just continue to deliver on those things,” and that would still be exciting, right? So, we’ve got agentic commerce, and if that really works, there’ll be a decent ecosystem of lots of building on top of that and a lot of additional opportunity for people.

We’ve got a new video platform that is intended to help creators monetize, too, and so there’s a lot of potential opportunity there for lots of other people. The third here is we just announced a bunch of partnerships to build more compute, and we’re going to do that with a bunch of players, including Microsoft and Oracle. That’s also going to provide a lot of opportunity, both for those partners, as well as people who are going to use that compute. And so, I think just executing on those things over the next 12 to 18 months, and seeing what everybody else is going to do with it, I think that is actually the thing that’s going to be really interesting and fun.

Vivek: It’s going to be really fun next year. Well, Jason, thank you so much. Let’s all give a big round of applause.

Jason: Thank you.

Vivek: Thanks.

Redefining Communication: Yoodli Founders on AI Roleplays, Confidence & Building in Seattle

 

In this episode of Founded & Funded, Madrona Venture Partner Patrick Ennis sits down with Varun Puri and Esha Joshi, co-founders of Yoodli AI Roleplay, an AI-powered communications coach helping people become more confident speakers.

Varun and Esha share their leap from Big Tech into startup life, eating noodles on secondhand couches while chasing a bold mission: making effective communication accessible to everyone.

They dive into:
• The origin story behind Yoodli, (and where the name came from!)
• How AI roleplays bring “exposure therapy” to speaking practice.
• Use cases ranging from sales training to doctors having end-of-life conversations.
• The decision to build in Seattle vs. SF with support from Madrona and AI2.
• Co-founder dynamics, culture-first hiring, and why they turned down a famous Silicon Valley VC.

This is a must-listen for any founder building with AI, thinking about culture-first scaling, or just wanting to speak with more confidence.

Listen on Spotify, Apple, and Amazon | Watch on YouTube.


This transcript was automatically generated and edited for clarity.

Patrick: Let’s start with a little bit about your your backgrounds? Why did you leave Google? Why did you leave Apple?

Esha: So my background’s in software engineering and product. I was at Apple for a couple of years based out of Los Angeles and London. It was a really fun experience right after college. Awesome to be able to lead some of their newest products at the time like TV+, which is their entertainment subscription business. While I was there and even before when I was in college, I really struggled to speak up. I still struggle with it. It’s not something that goes away, but with a lot of practice and intention, you can feel more comfortable with it.

And I remember at the time there were people of higher authority who struggled to say words, or I saw my friends and peers were really remarkable, but didn’t have… Rather, they were really remarkable; they had amazing ideas, but they didn’t have the confidence or training to use it. And so I thought there’s got to be a better way. What can we do with technology to help people, i.e., myself at the time, improve and be more confident? And what’s possible? I also always wanted to start my own company, so I figured this could be an opportunity to do that.

Varun: My backstory is that I grew up in India, came to the US for college, was at Google, ran special projects for Sergey Brin, who’s one of the founders. It was an incredible role, so much access and exposure, and meeting folks I had read about in the newspapers. I was the only person at Alphabet reporting to Sergey, so it came with expectation and pressure. From there, I worked at Google X, which is Alphabet’s moonshot factory.

My reason for leaving was, one, I said, “I’ve had this incredible opportunity. I want to do something really big with it and take some bets. Now, shame on me if I just continue to play the game when I’ve had some exposure to try and do something more with it.” And second, I’m very, very passionate about this problem. As Esha mentioned, her why is to help women speak with confidence. Mine is that I think there are so many smart people in the world, especially kids in India, who deserve to be on this stage more than either of us, who deserve that job more than the extrovert who might get it, for example. And they miss out because they don’t back themselves while speaking. So it’s how do we build technology that can help billions of people do to speaking what Grammarly did to writing or Duolingo to language learning, and impact lives when it comes to both professional and personal outcomes.

Patrick: It’s interesting. You’ve always emphasized to me that over 280,000 years of homo sapiens history, humans have always struggled to communicate. And they’ve come up with various ways to coach and get better, but it’s always been done the same way until recently. Is AI really the reason why now?

Varun: I actually don’t see Yoodli as an AI play, per se. For us, it’s how do we solve the problem? The problem is that people doubt themselves before important conversations, interviews, sales pitches, salary negotiations, podcasts, etc. It’s more of an exposure therapy question. How do we normalize the process of recording yourself, watching yourself, and cringing? You do that three times, you’re going to get better. I think AI is one tool that enables more of that exposure therapy, but if it weren’t AI, I’m sure there would be another way to enable this gamified exposure.

Esha: To add to that, if you think about what options are today to better yourselves speaking-wise, that ultimately culminate in the act of exposure therapy. You’ve got coaches of any kind, speech coaches, interview coaches, career coaches, sales coaches. You have Toastmasters and various clubs and opportunities to help with that. All of these are really great options, but they are inconvenient for a number of reasons. Number one is cost, two is physical location, and three is just getting up in front of people and speaking. These three ways pose challenges for folks who don’t have access to resources, don’t have access to transportation even. And so the thinking is for those folks who need a safe space at least to get started to practice, nothing revolutionary with the idea of practice. You think about any sport, any new skill, you have to practice and you have to get yourself over that initial hump. And I think with communications and communications coaching, it’s no different. Practice in the case of speaking, see yourself, hear yourself, cringe, and then if you can get over that initial hump, you’ll get to a place of normalizing this activity.

Patrick: And you have a very varied customer base. My 93-year-old mom loves Yoodli because she goes to the Elks Club events, and occasionally she has to stand up and give a short speech. And that really hit home. And I realized your market size is literally the 8.2 billion people on this planet.

Esha: When you say it like that, it’s a big number. And it is, honestly, it is fascinating to think that, but it’s true. We’ve got individual customers who are using Yoodli for a big presentation tomorrow and an interview. And then we have companies using Yoodli, too. This could be for sales training. It could be for conference prep for big events coming up. It could also be for helping your partners at a company who are selling your products get up to speed. So there are varying use cases. The tool supports all of the three, and we have different ways of adding value across the board to make sure that everybody has a good personal experience. But then also companies and coaches are able to use the tool to understand how their clients are doing and how they’re progressing over time.

Varun: And just as an example, the kinds of use cases that have been really fascinating are there’s an obvious focus on GTM enablement. You need to train your sales, your customer success, your user-facing teams on how to pitch with confidence using your company methodology. Or we have healthcare organizations, hospitals, that are using Yoodli to train doctors on difficult end-of-life conversations. That is so impactful. There’s one thing to train a sales team, an incredible use case. There’s another to teach a doctor how to have challenging life and death conversations with patients. And we’ve learned from our user base that people have made Yoodli their own. Just last week, I was speaking to someone in India who had a stroke at a really young age and has had a very deep stutter since. And he says Yoodli is his private practice environment. We’ve heard of applications and people using Yoodli for dating. We know of investors who are using Yoodli to practice before the LP pitch. It’s super fun and interesting.

Patrick: That brings up a great point. When you’re doing startups, everyone tells you it’s all about focus. That’s actually not true all the time because if you focus too much, you don’t capture the broader opportunity when you have a very powerful capability like what you have. So how have you thought about that, just nailing one particular wedge versus going after the North Star of 8.2 billion people?

Varun: The way I think about it is, we are building a category called AI role-plays. Our laser focus on the enterprise is GTM enablement. That’s what we know we do really well. We quantify the impact and the outcome, but then we have a consumer tool where consumers are making Yoodli their own, and we have folks approaching us for a whole host of use cases. So I think it’s possible to build a horizontal platform the way we’ve built it, fingers crossed, while having specific verticals that you focus on from a GTM standpoint.

Patrick: So soon, I want to get into team building, and how you built the company, and why you chose Seattle. But first, let’s ask the question that everyone is no doubt wondering. Where does the name Yoodli come from? It’s such an awesome name. What’s the etymological root of Yoodli?

Esha: Well, okay, we have a couple of different stories. I’ll give one of them, and then Varun, you can give some other ones. When we started, we were an AI-powered public speaking coach. It stems from our love of public speaking, but also this really visceral feeling and pain of, oh my gosh, it’s scary to get in front of people. And so when we came up with a name, we’re like, well, what is similar to audio and voice and sound? Yoodelay-hee-hoo. The Sound of Music. I think a lot of us grew up knowing what that is. And so Yoodli is like that. That’s reason one.

Varun: Reason two and three that we publicly talk about are, well, it has two O’s and an L. It sounds Silicon Valley-ish, it has Hulu, Google, etc. The SEO and trademark was relatively easy. The truth is, all of that’s a lie. We say that publicly because it makes a good story.

Patrick: By the way, this is public, just so you know.

Varun: I know, but it’s a podcast. I’ve not said it enough times. Most people enjoy the story. It was my freshman year of college. I had just come to the US. I was looking for my roommate, Tyler. I was a couple of Fireball shots in, and I couldn’t find Tyler. There was another Fireball shot, and I was like, “Tyler, where are you? Yoodoohoo. Yodoohoo. Yodoohoo. Yoodli.” And Yoodli just became this chant within our friend group. So when we would go to a party or we’d be looking for each other, we’d just say, “Yoodli.” I mean, in our friend group for fun, we keep saying, “Yoodli.” When we came up with the idea for our company, our thought was that it would be so cool if grown adults said Yoodli with a straight face. And some of our favorite moments at Yoodli, even today, we’ve been getting some press, et cetera, and investors write about Yoodli, is when a college buddy I haven’t heard from in a couple of years will message me saying, “I cannot believe my organization is saying go use Yoodli. Like, I cannot take that seriously.”

However, the way to validate this is when we came up with the idea, we wanted it to be true to our personality. So our best friend, Andrew, we’re all part of the same friend group, came up with our first little jingle. So if you look at any of the Yoodli videos, it has whatever the formal stuff is and then, at the end, it has a Yoodli logo and 30 seconds later you’ll hear Andrew saying, “Yoodli.”

Patrick: That’s wonderful. So let’s pivot now to the company building.

Esha: Can I actually add one thing quickly?

Patrick: Of course.

Esha: So there’s some part of the population that knows what Yoodli is and can pronounce it. There’s a larger part of the population that does not know how to pronounce Yoodli so much so that every time we get on calls, we know, okay, the first couple of minutes we’ll be like, “Is it Yhudley? Is it Yoadly?” And we have to be like, “It’s Yoodli.” With a smile on our face. That led us to create a roleplay on the Yoodli platform to educate others on how to say Yoodli.

Patrick: Because you wouldn’t change the name.

Varun: It’s too personal, Patrick. No matter what marketing expert tells us to change it, we just can’t. Our friends will, at this point, disown us.

Esha: And people come to us and they’re like, “You’re a B2B SaaS platform. Don’t you think you should have a little bit more of a sophisticated name?” And I was like, “Well, Hooli in Silicon Valley is far from sophisticated. Google is a made-up word. So what’s the big deal?” And so here we are.

Patrick: I think that’s great. We always talk about staying true to your North Star. You have a vision, and the name is part of that vision. So let’s get to that. At the very beginning when you’re starting the company, it’s all about relationships and who you’re going to involve in a company. Then it’s about where do we build the company. So let’s start with the team. So, where did you two meet, and where did you two first start working together?

Esha: We met about 10 years ago at a college internship at Intuit in Mountain View. And it was a college internship, so we’re coming to this first experience in tech with these bright eyes and bushy tails. And being in college at an internship, we had a lot of fun.

Varun: We were sophomores, so we weren’t as worried about getting the job offer. Everyone else was more trying to convert, Esha and I. I think we got the internship at the last minute, so we were just there having a great time. We became very good friends. The Intuit crew is actually some of our closest friends, and we’ve been friends since. Our entire friend group was digital nomading during COVID, and around one of these COVID, socially distant dinners, we came up with the idea of Yoodli. We tried to convince some of our other friends as well to join, but it was just Esha and me who jumped in.

Patrick: And when that conversation happened, you were at Google and you were at Apple, or were you leaving?

Esha: We were employed.

Patrick: Okay. So it’s a big step. You go from working at famous companies, probably making decent money, good benefits, to a startup where initially you’re making nothing and you’re eating cans of beans in the basement That type of thing. Is that what you did initially?

Esha: Similar to it. It was probably cans of Maggi noodles that our parents had given us, sitting on couches that were bought secondhand, figuring out what we’re going to do and working through the nights, working, iterating on random things. But yes, that’s true. A lot of that was coming to Seattle, having a very intentional decision of leaving where we were before. It was the pandemic, but prior to the pandemic, I was splitting time between London and Los Angeles. Varun was in South Africa or was in sub-Saharan Africa. And so for us, it was a big move to come to Seattle and to start this company. And we were really attracted to the opportunity here of there being a number of great investors, but also the AI Incubator, where we figured, okay, we’ve never done this before. We’re first-time founders. We’ve been told that there’s this AI thing that we could use to help with this. We just want to-

Patrick: This is the Allen AI Institute? Which Madrona has been a big part of through the years.

Esha: That’s right. But when I say AI thing, I don’t actually mean AI2, I mean there is some concept of AI, and we were learning about it, and we were saying, “Our goal is to help people become more confident communicators. How we do that is we’re going to figure that out.” And Seattle was a really interesting place with resources like the Allen AI Institute incubator, Madrona being a VC, and there being a tight-knit community space to work on this.

Patrick: Silicon Valley had all of that to an extent. So take me through your thinking. Was there something specific about, well, we love the sunny three months in the summer, and we love the nine months of rain and darkness. I mean, what else factored into your decision?

Varun: A couple of things. One, both of us had been in the Bay Area, and we wanted a change, and we knew we wanted to be close enough to the Bay Area. The second big decision factor for us was that we spoke to Oren Etzioni. Oren’s a close mentor and friend. When we’d just came up with the idea, he was very blunt with us. He was like, “Look, y’all are first-time founders. You don’t know too much about too much.” And if you all know Oren, that’s exactly what he’d say with all of the love and warmth. “Why don’t you come here and we’ll help build a company around you?” This is pre-GPT. “We’ll give you the resources to experiment with AI, we’ll give you the place to continue to test with new hires, et cetera, as opposed to a typical incubator, which is in and out in three months.” And given the stage in life we were at, we’re like, “YOLO. Why not?” So I think from having the idea to leaving the respective companies was a very, very short time period.

Patrick: So, fast-forward to after Madrona was involved, we were already investors, things were going well for Yoodli, and there was very famous Silicon Valley investor who wanted to invest, give you money, take a board seat, and you turned it down because it didn’t seem like it was the best fit. That must’ve been a tough decision. Take me inside what you were thinking about the pros and cons at that moment.

Varun: I can take this. I mean, we’ve had a few of those, which is exciting. Where we’ve had names we’ve heard about in the media all of our lives offer interest in Yoodli. Think for Esha and me, our why is we want to help billions of people speak with confidence. We are starting with GTM enablement. We are seeing great traction there, but we want to build a company that’s very true to who we are. We have this B2B thing and this B2C thing. We’re doing B2B, but we keep saying, “Yoodli.” We own our mistakes. We come to Madrona being like, “Hi, you all took a bet on us and here are 50 things we don’t know.” And sometimes have asked really silly questions.

In the company and the culture we want to build, we want to experience more of that. And at various crossroads in Yoodli’s history, be it with a dream candidate, a potential customer, an investor, whenever we’ve been at a crossroads on do we take this check or not, is this the right person to give the offer to, even if they would transform our trajectory? We’ve basically looked each other in the eye and been like yes or no and then we take a snap decision. So far it seems to have been okay. I think both of us are very intuitive, gut-driven founders.

Patrick: And so many famous companies throughout history have had two founders, whether it was Steve Jobs and Wozniak or Bill Gates and Paul Allen. I’m not putting pressure on you two. There are advantages, but sometimes there are disadvantages to having two founders. So, can you tell us some stories about times when it really helped to have an equal co-founder and times when there was some friction because of that?

Varun: I can give a couple. It’s Esh and I are very, very close friends, which means there have been so many times where I’ve just made mistakes. At team meetings, I’ve said something ridiculous, at a board meeting, or fundraising, I’ve said something that may not be fully true to us. We will go for a walk, and Esha will absolutely call me out. In the team setting, we’ll have a good conversation, we’ll go on a walk. And she’s like, “What is wrong with you?” It’s at the point where, when it’s my birthday, my mom and dad will message Esha saying, “Ensure Varun takes the day off.” It’s so personal. I don’t see how I can do Yoodli without Esha. It’s just so much more fun this way. We have our friends who are equally invested in the success of Yoodli. Our fights, we rip each other’s head out with respect, et cetera, as siblings would, but we know we will have each other’s back tomorrow, no matter what. At least that’s how I feel.

Esha: I feel the same. What’s interesting is, as a founder, you have to make a lot of decisions, which is a vague thing to say, but decisions about who to take money from, who to hire, leaders to bring on, strategic direction, but also, being first-time founders, there’s so much that we didn’t know and still don’t know. So we come across these decision points and we’re like, “Well, what do we do?” And we started a company before ChatGPT existed, so now we can be like, “What do we do? Let’s both search on ChatGPT.” But before there wasn’t that, so you’d Google to search for, but you also are like, I want to talk through this with someone. I’m going to talk to somebody who also doesn’t know what they’re doing, but at least we both don’t know what we’re doing together.

So there’s this idea of having somebody who cares about something as much as you do with you, which is a little bit hard to do sometimes with an investor, just keeping it real, or with a candidate or an employee. And so that’s been one of the biggest highlights is having somebody so you feel less alone in the founding process.

Patrick: So you’re first-time founders and you’ve been doing a great job from the beginning, but where was that moment where you felt like, you know what, we may be first-time founders but we got this, we understand what we’re doing and we’re not going to be as nervous as perhaps we were at the beginning?

Varun: I think it’s gone through ebbs and flows. There are times I’m like, “Oh my God, Esh, we figured this out. This Yoodli thing is taking over.” Then just the next day I’m like, “Oh my god, we know so little.” There are many moments mostly around convincing people we really respect and admire to take a bet on us. I would say landing Google as our first customer or in the early days getting Google on board was such a voucher of approval when we were just starting the company. Patrick, having Madrona join us, was, “This is awesome, Esha. We likely know what we are doing. People like Patrick and Matt are willing to take a bet on us.”

The biggest one for me is getting the right hires, especially leader hires. Our approach to company building is we don’t have all the answers, but we will try our best and be really honest and vulnerable. And in doing so, we’ve brought on Derek who’s our CTO, who has built and led teams before and understands the engineering scaling side of things. We just brought on Josh to run our revenue function. Josh ran a lot of sales in CS at Tableau. He was a VP at Salesforce. I mean Josh is an industry expert and we often joke that convincing someone like Josh to join us is one of our biggest professional accomplishments because we will learn so much from him.

Esha: And have been learning so much. And also another piece of the founder experience that can be really humbling is to bring on folks that are wholeheartedly better than you in some function and then say, “Okay, we’re going to make space for you to come because I think you can crush it and you can handle this part of the business in a way that we can’t.” And so time and time again, almost for every single function or in some cases many big decisions, we’ve brought on folks to help with that. And then we’ve moved to another part of the business, still overseeing and being really involved. But there is a really interesting internal dynamic that you have to reconcile with and be like, this is what’s right for the company, therefore it makes sens,e and I’ll go focus on something else.

Patrick: That’s great. And how do you balance when you’re early building the team, you tend to reach out to people you know to hire them, but as you grow, you’ve got to do more of a broader search, often a nationwide search. And it’s not quite as personal, and I don’t mean that in a bad way, but you’re someone who perhaps they weren’t there for the first four years, they don’t know the origin of Yoodli, although you make them watch that video over and over again, I’m sure. But how do you adapt your recruiting and hiring strategy as you grow and as you get bigger?

Varun: We are still figuring it out. So we were 10 at the start of this year. We are now 30, we are growing to 50 at the end of this year. We have frameworks in place to scale ourselves, whether it be through investor networks, recruiting agencies, folks on the team who are specialists in hiring specific roles. That being said, Esha and I see our roles as being culture custodians or at least ensuring people who come in will expand our culture. We are very closely involved in every hiring decision. In most cases, we’ll do the initial 15-minute vibes check, and we’ll get on a call with a candidate and say, “Look, we want to tell you everything that’s awesome and everything that’s not about Yoodli and see if you’re still excited.” And then based on that, they go through the rest of the flow.

Esha: There was a period of time when we were pre-product-market fit, and we were figuring things out. And the reason I bring that up is because it gave both of us time to actually clarify and articulate what the values of this company are and whether they are different from personal values. And I think the answer is not really. It took us the first year and a half to be like, “Well, this is our inclination for who we want to bring on.” And now after having some data of who stayed at the company, who’s left either for performance reasons or culture reasons, we can clearly articulate that.

And so to bring that full circle now as we’re hiring people relatively quickly and we’re doing vibes checks and also skill checks, we have that written out into a values doc, which seems like a pretty straightforward thing to do, but the act of getting to that doc and that list of core values that we can then now use and repeat and show a candidate actually brings that more front and center in the conversation much sooner than it being an ambiguous topic that is hard to clarify.

Patrick: Have you found the Seattle market and the Pacific Northwest market deep enough that you’ve been able to hire locally or have you started recruiting people from out of region?

Varun: Absolutely. I mean the talent pool here is incredible. In fact, we do have some remote folks, but our entire focus is on building an in-office culture, and we fly remote folks in. We have been more than overwhelmed with the quality of candidates from the hungry new grad from UW who’s knocking on our door to someone who’s done Amazon, Microsoft from an engineering infra standpoint to incredible startups. I mean, we are right next to places where Lexion and Read AI and WhyLabs were. That is so much knowledge sharing.

Patrick: Yeah, that’s great. I think the Pacific Northwest ecosystem has really developed certainly over the long time that I’ve been here. When I first moved here in 1998, seems like forever, but when we were recruiting people from out of region, a real issue was the trailing partner or the trailing friend issue. It’s like, well, can you find a job for my friend? Can you find a job for my partner? And Seattle wasn’t quite as deep back then, so it was harder. But now when you need to move people out of region, you can. But the reality is you don’t need to because such a depth here.

Varun: I mean also for those of you all who haven’t visited the AI House, that’s where we are headquartered. The place is bumping. There’s events happening every day. There are maybe 15 startups in the area. It’s just so much fun. And I feel like one way in which Seattle is very different from the Bay Area is everyone cares. I don’t know if other folks have seen this as well. Seattle just has this incredible community where if on LinkedIn I message anyone being like, “Oh, you went to UW? We are working with UW.” Or “Hey, I’m in Seattle too, would you be down to chat?” Seattle gets a bad rap on the Seattle freeze when it comes to making friends. I don’t know if that’s true or not, but on the professional side, I’ve seen the exact opposite.

Esha: It’s lovely walking through a building like this one and on different floors, people who are part of the broader tech ecosystem and even just walking here running into somebody who is also working on their own company. And that connection is there because of events that investors, VC firms put on that people in the broader community do. And that’s really lovely.

Patrick: So, back to recruiting. When you’re recruiting, let me ask you to rank order four attributes, and I think I know what your answer will be, but integrity, work ethic, relevant experience, and raw intelligence.

Varun: Integrity, work ethic, raw intelligence, and then

Patrick: And relevant experience.

Varun:… and then skills.

Esha: It also depends on the hire. If it’s a leader hire, relevant experience is higher up, but I think it’s always going to be integrity. What was the second one?

Patrick: Hard work or work ethic.

Varun: I just, and maybe this is naive, but one and two, I mean one is table stakes. There’s no question.

Patrick: Yeah, you can’t teach integrity, and it’s really hard to teach work ethic also.

Esha: If those things are not there, then you get into conflict or you get into really hard decision making, and what seems fundamental falls and drops, and you’re like, whoa. There’s a broader, bigger issue here around how we see the world and how we make decisions in hard times. That is a problem. Without figuring that out, it’s difficult to actually solve the main reason why we came together anyway.

Varun: Yeah, I mean, one of our values, and we ask candidates this point blank, is to disagree and commit. Give us examples of when you disagreed and committed. And if we get some, “Well, I was perfect and then I realize I’m slightly not perfect.” We’re like, “Okay, give us three more examples.” So we can really get to the root. One of the things we look for first is people to be like, “Here’s an example where I really messed up and own it, and my God, I botched it so hard, and I felt terrible, and I will never do this again.” I think those are the kinds of stories that resonate with us the most. When I’m talking to candidates, I often say, “Look, we hire for humility, authenticity, vulnerability, and then the rest.” And I think we do the same thing in our interviews.

Patrick: So let’s finish up with some fun lightning round questions. We often learn the most from our struggles and when things aren’t going well. So what do your customers not like about the product and what are you working on to remedy that?

Varun: We often get the question of — is Yoodli doing everything for everyone? That’s not a product question; that’s a marketing question because we are trying to build a category. When they use a product, they love us, but often from the outside, and competitors will position against us in this way, where, “Well, Yoodli, they’re for manager training, they’re for interview prep.” And we’re like, “Huh, go use the product. It’s actually GTM enablement and more.”

Esha: Some of our customers ask for their learners’ practice sessions to be automatically shared with them for visibility, but also accountability of practice and then being able to see the scores change over time. And our answer is, “Nope, no can do. We’re not going to automatically share people’s recordings with managers.” We started off as a judgment-free personal private coach, and so that’s what it’ll be. Of course, we’ve got sharing functionality, and we make it really easy for learners to do that, but we will never automatically share recordings. Maybe a conversation intelligence tool, we’ll automatically do.

Varun: Oh, one more good one. We’ve had many customers ask us, “Why don’t you use AI for evaluation? You use AI for interview prep. Why can’t we just use AI to pick which candidate to hire?” We’ve lost so much money saying no to deals like these. We’re like, “No, we think AI is good. It’s not perfect. A human should be the decision maker. Have a human in the loop, not AI making final calls.”

Patrick: So what’s one thing that you wish you started doing earlier as founder, CEO?

Esha: I think probably posting content and establishing ourselves as the company is thought leaders in a space. Because what we’ve seen some really good founders do, second time, third time founders, is even before they have a product, before they have customers, they’re just writing about establishing themselves as a thought leader, but establishing the brand as a brand to know. So then by the time the product is released or you’ve brought on customers, you seem bigger than what you are, just given the knowledge you have. And that’s a tactic I’ve seen done really well, especially among B2B companies. That’s one small, albeit significant thing I’d do differently.

Varun: One for me is, and I think it’s taken us time to get to the confidence to be able to do this, go to customers with a very clear point of view of what we believe right or good looks like. A good example is we are like, “Look, the future of is experiential. Nobody’s going to, or very few people are going to learn from books or passive videos. It’s all going to be these roleplays.” Now that’s a provocative thing to say. We aren’t a hundred percent sure if that’ll be right or not, but we are starting to take bold bets, put our chips on something, and be like, “This is what we believe to be true.” And we are making proclamations around that until we are otherwise disproven. And all with humility, of course, and being willing to have strong opinions loosely held. I think when we first started Yoodli, we had a great idea with a lot of opinions. Now it’s everything we are doing is in direction of those few key opinions.

Patrick: So let’s finish up with another personal question and topic. You and I had an interesting conversation a couple of days ago about what it means to be a founder and who’s a serial founder and who isn’t. And many founders are like, “Oh, I’m a serial founder. That’s all I’ll do for the rest of my life.” And it was interesting the way you approached it. You’re like, “Hmm. I don’t think of myself that way. I think of myself as the founder of Yoodli.” And I thought it was a great answer because it’s from the heart. But please elaborate on that, and I know you both may have a few words on that to finish up.

Varun: Yeah, so each context on this is Patrick and I were just chatting, and he’s like, “Varun, what happens? You love Yoodli, you want to do it. What happens after?” An honest answer, I think I’m a really good founder of Yoodli, and I want to do Yoodli for a really long time. My running joke is that we are going to take Yoodli to the moon or down to zero. I think different founders are incentivized by different things. I hope we make a lot of money, whatnot, but it’s more about how we build something that so many people find value in. I love being the guy at the airport wearing my Yoodli jacket, and someone comes up and is like, “Oh my God, Yoodli. I used Yoodli to land this job.” Or my mom’s at a dinner party in India and someone goes up and says, “Your son works at Yoodli?” It’s just such an incredible feeling.

The long and short of it is, I want to do Yoodli for as long as I can, and I have too much fun with it. I don’t think I have it in me to do multiple startups. I don’t know how I’ll get this attached to anything else. I don’t know if I’ll be good enough at any other idea. And boy, if and when Yoodli does end, I’m on a sabbatical. I am peacing out for a long time.

Esha: Honestly, no additional notes. Yeah, I feel the same. It’s funny because people are like, “Oh, I’m kind of tired of this idea. I’ll move on to the next one.” I’m like, “First of all, don’t you want to take a break? Are you tired?” But second, going back to what he said about, Varun said, “I think I’m a really good Yoodli founder.” I feel very passionate about what we’re trying to solve. And so much so that three years, two years before starting Yoodli, when I was thinking about what should I start a company in? I was always thinking something with confidence and communication because I cared about it. And so in a lot of ways this is that, and I, from years before Yoodli, I felt that way. Even throughout Yoodli, I still feel passionate about it, and I want to work on it for as long as we can. And the next idea, it’s too difficult to think that far in advance, but I’m also like, yeah, I would like to relax for a second in between and really process everything that we’ve learned.

Varun: And we want to take this all the way. Too many people struggle with this problem. We could spend all of our lives together and still not solve it. Let’s make a real crack at really advancing the world with this problem.

Esha: With 9 billion people and a lot of folks as individuals, maybe not working, it is a long journey to world domination. Phase one could be this B2B SaaS approach, and that’ll have to be multiple stages, of course, I’m sure. But to circle back and help those people who don’t have access to resources, that is a long journey. So here we are.

Varun: Yeah. And we’re in a stage V0 dot crap. It’s not like we’ve cracked anything yet.

Patrick: Well, that’s a great way to end it. Really appreciate your time. Know how busy you are. I know everyone enjoyed listening to your type of stories and a lot of great advice, and congratulations on the success you’ve had so far with Yoodli.

Varun: Patrick, thank you for having us, and if this isn’t evident, thank you to all of Madrona. I mean, our running joke is Madrona, the third co-founder. Anytime we have any questions. This morning for instance, we opened a new job rack. I kid you not, I emailed 25 people from Madrona on the same thread, being like, “Hi, can you help us hire?” I don’t know if this is typical of VC firms, but just the number of, the frequency, and the sometimes triviality of questions with which I reach out to you all, it’s just incredible. We feel very grateful.

Patrick: Well, it’s my third VC firm, and I will say no, this is not typical of VC firms. We work really hard at Madrona at that because we want to earn the trust and respect of the entrepreneurs who are at the center of everything. And I wish we had time to name all 25 people because there are 25, but certainly Matt McIlwain and Rolanda Fu, we’re there 24 hours a day for you with the rest of the team. Well, to finish it up today, thank you both so much for joining me.

Varun/Esha: Thank you for having us, Patrick.

 

From Google and Amazon to NewDays: Why These Tech Vets Bet on AI for Dementia

 

In this episode of Founded & Funded, Madrona Managing Director Tim Porter sits down with Babak Parviz and Daniel Kelly, co-founders of NewDays, a platform purpose-built for older Americans living with cognitive change to help them reclaim abilities, preserve independence, and keep being themselves.

Babak and Daniel share how their experience at Amazon and Google led them to apply AI in the most human way: helping people with dementia and mild cognitive impairment (MCI) with scalable solutions.

They dive into:

  • Why we chose to build at the intersection of AI and cognitive health
  • How to translate clinical science into accessible daily experiences
  • What “meaningful velocity” looks like inside an early-stage company
  • How to balance mission-driven purpose with commercial viability
  • Lessons learned from scaling startups and building impactful teams
  • Why the next wave of generative AI is human-centered, not model-centered

Listen on Spotify, Apple, and Amazon | Watch on YouTube.


This transcript was automatically generated and edited for clarity.

Tim: Well, let’s start with the obvious. This is a massive problem, and it’s a deeply personal one. One in three people over the age of 65 in the US is dealing with some form of cognitive impairment, dementia, or early-onset Alzheimer’s. We all have someone in our lives, either a parent or a relative, an aunt, uncle, a close friend, and we know it’s heartbreaking because there’s really no treatment to date, there’s nothing you can do. You’ve now identified a treatment path that is actually validated, and technology, specifically large language models, has made it possible to go build a service to address this. But yet it’s a big step to decide to start a business. With that big problem, what made you say,” this is the right ‘why now,’ let’s go start a business to address this opportunity or this big problem?

Babak: Maybe to take a step back and share where this all came from. I was at Amazon for a number of years, and one of my roles was to figure out what the company should do next. So we ran many different investigations in many different areas, and one of the investigations that we ran was about aging and what happens to older people. We looked at this nationally and globally, and we surfaced many problems, and this was, I would say, one of the most daunting areas that we looked at. But it was really hard to find radically new solutions to these problems that we could believe or have ourselves believe that they could move the dial in a meaningful way.

And as you mentioned, if you think about the scale of the problem, today, if you look at the population over 65 in the United States, 11% of them have dementia and on top of that, 22% have something that’s called the mild cognitive impairment, or MCI, but that’s truly a euphemism. Someone with MCI is impaired enough that they may not be able to pay their bills anymore. So these are very substantial cognitive issues. So one third of people over 65 are battling with these cognitive changes. And we all know that our ability to remember things, our ability to reason about things, they really, to a large extent, define who we are. So once they start to go away, we feel like we are fading away, and people around us also feel that the person is fading away.

So there’s substantial emotional toll on the person. And after that comes financial issues, many health issues. So these are very, very serious things that we have to deal with. And even though we just say, okay, one third of people deal with that over 65, but if you think about any person living in the US today, they are very likely to have a spouse or parents or siblings, so with the likelihood of 90%, this situation would hit one of us sooner or later. So basically, it’s going to hit all of us.

And if you think about what solution is available to people today, if I go to a doctor and get a diagnosis of MCI, for the most part, the next step is good luck. There’s not much you can do. There is no drug today that can cure these diseases. There is no drug today that can even stop these diseases. So this is highly motivated both Daniel and I to do something about it. And the way that we went about doing something about it was to step back and see where it is that we have solid clinical evidence of an intervention that can help people? So we did all the background research to surface randomized clinical trials, and that’s the gold standard of medicine, that showed any form of intervention that was beneficial to people.

So we went after those, but we realized none of them have really scaled because they’re limited by the availability of trained expert humans. And then we realized that now we have also radically new technology available to us in the form of generative AI. So we put the two together. There was a massive need, unmet need, and there was a radically new technology available that could bring medically proven interventions to a large number of people, to put them all together and decided that this is the right time now to go after this problem and help many people that have no recourse otherwise.

Tim: Amazing. Babak. I was not aware of this research. And just in simplistic terms, these therapeutic conventions really boil down to immersive conversations frequently. And there’s a clinical rubric behind it; you guys should describe that more. How does that come together in a product?

So you had this insight, there’s real research that underpinned it. There’d been no ability to turn that into something that scaled before. Daniel, tell the audience what actually is this thing? I’m sure a lot of people are like, “My gosh, I have people in my life I want to introduce this to.”

Clinician-Guided Conversations, Powered by AI

Dan: Absolutely. I like to think of NewDays as a therapeutic intervention for people with MCI or dementia. And really, that comes down to, so it is a telehealth clinic where you meet with a clinician that has expertise in cognitive therapy and then we kind of amplify the ability of that clinician to work with you through these conversations that you do with a large language model.

As far as the product goes, there is just an ability… Well, really all of it is kind of based on these three different methodologies of cognitive stimulation therapy, cognitive rehabilitation, cognitive training. And I think that those methodologies really spell out a range of conversations that can be beneficial for people with MCI and dementia. On one end of the range is just casual conversations, long-form casual conversations. I almost think of this as like going for a walk. These long-form casual conversations are just good things for you to do to maintain your ability to continue to have conversations with the other people around you in your life.

And within our product, these casual conversations are a good way to promote reminiscence about past memories. They’re a good way to practice verbal fluency, so finding words that you want to use to express yourself. They’re also a fantastic way to reinforce the concepts that you’re working through with your clinician in the telehealth setting. Repetition is important for the memory of someone with MCI and dementia, and these kinds of casual conversations are a good way to promote repetition.

Maybe on the other end of the spectrum are more challenging conversations, so these are maybe like doing sprints. These are designed really as a stimulus to stretch your ability within a particular cognitive function, and then that stimulus becomes something that your brain has to respond and adapt to. And then ideally, those challenging conversations would happen in a context that is as close as possible to a real-life scenario, so that you are challenging yourself in these conversations. But then in your real life, there’s a quick transfer between realizing like, “Oh, I was just working on this same skill through NewDays, but now I find myself in a similar situation in my everyday life, and I can apply the same methodologies there.”

Babak: And if you think about all of them, they’re highly personalized and they’re delivered in a conversational form and they require someone who’s trained to deliver these therapies. So they’re highly dependent on the availability of that trained individual to work with a patient to deliver the therapy. And that has been the issue with scaling these interventions because we do need millions of trained professionals to deliver these therapies. They’re unavailable, and even if they’re available, the cost of the highly trained humans to deliver these therapies would be astronomically high.

Clinical Foundations of AI for Dementia

So that’s why these therapies did not scale, and that’s where our NewDays.ai comes in. We allow the scaling of these therapies that are highly personalized and conversational through the use of generative AI or large language models. So you mentioned some of the clinical trials that we have anchored our work on. One of the most exciting ones that we came across was a study that was led by Professor Hiroko Dodge. She’s a professor of neurology at Harvard University. This was a long study, took many years to run it, but the results are pretty fascinating.

What’s the most amazing result about this study is that they managed for the population that was participating, this is a randomized clinical trial, registered, they managed to increase, this is highly unusual, increase the cognitive score of some of the participants. And what this practically translates to is to push back the symptoms of cognitive decline by six months or more. So this is really incredible for giving people time with their cognition.

If you look at those conversations, they look like normal conversations that you might have about the particular topic. Let’s talk about the Second World War or something like that, but under the hood, they are designed to encourage reminiscence. They are designed to challenge the person’s vocabulary in a particular way, and they’re designed to challenge the person’s critical thinking. So even though on the surface they look like normal conversations, when they’re delivered, there’s actually a purpose for these conversations.

So we saw this, we got super excited about it. We licensed the clinical methodology exclusively for our company. So that’s actually one of the things that we deliver through our AI system is dual steps of conversations. So what we are building, as Daniel mentioned, is a system that has two parts. One is that the patient interacts with the clinician on video. That’s not as frequent, could be once every two weeks or once every month, but every day of the week, the patient is interacting with AI.

We still have the human expert in charge, but by using AI tools that are augmenting and amplifying the human expert, we can do 20 times more. So there is the patient-clinician interaction, there’s a patient-AI interaction, and very importantly for us, and that’s a lot of technology that Daniel is building, there’s also the AI clinician loop of informing the clinician and the clinician controlling the AI.

Tim: Great blend of empowering the human to do more, the therapist, and using AI and technology to create a great patient experience, great user experience. And everybody has played around or done more than play around with ChatGPT. It’s incredible, but it can also get off the rails. There are a bunch of problems with it. This is a lot more than just slapping a voice front end on ChatGPT in the background. Maybe talk about some of the things that were hard problems to go from, yes, you can go interact with ChatGPT to having a delightful and reliable exercise and solution you’ve built.

Making Conversational AI Reliable

Dan: These days, I really kind of break down voice-based interactions with an LLM almost into two categories. There’s almost like verbal interactions with an LLM, and then there’s conversational AI. And when I think about verbal interactions with an LLM, it’s almost like I have a query in mind. I have an objective of why I’m coming to this LLM, and I know the result that I’m looking for. And sure, I’m choosing to use my voice to interact with it, but I’m really going to rate that experience based on whether or not you gave me the information that I was looking for in a relatively simple way.

When it comes to conversational AI, I think it’s just a much different problem overall. Users will come to it without a particular objective in mind, and a lot of the problem shifts into this space of having an engaging conversation. And part of it, there is really the semantics of the conversation, and there’s also the mechanics of the conversation. And so the semantics of the conversation is really the content of that conversation, and is the LLM responding in an engaging and interesting way?

I think LLMs are reasonable at this today. A lot of times, though, they can become quite repetitive, and this is in how they’re trained, and it’s also in that the conversation history can kind of shoot the model into responding with the same patterns every single time that it responds. I don’t think that’s representative of what a true conversation is like. I think there’s a lot of variety in conversation. I think that when you have interactions with an LLM where the content is kind of formulaic and repetitive, you start to lose a lot of those engaging characteristics that would make you want to have a long-form conversation with the system overall.

The other piece is the mechanics of the conversation. This is kind of how things are said. I think that there are just a lot more of both of those components involved. When it comes to conversational AI. The content is really important, the variety of the conversation is really important, the mechanics of the conversation are really important, and all of that is just a very different scenario than sitting down and bringing a query to ChatGPT, but just choosing your voice to interact with ChatGPT.

Tim: Very complex system to make it work reliably and really human-like. Babak, you mentioned earlier that there are a lot of clinical trials that underpin that this type of therapy works. You’ve actually licensed some of that research, and then that gets baked into the product. Maybe explain a little bit more about, it’s both interesting technically, how you take sort of this private data and incorporate it into this type of system, but it’s also just super important about the fidelity of what we’re doing. It’s not just having random conversations. It’s actually underpinned with the type of approach that a therapist would take with you.

Personalization, Memory, and the Clinician-in-the-Loop

Babak: So I was at Amazon when we launched Alexa. Daniel and I were at Google when we built the voice interface for Google, which started with, “Okay, Google.” Well, at first was, “Okay, glass,” and then became, “Okay, Google,” and everyone used it on their phones. But, so that the state of the conversation with an AI system up to very recent times was, “What’s the time?” It’s 8:30.” That was the end of the conversation. So multi-turn was very difficult. Now with the advent of the new LLMs, especially ChatGPT, that is widely available, you can have a multi-turn conversation with an LLM, but these are not really optimized to hold a meaningful half-hour-long or longer conversation.

So one of the core technologies that Dan has built is really to enable a long form, half an hour or longer conversation with an AI system that feels natural and engaging. This actually is extremely difficult. It’s done with many agents and many models under the hood. So the technical implementation is quite complex, but we’re at the point that we can actually hold long-form conversations, unlike any other LLM system, which is not really optimized for this purpose. The other point that’s important is that we need to get to know the person and personalize it as we have these conversations, because these are meant to be daily.

So, Daniel and the team have actually had to build a very specific type of memory for the system that begins to learn about the individual, which is very different from a generic memory. Especially for the population that we are dealing with, because as an example, we may hear from the person that, “I don’t have a brother.” Two weeks later the person might say, “My brother said X or Y.” So then the question is what would the memory do about this fact? Because the previous fact that we stored about this person to personalize is that our knowledge is that this person doesn’t have a brother. Now they’re saying their brother is saying something. Is this because of their dementia or is there something else going on?

So the memory actually that’s built for this system that’s optimized for the population that we’re serving is also very specific type of memory that gets to know the user and it sort of reflects on what the conversation is doing. So that’s the second big difference with something very general like ChatGPT. So long form conversations, very specific type of memory for personalization for this population. The third part is that we are having very frequent conversations with our users. We need to report back to the clinician what’s going on in these conversations.

So periodically, our AI system publishes a report to the clinician of what is really happening in these conversations. Is it something that they need to pay particular attention to? And that’s actually something that no GPT system at the moment has. And then our clinicians also give guidance back to the AI of what kind of conversations to have. So at the moment, this is the only system that we know of that is under direct supervision of a clinician. And what our team had to build the ability to deliver these conversations while maintaining the particular therapeutic reason behind these conversations. So those are the things that we are doing right now. Super exciting.

Tim: If you put it in the perspective of the company, and what are the moats that you’re creating, or a little bit, what was the investment thesis? You guys are an amazing team. This is a really big, important market. You’re building really hard tech that somebody else can’t go build. You have proprietary data that’s getting incorporated into it.

And then there’s this human piece about having an actual, the telehealth part also, we decided that was really important and that’s a big lift. You’re operating in three states right now. You’re seeing patients in those. Maybe say a little bit either about other parts of the what’s hard moats or just like why the heck do this human part of it? Why not just launch this AI app out there? That sounds easier.

Why Humans Stay in the Loop: Care, Safety, and Trust

Babak: The human part is really important. This is actually a major part of our thesis that these types of interventions, for the foreseeable future, they should be supervised by human experts. So we don’t want to just run the AI open loop without supervision. We’re always going to maintain the human supervision for this. It gives us confidence that we’re doing the right thing. I think it also gives our users have more confidence that-

Tim: Big time.

Babak: Yeah, the right thing is being done for them. So there’s a pilot in this plane, so it’s not an autopilot.

Dan: I think I see what there’s this saying within the healthcare industry for startups that services save lives and software improves efficiency. And so I think when you’re in healthcare and you’re dealing with someone’s cognitive decline and they have now been diagnosed with that condition, just having a person to speak to about that condition is amazingly and powerful versus, “Hey, I’m just going to go interact with a piece of software every day without any sort of oversight from another person,” I think would just never be able to get off the ground in the same way.

Tim: You all, the product’s live so people can go try it. You can tell others in your life to go try it. Newdays.ai. You built an incredible amount in a short period of time, you’re barely six months into this, and the product’s live. This isn’t the first time you’ve done that. You’ve actually gone from an empty room to astartup before. We’ve referenced Google and different things here in the conversation. Maybe back up, talk a little bit about how did you all decide to work together? What was your history before that’s made you say, “Yeah, let’s go jump into this big problem and do it together as co-founders”?

Dan: Babak and I have known each other for 17 years now. It’s been a while. We’ve worked together a bunch of times. That part of it is absolutely fantastic. I think that the reasons to go do this company, we saw cognition as this huge problem that we should try and go make a difference in that space. We saw LLMs and generative AI as this transformative technology that could be very beneficial there. We both have, Babak has family histories of the disease, I have family histories of the disease. It kind of made a very personal connection for both of us as well to something that we would be just invested in working in because we know how much of a problem this can be for patients and families within the space.

And then, yeah, I think that as far as co-founders, there are aspects of both working at Google, Google X and Amazon Grand challenges. I mean these were very much startups within these giant companies anyway, so in some ways Babak and I have worked together within startups for a long time now, but I can also say that doing a startup and co-founding a startup with somebody that you know so well and you just trust 100% is fantastic because there’s so many things to go do, there are so many problems to solve.

And that if you just know that the other person is capable of getting those things done, it just makes the whole process that much more seamless because you trust each other and you just say, “All right, I know that these are the problems you’re working on and these are the problems I’m working on and we can chat about those problems, but I also 100% trust that you’re getting these things done,” and that makes… Startups are always a journey, but that makes the journey just a little bit easier.

Babak: Plus one to everything that Dan said. I would say both of us are highly motivated, not to do something that’s cool and interesting, but to do something that’s meaningful. So something that’s not a niche product, something that’s genuinely good but also can help millions of people. That really highly motivates both of us,, and we both like velocity of execution because we’ve built a number of things from the ground up from robotic surgery to Google Glass to e-commerce services to human-machine interfaces.

So many things, many times from really the grounds up, like an empty room with nothing. And we’ve built these products, launched them. Some of them have been quite successful, some of them have been less successful, but we’re not afraid of building from the-

Tim: Like all startups.

Babak:Yeah, So we are not afraid of building from the ground up. We love velocity, and we love working on something that’s meaningful. I also have to add that we want to make this a successful business. Because what we’ve learned, I actually do believe in philanthropy and I think in our personal capacities we are engaged in that, but I just put that aside. In order to truly actually change the world with a solution that does good for people, we need to make it financially viable. That’s the only way for something to scale. So we’re building this, even though our intentions are as we mentioned, is really to do something good for the world first and foremost, but we know that we need to make this a successful business in order to have the impact that we want to have. So that’s equally important for us because otherwise it’s not going to have the impact.

Tim: In these early days, I think you’re really focused on initial users getting lots of people to try and experience the product and we’ll come back to that aspect of it, but I think in healthcare, there’s always this, how do you get distribution? There’s multiple ways. It’s early days, but maybe how are you thinking in a really customer-centric, user-centric way about ultimately how do you go to market to use that startup term here for this type of service?

Babak: Yeah. I would say our GTM is still a work in progress. We haven’t exactly nailed it.

Tim: Absolutely.

Babak: We have two things that we are doing. So we’re doing D2C, so we’re trying to remove all barriers for people to access us as fast as possible and with a minimal amount of effort. So there’s a D2C flow to us operating and there’s partnerships that we’ve had a number of conversations with different entities for partnerships that would not be of a D2C type. It would more be of some form of enterprise. So they’re both in progress. We’re trying to figure out which one is the best way.

We’re pursuing both at the same time because they’ve both shown promises and they’ve both shown challenges for us. So we’re learning and we are experimenting with. What I would say is that something that has been highly motivating for us is we recently exited our beta, is the feedback from our beta users. Because as you know, a startup has lots of ups and lots of downs, and every time we get feedback from our beta users is really having a shot of espresso for the whole team because we actually see that we’re helping people, and this has been really meaningful for us. Just in terms of-

Tim: Let me break in on that because I want to talk about this. I mean, Amazon’s famous for work backward. From the customer, I’ve been super impressed by how you’ve gone about it, not just like, of course, we need to go get user feedback, but you’ve had a real process leading into the beta and now the GA is going to lots of different users. Maybe just talk about how you’ve approached that around being really customer driven and getting feedback on how the product’s working as well as any aha moments from hearing this feedback from end users.

Babak: I would say that’s both. Daniel and I have spent a number of years at Amazon, and that’s something that Jeff is amazing at drilling into people, to be customer-obsessed, and that’s really a superpower of Amazon, and that’s learning directly from Jeff of how to be like that. So we tried to.

Tim: I think it’s been a superpower for NewDays too.

Babak: Yeah. So we’ve tried to bring basically that part of the Amazon culture into the company and some of the mechanisms that Andy and Jeff have built at Amazon. But yes, from day one we have been very customer and user obsessed and maniacally basically ran user studies, ran beta tests, measured CSAT, customer satisfaction score, to make sure that we are building something good for people and people actually like it. And I’m happy to report as of now our, CSAT is really high, so it’s measured in the scale up to five. For our AI exercises, our CSAT at the moment is 4.55, which is very high. For our full clinical service is 4.8.

Tim: Wow.

Babak: So people love this service. So that’s the numerical measurement but also the anecdotal, what we are hearing from people is that they are actually seeing really meaningful results. So that has been extra motivating, an extra motivating factor for us to continue. And encourage us to exit beta and make the service more widely available.

Dan: I do agree, as MCI and dementia progress, people tend to socially withdraw a little bit. They don’t seek out social interaction the way that they used to. And there was one participant that was engaging with the exercises and he used to be this really outgoing, gregarious individual and has definitely pulled away from that, but as soon as he started interacting with the exercises, all of a sudden that aspect of his personality came out all over again and his care partner was there with him as he was interacting with this system and her mouth just hit the floor because she had not seen this aspect of her husband in a very long time.

And was almost like these exercises are just kind of a safe space. They can interact with them without all of the social pressure or other potential drawbacks to the kind of fear and worry that come from interacting with other people. And just hearing that sort of feedback and seeing those sorts of things is incredibly empowering.

Tim: So powerful. For me, just thinking about it, it’s like, okay, is someone going to really want to have this in-depth conversation with an application? And not only have you seen they do, this idea of maybe you’re self-conscious in groups of people and unsure of yourself a bit, but that hey, there’s no downside here and then that kind of really builds on itself that you’ve seen that multiple times with your early users.

Dan: Absolutely. I think that there is a certain amount of interacting with another person where there’s some amount of concern that you may repeat yourself or you may not just come across as the person that you think of yourself as. And so that may lead to the decision to just refrain from the conversation and not participate in the same way you would’ve. But when you have a system like this, there’s no judgment there. You can make mistakes. It’s the place that you’re supposed to practice in order to continue to have those interactions with the other people in your life.

Babak: Again, to complement what Daniel was saying, we’ve heard this multiple times now from our users that they feel like this is a no-judgment space. So when they talk to Sunny, our AI character is called Sunny, they feel like they could just talk and they’re not worried about being judged. And we love that because we want them to have these conversations because we kind of know the therapeutic results there, but we want to encourage them also to socialize more.

So it’s not really meant to be a substitute for socializing with other people. It’s meant to be a safe place to practice and get better, so they feel more confident, so they can actually socialize more. So the more they socialize, the better, and this is a place for them to practice and build confidence, and hopefully they feel overall better about themselves and their abilities.

Dan: The other place where we’ve heard really great feedback is from the care partner as well, and that the person with MCI or dementia can sit and have a conversation with our system and it’s almost half an hour to 45 minutes of respite care for the care partner where they get a break and they don’t have to be so involved in the care and management of their loved one.

Tim: You mentioned earlier you’re building a business, people might be wondering, it’s early days and this point of, which I’ve really encouraged too, is to reduce all friction to getting users to experience this thing, wanting to come back, and then those positive references I think build on the system. Eventually, you’ve got to make money. It’s early days, you’re experimenting, but how do you think about pricing a product like this? Some days as an insurance could pay for type thing? How do you think about that aspect of the business?

Babak: So we have two components now. There’s access to AI and all the exercises overall, the AI platform. And then the other one is the clinical visits. So we already accept insurance for our clinical visits. Then that really helps lower the cost, the actual cost to people that participate in the program. The AI access part is out of pocket and we have a subscription for that. There’s a cost associated with that, at the moment is priced at $99 per month, which appears to be fairly affordable for the population that we are targeting. And every day actually you can go in and do another free conversational exercise and then if you’d like, at some point you can upgrade to become a full member of the clinic. But we’ve definitely tried to make it easy for people to access the system.

Dan: Maybe one thing I’ll add on to that is that an additional feature that we think is really important that we’re building towards now is just providing people feedback about their cognitive health. We have to do this. There’s a lot of UX associated with providing feedback, but I think one benefit to the feedback is, or maybe my general impression is that I’m not totally convinced people just want to go sit and have a friendly conversation with an AI for no particular reason. It’s not like you’re walking down the street and you see a stranger and you just stop and talk to them for 45 minutes for no reason.

I think that there is, you want to know what is the motivation or justification for me participating in this particular program. And I think one thing that’s top of mind for all of our users is, is this helping? How am I doing? Et cetera. Feedback along those lines? And so if we can provide feedback, it’s just another piece of the puzzle as far as going from casual conversations to complex conversations, all of that overseen by the memory and the things that we’re learning about you to make it very particular to your cognition and your life experiences. But then also using all of that information to provide you feedback to continue to give you justification for coming back to these exercises to understand that process overall.

Tim: I mentioned how fast you’ve built the service and the company and you’d never do that without a great team. And part of being able to move fast was that you hired a great initial team quickly, and hiring is hard. Maybe talk a little bit about how you were able to do that and the type of culture overall you’ve tried to instill in NewDays here from the inception.

Dan: I mean, we definitely leveraged as much of our personal networks as we possibly could in order to find people that were interested in this space. It’s nice having-

Tim: The mission’s important though.

Dan: Yeah, the mission is kind of one of the big selling points for coming to work for this company, is that if you are the type of person that wants to take your technical expertise or your business expertise and apply it towards making a real difference in someone’s life, then this is just a fantastic company for you to come look at.

And so yeah, leveraged personal networks to the greatest extent possible. And have also just been super scrappy about finding fantastic talent to help us go build things as fast as we possibly can. And we leverage, we have people here in Seattle, we have people in New York, we have people in Argentina, we have people in Indonesia, and they have all just been… Every day I’m impressed with the work that they’re getting done. And so yeah, I think it’s just a fantastic team so far.

Babak: I think that’s an accurate statement that every single person that we tried to recruit, we actually successfully recruited. So that has been our track record. And hopefully we can-

Tim: Knock on wood.

Babak: Knock on wood. Hopefully we can maintain that. But pretty much everyone that we wanted to get, we got. Because I think part of it is because they really found not just the technology exciting because this is obviously cutting edge technology and AI, but the mission is meaningful so that both components to work on cutting edge technology, which is the hottest technology of the day, but use it to do something that’s very meaningful and humanly relevant. I think that the combination of the two has been a good combination. As Daniel mentioned, our team is quite distributed. We try to fly people in with the regular cadence, so we physically actually get together with the regular cadence and that has been helpful in building the company culture.

And that’s something that I got to confess that that’s different post-COVID because I guarantee that before COVID, if Dan and I were building this team, we would’ve insisted everyone to be in the same building in Seattle. This is different. So the post-COVID world is different. We are learning how to operate in a more distributed way. So this has been in a sense liberating because we can recruit from all across the US and globally. As I said, we have people in three different continents right now, but we also, we have to be even more intentional about building the culture of the company to build trust and build velocity. And part of it still relies on flying people in to spend some time physically together every month.

Dan: I agree that it’s the post-COVID world is just totally different. Babak and I chatted a lot about this and we both agreed, we’d been part of a hundred percent remote teams during COVID and agreed that that didn’t work. We’ve also been part of teams where everybody’s in the office five days a week and you almost get some of the most toxic environments, competition between different teams. And so it’s not just like being in the office all the time is the solution either. I think ultimately the solution is you have to put work into culture, the company culture, just the same as you have to put work into the technology you’re building or the business you’re building. And if you’re not doing that, then probably the culture’s going to get away from you.

Tim: You raised a nice round of funding early on to start the company. Madrona was, we were grateful to be able to invest in General Catalyst. Holly Maloney has been an amazing board member and investor, lots of healthcare experience, but still the amount of resources paled compared to doing this inside Amazon, doing this inside Google where you did start with an empty room, but there’s more resources there if you need them and can make the case. What other learnings generally, advice for other founders, how to get going, how to move fast, have you learned? Probably some things you’ve learned the hard way here, even in these first months.

Babak: If you get a no from a VC firm, do not necessarily get totally discouraged because that, premiere VC firms see a lot of pitches and the success rate in these firms is 1% or below. So that no doesn’t necessarily mean that your idea is bad or you’re a bad entrepreneur. That no might actually mean that the answer is no at that particular time, and maybe next year the answer is yes. Or there’s a lot of things going on inside the VC firms that might actually result in a director saying no. So I would say to an entrepreneur, if you hear a no, don’t get discouraged. So continue if you’re convinced that your idea is good.

But the other thing that I would recommend is to maintain execution velocity. Because if I think about the startup, Daniel and I have been also part of one of the biggest companies in the history of the world actually. Big companies have a lot more money, they have a lot more resources, they have a lot more people, and actually good and smart people. So they have everything going for them. They have more access to customers. Everyone they call will pick up the phone. The only thing that the startup has is the velocity.

So as a founder, if you find in a situation that your velocity is starting to stall, that’s a major red flag. So we have to have the velocity. That’s the only way actually to be able to be competitive in a place that you don’t have massive amount of money, you don’t have massive amount of infrastructure, people, all that. So I would say paying attention to velocity is super important and that’s part of the culture building in the company that we have to be intentional. The company actually incorporates velocity of execution as part of the core of the values of the company early on.

Dan: I was having this conversation with Trevor, our CTO the other day, and we were talking about moving fast. And I gave this example of there may be a problem that’s posed where someone says, “Which one of these two objects is bigger? The bus or the motorcycle?” And our engineer instincts are like, “Oh, I know how I can figure this out. Here’s my process description of how I’m going to do it, and I’ve calibrated all my measurement instruments, and I know exactly the experiment that I’ll do.” And then someone else walks up and just says, “Yeah, the bus is bigger,” and you move on.

And I think that when you’re doing a startup, there are a lot of times where you’re facing those sorts of problems. The bus versus the motorcycle, it’s just harder to recognize because you’re working in a brand new space. But there are just a lot of times when you almost have to turn down your engineer instincts a little bit just to move faster when the solution is actually obvious. And it would just help you to move faster and go figure those things out quicker rather than dialing up all of the engineer processes that you know how to do and have been beaten into you for a long time.

Tim: And then jump on the motorcycle.

Babak: So I would add also one thing that might be counterintuitive, but I would say a bad decision is better than no decision, because a bad decision would allow us to proceed and figure out the problems, and the decision was incorrect and course-correct and get to the right decision, but indecision and making no decision would basically torpedo the operation.

Tim: Well, thank you both so much. So excited about the future of NewDays. It’s early days, but the results point to a future where AI for Dementia complements clinicians and restores confidence. Appreciate the opportunity to work together and build this service. So congratulations and thanks.

Babak: Thanks so much for having us.

Dan: Thanks, Tim. This is great.

Learn how AI for Dementia supports patients and care partners at NewDays.ai.

The Future of Biology Is Generative: Inside Synthesize Bio’s RNA AI Model

 

In this episode of Founded & Funded, Madrona Investor Joe Horsman sits down with Jeff Leek and Rob Bradley, co-founders of Synthesize Bio, a foundation model company for biology that’s unlocking experiments researchers could never run in the lab.

Jeff, chief data officer at the Fred Hutch Cancer Center, and Rob, the McIlwain Endowed Chair in Data Science, share:

  • Why a startup the right fit for generative genomics
  • How generative genomics could reshape research, drug trials, and more
  • Why now is biology’s “ChatGPT moment”
  • What makes Synthesize a true foundation model for biology (not a point solution)

Whether you’re a founder, biotech innovator, or AI researcher, this is a must-listen conversation about the intersection of AI, biology, and the future of medicine.

Listen on Spotify, Apple, and Amazon | Watch on YouTube.


This transcript was automatically generated and edited for clarity.

Joe: Maybe we can start off from the beginning. What is the founding story of Synthesize? Take us back to the conversation that kicked off this company, and why we’re having this conversation today.

Jeff: Rob and I have known each other for a long time. We’ve been academic colleagues for probably 20 years and have followed each other’s science. I moved back to Seattle about three years ago, and as academic leaders, we were running into each other in the halls. This idea of building a foundation model for biology was something that was on both of our minds. We started talking about it, and there was just enough information out there and just enough of an idea out there that we felt like we might be able to take a crack at it.

And so we started thinking about it a little bit and then immediately emailed you and Chris here at Madrona and said, “We need to talk to you right away.” Because we sort of felt like this was the convergence of the moment in time when there was the right kind of data to make this possible and the right kind of technologies to make it possible. And both Rob and I were really excited about trying something, taking a big swing, and trying something cool.

Joe: Maybe to take a click up in elevation. Today, there’s this Cambrian explosion of AI and bio. And I’ll maybe put “AI” in quotes here. There are lots of people who are building things that are probably equally well done by linear regression. But as you point out, there have been some amazing advancements in things like protein design.

There’s a Nobel Prize; lots of companies are now using this to develop drugs or to develop drugs in collaboration with the biopharma ecosystem. But you’re doing something different at Synthesize. I’m curious why do something different, but also how are you thinking about where this fits into the broader ecosystem of everything from people in your labs getting their PhDs and masters, all the way through big pharma trying to find the next blockbuster drug?

Jeff: When we set out to do this project, we didn’t think about solving one particular biological context. There are certainly specific problems we could have tackled, and we both have in our academic careers, which are very highly contextual to a specific disease in a specific set of people. And our goal from the start was to try to do a model that would, in a similar way to large language models, have enabled many applications, we wanted to build models that would enable many different applications. And so I think that’s one difference of why.

One way in which we distinguish ourselves from a lot of different groups is that we’re trying to build this broad-based foundation model with lots of applications. And so whether the pharma company of interest is interested in cancer, or whether they’re interested in neuro, or they’re interested in cardiovascular disease, our model has capabilities in all of those areas. And so for us, it’s less about one specific target and really building that foundation that lots of people can use to accelerate most, if not all, of science across drug development. And so we’re excited about putting it in people’s hands and seeing how they can try it out and use it in a lot of different contexts.

Rob: I think there are a lot of reasons from computer science and machine learning literature to think that modeling the most diverse possible training data set will give the best results. It’s clear that with language models, this is the case. For a long time, people made really focused ML models, and they made progress. But it just turned out that taking the biggest corpus of text you could and then shoving it into a model and letting it model it.

Joe: Yeah. The bitter lesson.

Jeff: It’s pretty tough if you were working on that very specific application, then to have these foundation models come in and just do better on all of the specific applications.

Rob: It’s incredible. Right? I think very few people would’ve predicted this. That like, if your goal, for example, is to help people with legal contracts, maybe you’ll do better by starting with modeling the entire internet.

Jeff: It was not intuitive that that was going to be the solution. And I think probably the same for us.

Rob: Exactly.

Jeff: If you care a lot about this gene in a particular brain of a particular human, it’s not clear that modeling the entire corpus of gene expression data ever collected is the right way to solve that problem.

Rob: Exactly. I mean, I’ve built various ML models throughout my career in my highly specific scientific problem. And some have been useful, some have not. But I would say none of them have made me say, “I can do something I couldn’t do before.”

Joe: Can you explain how this fits into how we think of biology and maybe the central dogma biology that we all learned in high school?

Jeff: To recall your high school biology — everyone has the same DNA sequence in all of the cells in their body, more or less. But then there are little chunks of DNA sequence that code for something called genes. Those are encoded in RNA, and we measure those quantitatively. And so your heart cell is different than your brain cell because you have different abundances of these genes expression, which ultimately get translated into proteins and do functions in your cell.

And RNA is a molecule that’s much easier to measure than something like protein. So there’s a large abundance of measurements of this molecule. And so we, when we were thinking about this idea, we were really focused on where is the capacity right now to generate the kind of training data that would enable us to build a foundation model that really spanned a lot of different applications. And for us, that was RNA.

And it helps that both of us have a lot of experience scientifically, both of our careers are kind of built in that area. And so that was where we started to focused at the beginning. But the nice thing about an RNA molecule is it’s dynamic. It responds to environmental stimulus, it responds to drugs, it responds to what you eat in the day. And so you actually get this readout of biology from this molecule. And so if we can model it, if we can generate data that looks like realistic data from humans, we actually get a real window into the biology of what’s happening in those humans.

Joe: Totally. Like real-time biology?

Jeff: Yeah, exactly.

Joe: And so this is an RNA model that you’re building. Where does this fit into your everyday experiments and the biology you’re trying to unravel at the end of the day?

Rob: I think it’s helpful to always think about analogies with other AI models. At this point, we’re all pretty familiar with large language models, right? And these are generative AI models in the sense that you ask them to do something and then they generate something for you. So a large language model, you give it a prompt, you ask it to do something, and then it creates a bunch of text. And there’s something that’s really useful, lots of people we communicate that way.

We write lots of text. So we wanted to do something like that for biology. What scientists do; what biologists do most of the time is, we generate data then we analyze the data and we do experiments. We might run a clinical trial, and then we get the data and analyze it. And so we wanted to model it was the analogy of an LLM, but for what biologists actually do every day, which is to generate data to do an experiment, get the data to inform the next one.

Joe: And so in building this platform, how do you think about the problems that need to be solved? What can’t be done today that is going to be possible in this new world of Synthesize?

Rob: I think about my own personal scientific experience as a person working at the computer, at the bench, running a lab, et cetera, where I, and people like me, people like Jeff, we’re constantly faced with the need to make decisions when we don’t have enough data. And sometimes that’s because we just don’t have time to get the data. Sometimes it’s because it would cost so much money that it’s not possible, but a lot of the time it’s because the data is just not reachable.

There’s no way we can get it. Now imagine, for example, somebody’s developing a drug to treat neurodegeneration and that drug acts on cells in the brain. There’s no way that you’re going to be able to look inside and see what’s happening in the cells in a patient’s brain who’s taking this drug. But we need this information to make a decision. And so scientists are constantly faced with this impossible task. We have to make decisions. Do we proceed with the drug development when we just can’t get the data that we need? And so we wanted to build a model that would let us get these data.

Joe:. I quite like when all my brain matter stays in my brain. So-

Jeff: That’s the right place for it to stay. Yeah. Exactly.

Rob: Even if there are ethical challenges, you might not want to participate in this experiment.

Jeff: That doesn’t make it less important. It’s such a critical piece of understanding how a drug might function. And so a lot of these experiments that we want to do — an example from my background is the very first data I ever analyzed came from a study where they randomized patients to get either endotoxin or saline solution. And endotoxin is this horrible thing to get where you get really, really sick if you get it. And so people were randomized, so they had to kind of wait and see whether they were in the control group or the get sick group.

But they could only do it on a very small number of people. And they were trying to study the genomics of blunt force trauma. And so this is something that’s pretty hard because you can’t put people in car crashes. So we could do it only at this very small scale where it was just a few people that got randomized to get this really bad intervention. But if you can do that experiment in a model instead of doing it in a human, we could do it for hundreds of people, thousands of people.

There are no constraints around what the experiments you can do. And similarly, if you can sample people’s brains, if you can sample all the tissues in their body, you can get a much more comprehensive view on what’s happening in response to disease, in response to trauma, in response to treatment. And you can do it at a scale and a speed that’s sort of really hard to pull off in a traditional laboratory experiment.

Joe: So I think, as you’re right to point out, this is solving an impossible problems. These are inaccessible samples. These are experiments that are unethical to run. Was there a light bulb moment for you where you’re like, “Wait a minute, I think I can actually simulate these things. It’s maybe not intuitive to me that this would actually work.”

Rob: That’s an important question and something that we thought about a lot. We’re both scientists, I think scientists spend decades being trained to be

Joe:Skeptical by nature?

Rob: very skeptical. Very rigorous. So I mean, we really started out actually by thinking not about how do we make the best model possible, but if we had a model, how do we test if it’s working, what are the ways that we can assess is this doing something useful? Is it producing data that’s meaningful that would actually be useful to people like us? And we actually played this game, this is Jeff’s idea, which is a great idea, which is we started taking the model and simulating an experiment.

We take the model, we would have it generate data for an experiment, and then we put that data next to data from a lab from a parallel experiment. So data either from a scientist doing an experiment in cell culture for example, in the petri dishes or our AI model doing the same thing, but of course in seconds or minutes instead of weeks or months. And then put those data together. And then Jeff would send me these data and say, “Can you tell which is from a lab and which is from the AI data?” And I’ll say for a long time, it would take one or two seconds to say “That’s the AI data.” But there came a time when I couldn’t tell.

Jeff: And the story within the company is this is reinforcement learning with Rob feedback. He was one of the best people at picking out which data set was the AI data, so and which data was the lab data. And sort of once we could get it past Rob, we were like, “Okay, we’re kind of onto something here. We’re at a point where these data are really looking like what you would get from an experiment. They’re sort of indistinguishable.” And a really important point that is actually Rob’s point is when you’re measuring machine learning models, you usually do them in bulk measurements.

You’re measuring their overall accuracy, root mean squared error, things like that. But when you’re measuring a biological foundation model, it’s about what one gene in one environment in one context. And so the way that you measure errors isn’t in this bulk context, but it’s like looking at this particular receptor in this particular tissue under these conditions. And so we would look at areas that were very specific to Rob’s research area and have him look for the exact sort of genes that should be turned on and turned off in the right context. And if you’re savvy about this, you’ll be able to detect them pretty quickly if the models aren’t really accurately describing the whole distribution of what’s going on.

Rob: And I think this is a really important point, and it’s both a challenge and an opportunity, right? The challenge is that in order to build an AI model, train it, do inference, all these kinds of things, you need these aggregate statistics that describe how well the model is recapitulating the kind of whole shape of your data. That’s how you train a model. That’s how you assess it. But at the same time, exactly like Jeff was saying, much of biology, maybe even almost all of biology is about highly specific things. Right?

I’m an RNA biologist, but what I really know about is a couple of genes, and I know about how those genes interact. They make proteins interact with a couple of other proteins and this is my area of expertise. And the same goes for most other biologists and the same is true for drugs. Drugs tend to have ideally a few specific targets that they act on. The same is true for physicians treating patients; they specialize in specific areas.

And so we had to kind of merge these two goals, right? To have a representation of the whole shape of biology, of gene expression data, all these experiments that people have done while also making sure that we captured all the fine details. Because I can tell you that for me as a scientist, if somebody comes to me with a machine learning model and says, “This represents all the data really well. Look at my statistics” and then I look at the one gene that I’m an expert in and it doesn’t seem to understand what that gene is up to, this model is not useful to me.

Joe: That’s exactly what I did the first time you shipped me over the model. I was like, “I’m going to plug in the experiments I know.”

Jeff: I remember that. You sent us back exactly your area of expertise.

Joe: Here’s the genes I want to see.

Jeff: Exactly, these are the genes I want to see.

Rob: It’s exactly right. I mean, maybe it’s like if you have a large language model and it produces text, looks pretty good, but there’s four words that it always misspells. We, as users of language, are going to notice this and fixate on it.

Joe: So I want to come back to the data, but first I have to ask why a company? Right? You both are professors, this is your bread and butter of you could build this, get some amazing papers out. Why do this inside of a company as opposed to just, “Hey, my lab now does synthesize.”

Jeff: That goes back to the email I sent you and Chris right at the start was this was such a cool idea, and we wanted to get going on it immediately. We didn’t want to have to wait. And I mean, there are many amazing things about being an academic researcher, but being able to capitalize on a big swing idea on a very short time scale is a hard thing to do just the way the systems are set up. And so we wanted to move really fast and we wanted to go really big and it felt like the best context to be able to do that was as a company. I feel like that was what drove a lot of our interest in moving this way. What do you think?

Rob: I totally agree. Velocity and scale. We sent you this email, we had some conversations and then we were going. We were doing the things. And that’s exactly what this needs. And I think the second point is scale. We want to build something. Like Jeff mentioned, I mean we’ve done a lot of things in our career and it’s been really awesome, but we want to do something that’s going to affect a lot of scientists, maybe all scientists in biology. And to do that we need scale. We didn’t want to model just a couple of gene expression experiments. We wanted to model everyone we could access.

Joe: Yeah. I think this rhymes with a lot of what we see on the tech side where there is a moment right now. Do you think when we’re in the biology realm, is there an inflection point specific for biology. And kind of a two-parter here. Has life sciences and biopharma had this ChatGPT moment or is that still around the corner? Have we truly had the AHA as a field for this?

Jeff: So to answer kind of both questions at once, I would say I don’t feel like we’ve really had our ChatGPT moment in the sense that there haven’t been a lot of these models that have been deployed in a way that anybody could use them, access them, and build on top of them. Even people who are building sort of something that would be akin to a foundation model have tend to do them inside of a single company and not share them with other groups. And so I think there haven’t been as many swings at these sort of foundation models that anybody can use except for in one space where the protein space feels like there has been some of that work, like the protein structure, protein design space.

So some of our colleagues work in that area, and they’ve been based in a similar way on open datasets that then they built models on top of. And we’ve seen the sort of explosion of interest as people have made those available. And so I think we feel the same thing is possible in all the downstream consequences of biology past those sort of drug target identification with proteins and things like that. And so excited about really contributing to that. And I think that moment is coming though because there is the availability of these huge collections of data that’ve been supported by federal funding and lots of other organizations, and now there’s an opportunity to capitalize on really doing the same kinds of things that were done with large language models. And that’s certainly been our approach to this problem.

Rob: I think there are even closer analogies to be made with large language models. The thing that I find so inspirational about these large language models that we have now is not just that they can do things that I can do really well. Right? I mean, it’s cool that they can write text, this is very useful to me, but they can do things that I can’t do and I never can do. Right? They can translate between any two languages instantly. They can program way faster than any human ever can.

They can do these things that are just beyond human capabilities right now and are probably never going to be within the scope of human capabilities as we understand them. And this kind of by analogy, what I find equally inspirational is — it’s amazing that the protein structure problem has, in many ways, been at least partially solved. I think that’s incredible. But what I find truly inspirational is protein design, making novel proteins that didn’t exist before, that don’t occur in nature. And I think we can do the same thing in other areas of biology. That’s what we’re trying to do here for gene expression.

Joe: That comes back to the data where there’s no internet to scrape for biology. I mean, there are lots of papers out there, but it’s messy. Where do you think the field needs to go in data? Why do you think there is data sufficient to build what you’re doing at Synthesize and what is the foundation on the data that you’ve been working for the past over a year at this point to get to a generative model for biology?

Jeff: I think this is where picking the right molecule is so important. Gene expression data is amongst the molecules that you see in the central dogma biology, the one that’s sort of measured in the most conditions and been studied in the most context. And so there was a real opportunity to capitalize on the fact that the field in general has measured the experience of humans in a variety of different contexts and measured their RNA. And so while there isn’t an internet to scrape, there are a huge collection of existing experiments. The big challenge though is that it’s using other people’s data in the sense that there’s other experiments that have been done.

They aren’t normalized and synthesized to be worked together in one specific context. And so it’s a huge amount of both intellectual work and engineering work to bring all these data sets together and set them up in such a way that you can actually train a model on top of them. And so we’ve been capitalizing on the sort of ability of our team to bring together a large collection of data sets, the expertise that they have in really synthesizing and normalizing the metadata so that the descriptions of those experiments are common and unified across thousands and thousands of human experiments so that we can build models that understand the contextual representation of gene expression across different conditions, across different tissues, across different treatments.

Rob: I think here we really have to give a shout out to Jeff who saw maybe not the exact use of training generative AI models, but certainly the importance and potential of creating harmonized standardized data sets a long time ago.

Jeff: So my lab has been doing that for a while and I didn’t realize it was going to be a training set when we started. We were normalizing and synthesizing data largely for reproducibility, for helping scientists do their work. And that work sort of was where we built our original prototype, was on those data that I had developed in an academic context.

We were able to use those to build our first prototype of the model that we showed to you when we got together. Ultimately, our team has gone wildly beyond where we were when we started with that data set, especially on the side of sort of normalizing and harmonizing the sort of descriptions of the experiments. But yeah, that was sort of part of the reason why we had the AHA moment is we had already been thinking about these big collections of data that we were using in an academic context.

Rob: One of the things that’s been interesting about building this big proprietary data set where we have pair gene expression experiments and then this highly curated metadata we put together is we can see really unexpected things. So one thing that we noticed is that a surprising fraction of experiments are closely related to ones that our model understands well from seeing in the training data. So the way we, this is kind of getting into the weeds, if you’ll bear with me, but we really were into scientists skeptical, et cetera. We wanted to validate our model and so we thought, “Okay, we want to predict future gene expression experiments, so let’s validate it doing that.”

So we picked a date and that was our training data cutoff and all data that is in the public domain generated before that date we trained on and everything that was generated subsequently by scientists and then deposited in public archives, we called our validation set or test set. So we never looked at that data, right? It was totally holdout data and those were future experiments for the purposes of our model. And we can go into how well our model did on that data, which I think did very well and surprised all of us. Really kind of amazing.

But the point I was going to make was about metadata. And one thing that’s really interesting is because we created all the metadata, we could look at statistics aggregated across experiments. And one interesting thing to note is that approximately 95% of all experiments conducted in the future after our training data cutoff date were either in biological context like say primary tissues or cell lines or involve specific chemical perturbations like small molecule drugs or biologic compounds or gene knockdowns, perturbations to genes with CRISPR, et cetera, that we’d seen before. So the key point is that 95% of all future experiments were in a domain that were very close to our training data where we had extremely strong reasons to believe our model not just might perform well but should perform extremely well.

Joe: So can you maybe give a specific example of how you would use either in your labs or someone in your labs or someone in the biopharma ecosystem would use this? So I think putting a concrete example together would be useful for people.

Jeff: I’ll say from my lab, so my background is in biostatistics, that’s where I got my PhD, and so I end up helping people design studies all the time, whether those are clinical trials or preclinical studies or just research studies. And in all of those, you have to figure out what sample size to collect, which population to look at, how do you sample. There’s a wide variety of questions, and usually, you’re just making it up. You’re trying to figure out what it might be, and you’re gambling on a lot of things that you’re making a lot of assumptions about. And so now, we don’t have to make those assumptions. We can just generate the data from all these different circumstances and then we can pick the design that’s going to maximize our chance of success. So I think this is going to accelerate a lot of things where you have to design those studies in advance, and we can now kind of get a sneak preview of what the study’s going to look like before we ever do it, which we couldn’t do before. So it’s really exciting.

Rob: I’m going to give an answer that requires really going into the weeds.

Joe: I love getting in the weeds.

Rob: So one of the things that we’ve done that is super technical but also super exciting is that we have developed a model that lets you not just specify an experiment and then generate the data that results, but also add in data from a lab or a clinical sample and then see what might happen if you modify it. So what that might actually look like for example, is you could take an experimental description like say a sample of a particular cancer treated with a drug of interest and then simulate the gene expression result. That’s something our model can do that we’ve been talking about. The new thing, this in the weeds technical thing I’m talking about is that we can also give it information about what that sample might look like without the drug.

And this is super interesting because that information could, for example, come from a biopsy of a patient who has an active cancer and we’re trying to figure out which treatment course is best so we can give that information to the model and then it can give a patient specific prediction about the effect of the drug. And that is what I see as totally transformative. As you know, there’s lots of academics, lots of companies who are trying to do precision medicine and these efforts are incredibly important, incredibly exciting. But I think our contribution to that is going to be that we can have this huge model that can model essentially anything, but that we can also tailor the results to one person, to one sample.

Jeff: I think that kind of speaks to generally the approach that we’ve taken, which is we want to build these big foundation models that allow you to tackle many different applications. The two things we just talked about, if you go to talk to scientists and say, “Are these two things related to each other?” They’re super, super far apart.

They’re totally different academic disciplines. You would be talking to totally different humans. But our sort of underlying foundation model enables both of these kinds of applications and many others. And so if you ask me what I’m most excited about, it’s actually the things neither of us has even thought about yet.

It’s sort of what the grad student who’s staying up late and trying to figure out what problem, how to solve their problem that tries to apply this and can move forward a whole field that didn’t have answers before. And we don’t know what all of those applications are, but I think that’s so exciting about this to me is sure, we can come up with our ideas, but I’m excited to see what other people come up with.

Joe: On that, I guess let’s fast-forward 10 years. What do you envision the new lab is going to look like with tools like you’re building in the hands of every scientist having this out there for anyone to use? Where do you see this being game changing? How does it impact not only the day to day but I’m going to call it the year to year for the biopharma ecosystem?

Rob: Well, we’re really different scientists, so maybe we can just give our own answers. I would say for me the future I’m excited about that I want to help build is where there’s a seamless blend of what we’re calling generative genomics and wet lab experimentation in clinical trials. I think there needs to just be constant flow of ideas, data, everything between these different areas that a scientist in 10 years or hopefully we’re trying to move quickly, two or three years can do something like simulate a drug screen using our model in an hour in their computer.

Use that then to choose a cell line where they’re going to do an experiment in the wet lab. Get those results back and then use that to then inform, “Okay, we actually need data from this other system using the AI model,” et cetera. So there’s just this constant interplay between different sources of information and data.

Jeff: So my lab is largely computational and so I don’t have a wet lab like Rob does. So for me, it’s really empowering for all the students that work with me. We typically have to form collaborations with people like Rob to generate these data and sometimes that’s the only way to do a kind of experiment and sometimes there are creative ideas that we just can’t find the right collaborator for. And so this is empowering students and postdocs and site research scientists to try experiments that they would maybe not have the right collaborator for or maybe not have the right funding for, be able to go pursue that kind of wild idea that would be tough to pursue otherwise.

So I think I’m really excited about that enablement. The second thing that I think about a lot is how many bets we make. We have to bet in people’s time, in resources, and those bets are often made on relatively little information. So you sort of read the literature, you know what your friends are working on, and you’re sort of gambling that this next idea is going to be the right idea and that the experiment’s going to work out.

And as a person who collaborates with lots of different labs, I’ve seen firsthand how many of those bets don’t pay off and what the consequences are for science and both speed and the people that are actually working on those projects. It can have a huge impact on their careers and their lives. And so I’m really excited about increasing the win probability on every science bet that people have to take. They get to see a little bit in advance, they get to make better bets. Even if we increase that by 15, 20, 30 percent, that’s a lot of resources, that’s a lot of speed that we’re buying for the whole field.

Rob: Our vision is that what we’re building can be used throughout the research chain and the drug development value chain. The examples we just gave are of basic research or maybe translational research, but I hope we aspire for our models to be equally useful in a clinical setting. Earlier on I was talking about how scientists have to make these impossible decisions. Right? You have a certain amount of data, you’re not going to get any more. It’s not enough to know, but you got to make a call. And I think a great example is with clinical trials. Like phase one trials are designed to test drug safety. That is their purpose.

Nonetheless, if you can get any information on efficacy, that is going to be very useful. And so people find themselves in a situation where they’re using a trial that was not powered to make efficacy statements and you’re kind of reading the tea leaves on are there efficacy signals? Is this going to inform the decision that we make about moving forward? Which is a very important decision. Because if you move with one trial, you probably won’t move forward with another one, right?

It’s not just about that one drug, it’s about all your shots on goal. You could imagine, and we’re actively working on this, doing things like taking our model, conditioning. This is this reference conditioning, in the weeds thing I was talking about earlier, conditioning on the results from your small, say N equals 12 patients phase one trial, and then inferring what the results would be like if you’d run a fully powered trial with hundreds of patients. Now, of course, this isn’t the same as running a trial that costs a hundred million dollars, but it’s a lot more information than you had before. And it’ll-

Jeff: And cheaper too.

Rob: It’s a lot cheaper. And it’ll help you make a better decision.

Jeff: And if you’re going to make a $100 million bet, you might want to have some information before you make that $100 million bet.

Rob: Right? I mean, that would just be so useful to get more information to say, “I have a little bit more confidence now in either dropping this program or doubling down.”

Joe: Rob, Jeff, I really appreciate the time today. I want to leave it for you for the last minute. You’re building something for scientists that they can pick up today. Where can people go to learn more about Synthesize, to get access to your models, to start becoming the future of science where I can be empowered by a foundation model?

Jeff: First of all, thanks for having us. We really appreciate it and thanks for being such great supporters in general. It’s been amazing to work with you and Chris and Matt and the rest of the team here at Madrona. People can go to Synthesize.bio and access our models, whether they want to access them directly through our web platform.

We also have API access that they can get access to them in R and Python, which is where a lot of computational biologists live. So they can go access those today and they can go read our preprint about our GEM-1 model that’s online right now and they can see how we’ve carefully evaluated our results and make sure we’re being skeptical scientists.

Rob: Just to double click on that, we really want as many people as possible to use our models. We’re making them available for free right now so that anybody, regardless of where they are or what they’re doing, can experiment with our models, see where they work well for them and let us know if there’s areas where they don’t. Because one of the cool things, the opportunities about a model like ours is that it can be used for almost anything in biomedical research.

And so we’re still figuring out what are the areas where there’s going to be the most transformative impact. Just like with LLMs, I mean five years ago, I wouldn’t have told you the LLMs would revolutionize programming. I don’t think anybody would’ve said that. And we’re trying to figure out what are the areas where we’re going to see the most increases in velocity and scientific power from using the model that we built.

Joe: That’s amazing. I’m super excited. I think we’re in the very early days of this, and so I’m really excited to see where this goes. Like I said, not ChatGPT moment yet, but I think this is going to be impactful for the future of drug development. So thank you both so much for coming on today, and look forward to continuing to work with you.

From $1M to $10M ARR in 6 Months: How Fyxer is Winning AI Productivity

 

Fyxer.ai is an AI-powered assistant that helps users reclaim hours of their day by managing their inboxes, writing replies, and taking meeting notes. But Fyxer didn’t start with AI. It started with a decade-long executive assistant agency that gave them a treasure trove of data and insight that became the foundation for integrating AI.

Fyxer CEO and Co-Founder Rich Hollingsworth joins Madrona Managing Director Karan Mehandru on this episode of Founded & Funded to share how Fyxer went from $1M to $10M ARR in six months with just 30 employees.

In this episode, Rich shares:

    • Why great startups start with real-world pain
    • How there is still room for disruption with email
    • Why quality data matters more than quantity
    • The PLG-to-enterprise playbook that actually works
    • The rowing team mindset that drives their decision-making

Whether you’re a founder, operator, or investor, this is a masterclass in AI execution, startup focus, and scaling fast without breaking.

Listen on Spotify, Apple, and Amazon | Watch on YouTube.


This transcript was automatically generated and edited for clarity.

Karan: You and your brother Archie didn’t just stumble into this; you spent years in the trenches of this core problem before you made an AI product. So, walk us through that journey. What did you learn? What did that teach you, and how did that experience of owning an EA agency inform the blueprint for Fyxer?

Fyxer turned years running an EA agency into AI productivity tools for email and meetings — and hit $10M ARR just 6 months after launching. Launched by brothers Richard and Archie Hollingsworth.

Richard: The intention behind starting the EA agency, which we did in 2016, was to use it as a platform to build the AI solution, but we recognized the technology wasn’t there at the time. When GPT-3 came out, we realized that the technology was available, and we were ready to go, and we got together with our third co-founder, Matt, who’s our CTO, and drew from our executive assistant agency, a kind of data pool, and a series of customer insights that helped inform building the product, and eventually get us product market fit instantly when we released it. I’d say the main takeaway that we believe in the market is most people really underestimate the role of an executive assistant. They think that it’s about executing tasks, book me this meeting, schedule this appointment. Whereas in reality, the value of an assistant is determined by how proactive they can be for their boss. And so building a memory and a knowledge of the customer is the key to being able to unlock real value to the user. And we took that thinking from day one, and that’s why our product is A, built within the user’s existing workflow, and B, it’s built across the workflow. There are lots of point solutions in the market. There are very few products that are a full suite, and that we saw as an essential route to unlocking the real proactive value of an assistant.

Fyxer turned years running an EA agency into AI productivity tools for email and meetings — and hit $10M ARR just 6 months after launching. Launched by brothers Richard and Archie Hollingsworth and Matt Ffrench.

Karan: When did you know that you hit product-market fit? What did it feel like, and what advice would you have for other founders who are still trying to find that first wedge that takes off?

Richard: So, there are two moments that I think about. One was after launching the product, we moved to San Francisco for four months, joined an accelerator, and during the course of that, we grew by eight times in four months.

Karan: Which is where we first met.

Richard: Exactly. I think looking back at it, we had hit product-market fit, but at the time, it didn’t feel like that. We were just pissed that we didn’t hit 10X, to be honest. The real moment that I felt it was at the beginning of this year, when I was speaking to a user, and they told me that we had saved them from getting divorced that year, and that’s when I knew that we were onto something.

Karan: So, as you look back, and you see this explosive growth that Fyxer has seen, what do you think are the scaling risks or missteps that you think were crucial learning moments for Fyxer, and potential pitfalls for other founders to avoid?

Richard: I think the biggest thing that we’ve gotten wrong, and I see it happening a lot, is people spend too much time planning what’s going to happen if their plan doesn’t work, thinking about what they do in that scenario, rather than thinking about what happens if their plan works. And at the beginning of this year, we grew from $1 million to $5 million in ARR in two and a half months, and we hadn’t planned for that outcome well enough. And so we saw, for example, our customer support response time go from five minutes to five hours, and we found lots of similar scaling challenges in the process of growing so quickly. And that was a lesson that we won’t make… A mistake that we won’t make again.

Karan: I can certainly resonate with the thinking around what if things go wrong versus what if things go right. So, that’s sort of what venture capital is all about, too, because we have to believe in the art of the possible. That makes sense to me. So, as you look at Fyxer today, you’ve got capital, you’ve got resources, you’ve got a huge market, and so you almost have infinite choices in front of you, and where you can focus your energy, and your efforts. How do you make the decisions on where you prioritize, what you prioritize? How do you focus the team on the things that potentially matter, the things that you need to do today versus save for another day? What is the framework that you use?

Richard: So, I think this is where we’re exceptional. We have the fortune of having a map from our EA agency so we know exactly what we’re building, but not only that, we know exactly how to build it, because we can see the workflows that have worked historically, and we copy them wherever we can. The mentality we have within the company is, so in 2012 the British rowing team were training for the Olympics. They eventually won the gold. And the question that would always be a challenge to anyone who suggested doing something new was does it make the boat go faster? And we obsess about our end user who is a 55-year-old real estate broker in middle America, and our job is to make their email experience better. If the new tech idea, or the new direction people want to take the company in, or the new user they want to serve doesn’t do that, then we just don’t do it.

Karan: I love that. I love the analogy, too, of the rowing team. One of the things I’ve noticed, now being involved in the company with you, has been that you’ve balanced the bottom-up PLG motion with a top-down enterprise motion. So, you’ve got a lot of pre-users that convert into paying users, but then you’ve also got some really large customers like EXP, and others. How have you done that? It’s typically very hard for a startup to manage, and get both of them right, but how do you manage the bottoms-up PLG usage versus the top-down enterprise deals sales cycles?

Richard: We knew that this wasn’t optional. This is essential that we need to capture enterprise deals if this is going to become a billion dollar revenue company, not just a billion dollar market cap company, but the vehicle in order to do that very quickly at scale is using the PLG motion for individual users at enterprise companies signing up to the product, us building critical mass in the organization, and then taking it to their leadership in order to sign an enterprise deal. And we’re fortunate in our founding team, we have that skill set, and my co-founder Archie, who ran that motion from day one, and we managed to land a $1.2 million contract within our first six months, but it was all generated from individual user signing up, it expanding across the company, and then us capitalizing on that opportunity.

Karan: In some ways, the North Star, the ethos around product experience, and onboarding really haven’t changed that much. It’s just the nature of the deals that are coming in both from in bottoms-up, and the top down.

Richard: Yeah, we think about the user in both scenarios as exactly the same person — that 55-year-old real estate broker in middle America could be working for a large organization, or could be an individual broker, but the buyer of the software is different. And so there’s kind of a sales motion around that, and an onboarding experience around that we’ve had to allocate time for.

Karan: So, I don’t think you can go a day these days not reading, or hearing about how AI is disrupting everything, how it’s innovating, and how AI has permeated every sector of society. So, if you think about all of the stuff that you hear, and read about AI, and specifically related to the market that you’re going after, which is productivity, what do you think most people get wrong, or what do you think most people believe about the future of AI in your space that you don’t necessarily believe?

Richard: The biggest one I think is when I look at what’s happened since GPT-3 was launched, everyone thought that by now email would be a solved problem, that AI would’ve taken this off of our hands. And the reality is that just hasn’t happened at all. I remember when we started Fyxer, everyone said, “Oh, Google are going to do this really well. Microsoft are going to do this really well.” But the current competitive landscape is there are no startups I know who have gotten over a million in revenue, and Google, and Microsoft, I don’t know a single person who likes the products that they put out. So, our view on why that is that people believe that its quantity of data is what gets you the best results from the model, not the quality of the data that you train it on. And we have over index on quality. And I would say that’s the biggest differential we have to the rest of the market.

Karan:
Hearing you speak about email reminds me of that Mark Twain quote, which is the rumors of my death have been greatly exaggerated. And I think people have been talking about the death of email for years now, and it just hasn’t happened. It’s in fact gone in the opposite direction where we’ve gotten bombarded with emails. Let’s come back to Fyxer’s growth, and it’s very rare to see a company go from one to 10 in six months, and I know what aspirational plans you have for the company this year. As you think about thinking about hiring, and building a team that can capitalize on that opportunity in front of you, how do you think about hiring? How do you think about the people that you bring in to Fyxer in context of the growth plans that you have, and are seeing in the company? Are there frameworks that you use? Are there specific skill sets you look for? Are there specific types of people that you’re bringing in to Fyxer? Who are the folks that work well versus not? So, talk a little bit about your hiring strategy.

Richard:
Yeah, I think hiring is the one place that we really invest in looking in the long term as much as possible. Day to day at Fyxer, we’re obsessed by execution, and we’re focused on how do we move the company forward in the next day, in the next hour, and the next minute. And we focus the company very strongly there. Hiring though, we try, and look out as far as we can, and I think about it across two dynamics. One is if we’re hiring an exact person, I want them to have been there done that before. And so all of the folks that we’ve hired in the most senior roles have been a part of businesses that have scaled to as far out as I can see the next three years for Fyxer. But the crucial component is they need to be willing to do the very unglamorous work.

So, when at the beginning of this year we raised our series A, we had a team of only four people, and I had the opportunity to build this with a blank canvas. And the first people that I hired were an exec team, which really was quite an unusual, unorthodox approach. And with that it meant that we could have senior leaders build out their entire teams, which is fantastic, but it required all of them doing the grunt work, if you like, in their first three to six months. And them wanting to do that is the biggest important component for them getting the role. And on the other side, we like to hire people with less experience, but with a real growth mindset. So, hiring for the slope rather than the intercept. I know, is a mental model that you have. We very much lean into that, really, because we like the naivety, almost, of people having not got a playbook for how to do something, for thinking as much as possible from kind of first principles.

Karan: Yeah, I think that’s great, and I love that, and I can see it apply to so many real realms of society, including your work, and my work where pattern recognition, and having seen the story helps you for the first call it 70 yards, and then that healthy level of naivety, and first principle thinking is where the real magic happens. And I’ve seen that pretty much every day in my interactions with you and the team. So, kudos to you on sort of living that philosophy. Let’s go back to something more personal. Not many siblings would even think about starting a company together, yet you and Archie have done that, and done that successfully. How has that relationship either helped your journey as a founder, and how did you guys meet Matt and bring him on as your third co-founder? So, walk us through a little bit of that personal story.

Richard: Archie and I actually grew up on a farm. So, the nature of a family farm is that work, and family life are kind of intertwined very, very closely. So, it always felt natural to us that we would work together, particularly because we have very complementary skill sets. He is a born and bred salesperson, and my role as CEO — I’m a real generalist, so it just fits very nicely. He also has a very low ego, so he doesn’t mind his older brother being his boss as well. And I think that’s quite a key component to making the dynamic work. The journey then, so in the process of building our executive assistant agency, we knew that we wanted to build it into a tech company, and Archie met Matt, actually at a poker game maybe four or so years ago, and Archie was convinced that Matt was the person who should help us build that. And the reason for that is he’s a born and bred product engineer, and so has a 360 view of the engineering team, and we knew that he was very sought after.

So, we knew that we needed a very, very compelling pitch in order to persuade him to join us. And this is someone who the former CTO of Stripe described as the most impressive future CTO he’d ever met. So, we knew that we needed something really compelling for him. And the way in which we did that was taking what was a 500,000-hour time tracking data set from our executive assistant agency that could tell Matt exactly what were the key workflows to build, and then we recorded the screens, or we had our assistants record their screens of the work they did so that he could map the exact workflows. And so our pitch to him was, join us for instant product-market fit. Join us because you have a CEO with experience in building, in company building, and you have a head of sales whose experience in delivering multi-million dollar enterprise sales contracts in this market was what kind of convinced him that we were the right people.
Karan: And having personally interacted with both Archie and Matt. Now, I feel like I definitely resonated to a lot of the comments that you made. Matt strikes me as one of those people who is a man of few words, but when he speaks, people listen, and so do I.

Richard: Yeah, that’s very fair.

Karan: And then Archie, I’ve often felt if there was a picture next to the word grit in the dictionary, it should have a picture of Archie on there. Great. Well, to finish us off here, how about a few rapid-fire questions, if that’s okay?

Richard: Yeah, sure.

Karan: All right. Well, what’s the one thing you’d wish you started doing earlier as a CEO?

Richard: I wish we had hired people earlier. We pushed it too far in the end of last year where the four of us at the time, so it was the three co-founders, and on one employee were working seven days a week, 18 hours a day, and I felt like we paid the price for that towards the end of the year where we didn’t capitalize on opportunities that we could have done.

Karan: Got you. This goes back to your earlier point of planning for success versus…

Richard: Yeah, yeah, that’s a good point. Yeah.

Karan: How do you see the AI system category evolving over the next two to three years?

Richard: I see startups winning it, and incumbents continuing to flounder. I think they’ve got so much… There’s so much disruption being caused by AI that it is redrawing all of the existing product categories that we have that the incumbents can’t help but protect their core businesses. And as a result of that, are going to be unable to innovate to the same degree that startups are, particularly when there’s as aggressive, there’s so much market pull that you’re able to grow a business much faster now than historically has been possible.

Karan: Definitely attest to that. What’s one founder skill that you find yourself learning right now? What are you working on in yourself?
Richard:

We’re fantastic executors. We obsess over doing the unglamorous work. We need to get better at promoting ourselves, and that’s one of my main goals for this year.

Karan: Well, welcome to this podcast.

Richard: Exactly.

Karan: Other than Fyxer, what is the one other AI tool that you and the team can’t live without?

Richard: I think this is probably the answer for 95% of people, but honestly, ChatGPT. I mean, our team relies heavily on tools like Cursor, but the one that they can’t live without is ChatGPT. I think that has found its way into everybody’s workflow and become more and more valuable, over the last year in particular. I’ve tried every other version, Perplexity, and Claude, etc., but I just haven’t found something that works quite as well as GPT.

Karan: I’m a huge fan of this podcast, Invest Like the Best, and they always end the podcast with this question. So, I’m going to ask you the same thing, which is, what is the kindest thing that somebody’s ever done for you?

Richard: I probably just think of my wife. She’s done lots of nice things for me over the years. Last year, she moved our life, including our newborn son, to San Francisco at the drop of a hat, literally within 24 hours of me asking her to move to San Francisco for four months. She just did it without question, because she knew that it was important to me, and that has paid dividends for my life and my career. So, I’m incredibly grateful for that.

Karan: That’s great. I get asked a lot of times, what is the biggest decision I’ll make as a founder and CEO? And I always tell people, the biggest decision that you’ll ever make as a human being is the spouse that you decide to spend the rest of your life with. And sounds like you made a really great decision in that. So, thank you for sharing that. Well, that’s all we have time for, but I am so happy and excited that you and I got a chance to talk. Really proud of our partnership. Really excited to see what the future of Fyxer has as we partner on this journey together, and thank you for sharing some of your minutes in the day with us today. Thank you.

Richard: Thanks, Karan.