Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov
Hi, we’re Tim Berglund, Adi Polak, and Viktor Gamov and we’re excited to bring you the Confluent Developer podcast (formerly “Streaming Audio.”) Our hand-crafted weekly episodes feature in-depth interviews with our community of software developers (actual human beings - not AI) talking about some of the most interesting challenges they’ve faced in their careers. We aim to explore the conditions that gave rise to each person’s technical hurdles, as well as how their experiences transformed their understanding and approach to building systems.
Whether you’re a seasoned open source data streaming engineer, or just someone who’s interested in learning more about Apache Kafka®, Apache Flink® and real-time data, we hope you’ll appreciate the stories, the discussion, and our effort to bring you a high-quality show worth your time.
Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov
Being Wrong in the Right Direction with Caleb Grillo | Ep. 29
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
Tim Berglund talks to Caleb Grillo (Confluent / WarpStream) about his career in data streaming product management. Caleb’s first job: washing windows. Their challenge: reshaping Confluent Cloud’s billing and pioneering diskless Kafka to trade latency for huge cost savings.
SEASON 2
Hosted by Tim Berglund, Adi Polak and Viktor Gamov
Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
Music by Coastal Kites
Artwork by Phil Vo
- 🎧 Subscribe to Confluent Developer wherever you listen to podcasts.
- ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
- 👍 If you enjoyed this, please leave us a rating.
- 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
But we pioneered this idea that you can sacrifice latency in order to gain huge amounts of cost savings, specifically infrastructure cost savings. I mean, the thing that people need to understand about iceberg is that it is a it's a table format, it's a building block. Everyone's like, oh, good PMs are right a lot. You're you're wrong, like almost all the time. You just have to be wrong in the right direction.
SPEAKER_01Hey there, everybody, I'm Tim Berglund, and welcome to Confluent Developer, the podcast where we explore the fascinating journeys of software developers tackling complex problems. In this episode, I talked to Caleb Grillo, the staff product manager for Confluent Warp Stream. Caleb recounts his story starting as a PM in the early days of Confluent Cloud, talking about billing and a maturing platform, and then later joining an exciting young startup, WarpStream, in late 2023. WarpStream at the time was busy turning the world of Kafka upside down, and Caleb tells the whole story from the perspective of a data-driven product manager who is always wrong in somebody's eyes. Caleb tells a great story, and now we get to listen. Yeah. Hey, tell us a little bit about yourself. What uh what are you what are you doing these days?
SPEAKER_00Yeah, uh so I'm I'm Caleb. Um I'm our uh product lead for the Warp Stream product line at Confluent, uh, which is our uh BYOC product line. So bring your own cloud. Um we uh we'll we'll get into all the details of that, but yeah, I I run our our product management function um here. And uh yeah, so that day to day, it's everything's different uh every day. Um so no no two days are really the same uh for me, but I work I work across uh our engineering team to figure out you know what we're what we're building and whether it's the right thing for our customers and make sure that we're prioritizing all the right things. Uh but I also do a fair amount of sort of analytics work, um you know, digging into sort of you know how people are using our product. Um and I also work on sort of the business side, the GTM, um, making sure that our sort of commercials make sense and and sort of integration with the larger confluent business and all that. So it's a pretty uh that's a lot knee knee deep, mile-wide type role.
SPEAKER_01What do you what do you mean analytics? I thought being a PM was just vibes. You're saying you use data?
SPEAKER_00No, no, no, no. Yeah, we uh so you know, all of our consumption data and and all of our usage data goes into um confluence data warehouse, and so we're able to sort of um do like trend analysis on things like revenue and you know cost and make sure that our cost and margin makes sense and all that stuff.
SPEAKER_01Yeah, so it's it's good it's good to have that because there's I I've said there's probably a lot of reasons why I'm not a product manager, like you know, maybe qualification or something, but I think you'd be a great one, too.
SPEAKER_00Well, here's the thing though.
SPEAKER_01Here's the thing when you're a PM, like every decision you make is to a significant fraction of your stakeholders, not just wrong, but like how could you be that stupid wrong? Yes. So you're you're it's this this process of being idiotically wrong all the time. In somebody's eye, like 40, 50 percent of the people think that all the time. I got into a line of work where I I get like the constant affirmation that I need, like being on stage and stuff like that. So I think that's the it's really just an inner thing. You know, I I need people to tell me how awesome I am. You need the strength of character to do the right thing, even when they tell you you're wrong.
SPEAKER_00Oh you have that. I think to be to be a really great product person, you have to be a glutton for punishment. Yep. Uh and you have to everyone's like, oh, good PMs are right a lot. You're you're wrong, like almost all the time. You just have to be wrong in the right direction.
SPEAKER_01Yeah, and whether in terms of like objective results, you're you're right or wrong, you you could come up with some rubric for scoring that. Somebody thinks you're wrong all the time.
SPEAKER_00Yeah, oh yeah, yeah, yeah.
SPEAKER_01Yeah, you're just a guy, you're a moron all the time. And that I would I would struggle. Uh, what was what was your first job?
SPEAKER_00Yeah, my first job out of out of college was not um, well, my first job ever was uh washing was washing windows um in in high school. Uh I had uh my one of my brother's friends, I have an older brother, one of my brother's friends had a window washing business and we'd go around people's houses and wash their windows. That was my first job ever. It was like a summer job. Um my first serious job out of college, uh, I worked for an international development NGO. Um and so yeah, it was really interesting. Um, we did a lot of work for like USAID, the Agency for International Development, uh, but also for sort of UN agencies and other things like that. Um and so, you know, and it was really global, like there was there was there were projects in all kinds of different places and different different countries and continents. I didn't get to travel very much, uh, just you know, because I I moved on to tech uh before I sort of got into travel-related work. But um that role was kind of interesting. I was I was uh working on proposal budgeting. So, you know, there's like a technical portion of these, you know, sort of project proposals, uh, and there's a cost component. And so I was working on the cost component, which is figuring out, you know, how much does a Toyota land cruiser cost in Uganda, uh, you know, and putting that into budget proposals and things like that. Um so yeah, I mean it it was it was sort of interesting. I guess you could I I didn't at the time think of it this way, but you could draw some parallels between that and product management, you know, making sure that all the pieces fit together and make sense. Um but uh yeah, I mean that that was my first job out of college. Um yeah.
SPEAKER_01There you go. Window washing prepares you for product management only if your customers berate you about what a terrible job you did washing windows. Yeah, that's right. Um too slow. Too slow. You washed the wrong window pane first, and uh the pattern you used was wrong.
SPEAKER_00I read in a book somewhere that you were supposed to do it some other way.
SPEAKER_01There you go. This wiki this Wikipedia page. Um tell us about I think you kind of hinted this, but what's uh most interesting problem you ever saw to date?
SPEAKER_00Um yeah, so uh well I guess I can start with my my uh my character arc um because that'll that'll sort of explain my answer. Uh love that, yeah. I yeah, so so I I uh I got into um tech after a couple years of working for the International Development MBO NGO. Um I I got into tech uh in sort of a business operations role and then moved into product management a couple years later. Um and you know, one of my first product jobs uh was um dealing with a lot of the streaming data at uh at an e-commerce site that's not Amazon. Uh and so um we were basically building a data platform for uh for competitive intelligence data. So it was lots of high volume data um that was being like scraped from various places on the internet, uh, basically building a huge data lake of all that information. So teams like you know data analytics and data science could figure out which prod products were set were gonna sell well against others and you know, things like that, uh, who which sites were listing all these different products, how the review counts were correlating with uh you know, and review ratings were correlating with like listing those those uh products and surfacing them on competitor sites, basically trying to learn what was going on in the market from external data sources instead of just our own internal analytics. And that uh that pipeline, like the ingest pipeline, was Kafka. Um so this was 2017, 2018.
SPEAKER_01Um that was around the start of my my own Kafka journey.
SPEAKER_00Yeah, yeah, yeah, yeah. Similar uh similar tracks, I guess. Um and so you know, I I came to learn about Confluent when I was in that role and we solved a lot of interesting problems. Um, but I I kind of realized like, okay, there's like a more scalable way to do this, like to go to Confluent and work on the technology under underpinning this this whole thing. Um I just got super interested in that. And so I I ended up uh moving over to Confluent in May 2019. Um and when I started, uh we had a sort of a you you'll remember this, um, we had a sort of disjoint uh like product lineup. There was in the there was a cloud product that had gained a lot of traction, but that cloud product was not really what you would expect from like commercially from a cloud product. Um, meaning like it wasn't like AWS where you spin up an EC2 instance and you start being billed hourly for that instance and you spin it down and it you stop being billed. So there was there was like no possibility of self-serve anything uh in the in that version of our cloud product.
SPEAKER_01It's it's it's in our defense back then, it's very hard to make a cloud service. And so, yeah, that was super primitive. Yeah.
SPEAKER_00And it was it was like it was a proof of concept. Like it was like uh, you know, will people use a streaming platform hosted by somebody else? Yes, that was proven uh that they would, and it kind of it went too far.
SPEAKER_01Uh we got carried away, went crazy.
SPEAKER_00Yeah, yeah, it got carried away with the with the proof of concept, and and then it was time to make it a real product. Uh and so um when I came in at Complet, that was kind of the state. There was like, you know, the big cloud business was this um sort of interesting hybrid of like a hosted service, um, but it wasn't really what you'd expect from like a fully managed cloud product. It was sort of just we'll run your Kafka brokers for you.
SPEAKER_01Um Kafka on some servers is yeah, yeah, yeah, yeah, yeah.
SPEAKER_00And we wanted like the vision for Comple and Cloud was always to be way more than that. Um servers. And so yeah, and so so like the the biggest the the the sort of biggest blocker that we had, other than sort of the hard technical things of you know building a a real cloud service, um, we had the underpinnings of that, but like we were limited by the commercial model. Um and so that was sort of the that was the big problem that needed to be solved um back in you know 2019 was like how do we make the commercial model work um to in order to unlock sort of the self-serve you know cloud vision of a cloud product that everybody had. Um and sort of central to that was building out the consumption-based billing system, and so Confluent had a consumption-based billing system already, but it was for their self-serve cloud product. So there were two different products. There was the non-self-serve, serious usage, dedicated infrastructure hosted brokers product. Yeah, you like it. And there was another product, yeah, and there was another product that was kind of just like, you know, go sign up for this thing, put in your credit card, it's self-serve. That was all multi-tenant clusters. Um, so like there were all the underpinnings that we needed to sort of put everything together, but it was impossible to marry those two worlds together without the commercial model making sense.
SPEAKER_01Okay.
SPEAKER_00And so, in order to have a consumption-based commercial model, you had to have a consumption-based billing system that could handle you know all the different sort of cluster types you'd need and all the different infrastructure, like there's you know, Kafka Connect, there's KSQL. At the time, Flink wasn't there, but you need it to be sort of able to support all these different um products and different like cluster types, and you need to be able to meter throughput and storage and like all the different usage metrics. Um, and so building out the first version of that was kind of uh welcome to Confluent. Here's here's this huge project, you need to go figure out how to do it.
SPEAKER_01Now, a quick word from our sponsor. Confluent developer the podcast is brought to you by Confluent Developer the website, which has everything you need as a developer of data streaming systems. And it's completely free. We've got curriculum, hands-on exercises, tutorials, the online data streaming engineer certification are also free. A way to find a meetup near you, those are free, everything is there. I really want you to be successful in your journey as a data streaming engineer, and this is the site that has what you need. Check it out at developer.confluent.io. That's developer.confluent.io. Now back to the show. Um that was you and I were coworkers then, but remind me, were you a PM or an engineer? What where were you then?
SPEAKER_00I was a PM. Yeah, yeah, yeah. So I I I started I started a confluent as a PM. Uh and and so I yeah, that that was the first sort of thing to figure out. And that involved like it was hugely cross-functional, uh, which is a you know, hard thing in itself.
SPEAKER_01Always.
SPEAKER_00Meaning I had to work like with engineering, uh, but also all the other PMs with running all their other services, and also the sales team and like to and the marketing team, like the the whole way things happened at Confluent had to shift. Uh, it wasn't just, oh, let's just build a billing system and build it and they will come. Like we had to make everything work. The finance team, like we had weekly meetings with like you know, 10 different heads of X in the room trying to figure out sort of what we were doing. Uh and so that was difficult. It was also technically difficult because you know, like any metering, like at the time, there weren't these nice SaaS tools that you could buy to sort of solve the technical underpinning of your billing system. Like every company that wanted to do this had to build their own system.
SPEAKER_01Yeah.
SPEAKER_00Right. Yeah. Uh and so we had already built, you know, before I got there, there was there was like this sort of pre-existing metering system. Um, but it didn't support, for example, having multiple cluster tiers. We couldn't have like a dedicated cluster and a multi-tenant cluster sitting next to each other being built. Uh, so we had to like build that in. Um there's also, you know, this is getting into sort of boring business stuff, but like there's uh there's a concept of like a commit or like a minimum commitment that you know a customer would make in order to have sort of a long-term deal with Confluent. Um and building that concept into the billing system. Like we had to do that because your usage gets metered, but you can't just meter it at list price if somebody's getting a discount. So we needed a concept of a discount. Like there were like all these things we had to model in and sort of we were changing the wheels on the 747 as it was coming in for a landing, you know. That's the metaphor.
SPEAKER_01There's customers, customers using the platform at this point.
SPEAKER_00Yes, yeah, yeah, yeah, yeah. That's true. And then we had to come up with a strategy for migrating people over and making sure that the pricing made sense, and like, you know, we're just we're trying to merge all these worlds together. Uh, so it's very, very interesting and very, very high impact, uh, but also very hard um to get it right. And you know, we didn't, I don't think we got it 100% right the first time, uh, but we we built it in a way that was like that it could in theory support all the future iterations that you currently see in Confluent Cloud. Yeah, um, you know, and it's totally different now. It's like a completely different system. Um, the technical side of it was was also super interesting because you know the the infrastructure under the hood is you know throwing off all sorts of metrics. There, it's you know, you you can measure all kinds of stuff. Um, and figuring out like what what pipelines, like what parts of the observability pipeline would be suitable for billing was kind of a um, you know, that was a challenge in itself. Uh figuring out which metrics meant the things that they needed to mean when it comes to billing, because that's a different problem than observability. Yes. Um, yes, I mean it it it was it was uh it was a big undertaking, but I I think that it, you know, I I learned a ton. Uh it's a great, it's a great introduction to a company to be like, hey, can we just like change everything about it?
SPEAKER_01Uh yeah. Totally cross-functionally. Different agendas, different motivation, different incentive structures, different uh personalities, and and and you make it all work.
SPEAKER_00Yeah, yeah, yeah. The technical side of it also was was super cool. Uh, I don't want to minimize that. Like that that was um for me very interesting to figure out like how to translate all of these different requirements, very specific, like you know uh like gigabytes of cons of of rights, like rights to the cluster like means something specific. There's like six different metrics that you could potentially think would be you know used for measuring rights and what's the right one. Uh and so you know, there's there's stuff like that. And then you know, um different, like you know, I guess if uh somebody wants to be able to do per second billing, it's like, well, actually we want to aggregate things hourly and send present it hourly to cut like there there were there were lots of arguments about like what specific things we did and and we had to figure out like what was supportable with the current platform and what when you have things like you know bandwidth limitations and things like that, that implies uh a measurement regime and a time granular.
SPEAKER_01There's all kinds of things that that have to be true that are their own engineering problems to solve. And dirty secret in any given system that has a thing like that, you know, that might be over the whole day or something, you know, that it isn't necessarily like an instantaneous uh thing, it just depends on what you've built. And it's it's all its own way.
SPEAKER_00Yeah, this was something that was interesting that we always had to explain to people. It was like there's like a period of time where when you like log into the UI and you look at the billing screen, there's a period of time where the amounts can change. That like broke people's brains when you're trying to explain this to customers. They were like, and this was you know, 2019, 2020, you know, yeah, six years ago. Uh people would would look at that statement and be like, wait, but I can I trust your billing system and be like, yeah, yeah, we have all these correctness checks, and everything's like everything's good. It's just that there can be some late-arriving metrics. Like this is a massive distributed system of distributed systems. Like you can have stuff happen, but don't worry, like your monthly bill is fine.
SPEAKER_01Yeah, yeah. It'll converge, it'll converge.
SPEAKER_00Yeah, yeah, it'll converge, right?
SPEAKER_01So you um you ended up fast-forwarding a little bit at at Warp Stream. Uh how'd you get there?
unknownYeah.
SPEAKER_00Yeah, so uh I had been at Confluent for about five years. Um and actually it it came about basically because Rishi, uh the one of the co-founders of Warpstream, posted this blog post, which is now very famous, uh, called Kafka is dead, long live Kafka. And all the ideas that they were putting out there were just very interesting. They were basically saying that like there's this category of workload where you know Kafka is a super low latency real-time system. There's this category of workload that doesn't need to be super low latency, like it's not extremely latency sensitive. And I remember this, you know, for years and years. It was just like it wasn't even a thought that could cross your mind that you could differentiate a a product on latency. It was like, no, it's just fixed constraint. It's a real-time system, latency must be small.
SPEAKER_01It's um you know, is it is it zero yet? Well, keep working, you know.
SPEAKER_00Yeah, exactly. And it's like it there, there was never the insight of like you can differentiate uh on this dimension that like is a little bit counterintuitive if you're just glancing at it. But if you think about it for a second, it's like if you're running an observability platform for customers, you know, that's your product. You your your customer doesn't notice if it takes a few hundred milliseconds for your ingest pipeline to receive data. Like it does not matter in the grand scheme of things because there's this whole pipeline that needs to happen that needs to run to process that data before it can ever be displayed on a graph in your product.
SPEAKER_01And I I remember, I I think it was the summer of 2020. Like you could check me on that, but reading that blog post and just kind of warp stream coming on my radar and realizing there's 23. 23. Okay. Okay. Um realizing I don't know what the the like the demand elasticity of uh of latency really is. Like how many how many people if if you could give them half the price of Or a tenth of the price, or you know, do that, but it takes a second. How many people care? Like, I didn't know. And it was just it was that same revelation. Like, wow, that could be 90% of the market that just doesn't care. I and and I still don't think we really know that, but um it was very interesting.
unknownYeah.
SPEAKER_00Yeah, I mean, and and and I I I think that so I read that blog post. I literally just emailed like founders at warpstream.com. Nice. Uh and I was like, hey guys, like you seem to have some good ideas. You seem pretty smart. Uh if you ever need a product person, let me know. And they replied being like, Yeah, yeah, we're we're three people right now.
SPEAKER_01Like that was a valid, that was a valid alias they they had created.
SPEAKER_00Yeah, too early. No, no, I mean they said like in the the CTA, like the call to action at the bottom of the blog post was like, you know, email founders at if you're interested.
SPEAKER_01Okay, okay, okay.
SPEAKER_00Um, so I did. I was like, I was interested for a different reason. Uh just because it seems like a very interesting problem. And it was also, you know, very it was it was early, it was like early in a in a startup's life cycle. It's just I wanted I'd never done that before. I wanted to get you know that experience. Um there was a lot going for it, basically, in my mind. Um and yeah, so I I they they were basically just yeah, yeah, it's too early, like whatever, let's let's keep talking. Uh and then like a week later, it was like, just kidding, we're gonna raise a series A because VC's read this blog post and they want to invest in our company. Would would you like to talk more seriously? I was like, Yeah, okay, cool. Yeah, we'll do this. Yeah. Um yeah, so and and so I joined I joined Warp Stream December 2023. Um and yeah, about eight months later, we were, you know, in the in the early uh early talks with Confluent and ended up joining Confluent. And the way that happened um was pretty interesting. I mean it it at first, you know, that we didn't really know what to make of it. We're like, you know, what does Confluent really want to do with Warp Stream? Like they've they've got freight clusters, like what do they what do they want with us? Um so we all we had a meeting in Mountain View. We all went to Mountain View and and met with um with you know a bunch of people on the Confluent side. And we kind of got some confidence there that you know the intent was not to like put us on a shelf, you know, that we were a threat to their business or anything like that. Um we we felt like it was very additive and that like they they you know the team very much wanted like another line of business. The BYOC deployment model was like the thing that made it tick, and like that, you know, that that uh that was that was a really good process for us to go through. I think that if it was, you know, if if if it was a different situation with a different, you know, a different uh you know parent company, I think that it would be uh slightly slightly different dynamic for sure. Um, you know, and and you know, maybe maybe it would have worked out differently or whatever, but we didn't even you know think that way. We were just like, okay, Confine wants to add on this this new line of business, we're a good way to do that. Yes. Um and it was very additive, so it was it was good. Yeah.
SPEAKER_01Excellent. How would you I love I love that arc, I love that story. If you had to summarize the like the problem you solve, so kind of warp stream is the answer. What's the hardest problem you're solved? Well, warp stream. You as the PM there, yeah, um summarize that through that lens, like as a as a as a as a life scale problem challenge, cool thing that you're doing. Um what is it that uh you do? You you talked about being a PM, but like as a problem, what is the problem?
SPEAKER_00Do you mean like what what problem does WarpStream solve as a product or solving for okay?
SPEAKER_01That's that's actually worth in case there's anybody listening who doesn't know. Sure. Uh in fact, I'll just I'll I'll take a swing at that. You tell me you tell me. Okay, go for it. Yeah. Uh WarpStream is a diskless Kafka. So um it is uh a system that implements the APIs and semantics of Kafka, uh, but doesn't store data on any locally attached or network attached storage. Uh it's all stored in cloud blob stores. This is a trend in data infrastructure. Pick a form of data infrastructure, and you know, we can identify the one or two groups doing an open source project to disclassify that. Um and so it's a uh lower cost, higher latency way of being data infrastructure that also gives you BYOC characteristics. So all of the data is stored in, say, your S3 bucket under your cloud control, not under warp streams, not under Confluence. Like and I'll give me the warp stream pitch here. Kind of cool thing about it is uh it's a very thorough BYOC because like most BYOC implementations, the vendor can get in there and break glass if they have to and go in and fix stuff in the data plane. But here it's just there's no glass, it's it's it's out there and nobody can do anything. So yeah, it's BYOC Disclos Kafka. That's what that's what Warp Stream does. You're a PM there. Um yeah, yeah. So you've you've you've kind of solved that. I guess I guess that's the problem you've solved.
SPEAKER_00Well, no, there there's a ton. I mean, like, there's a ton of potential you know directions that we could go. There's infinite universes here. Um but my my role here is to basically help us pick the direction that makes our customers the most better off. Um meaning, you know, what else do people want to use our platform for? The deployment model provides some interesting possibilities for you know um the, you know, the like you were talking about. It runs in your account, in your VPC or in your data center, even uh, and just egresses metadata out. So with that in mind, like there's some use cases where maybe we could like lean into that with some more product features to make it um you know good for very highly sensitive workloads that you know you you wouldn't you wouldn't ever want to put on a cloud service. So that's one sort of category of things we can do. Um there's you know the the the scaling element of like the the scalability of the of the platform is another sort of thing we can we can sort of explore. For example, we recently sort of rebuilt our storage engine to be able to uh support um basically huge amounts of data retention because people there's a there we had a couple of customers who were saying basically that like we have infinite retention on compacted topics and we want to store data forever. We just expect this to keep growing forever. Um and so you know, we need you to be able to support that. And it's not it's not like okay, cool, now we can just like forever into infinity support all the data you could ever store in a Kafka cluster, in a warp stream cluster, but we like we 10x'd it. We made it so that you could store you know huge amounts of data. Um, and you know, there's there's more work to be done there to like figure out you know, under the hood whether that's like we we basically we extended the the timeline that we have because as a streaming platform, they're all always writing more data in, and if it never gets truncated, it just keeps growing. Um, you know, so there's like a there's some roadmap items there. I think the most exciting thing though, right now that we're working on uh is warp stream table flow. And so if if you're not familiar with what table flow is, um yeah, I mean this is an EA now.
SPEAKER_01We've we've released it sort of an early access uh we're we're speaking in um early to mid-November of 2025, if you're correct listening to this. I don't know when this will be published. Future timeline, yeah.
SPEAKER_00Yeah, it might be GA by the time this this uh this goes live.
SPEAKER_01Uh probably will be GA by the time this goes live. But yeah, long tail viewers in the future.
SPEAKER_00Uh yeah, yeah, yeah, exactly. Uh so so yeah, so WarpStream Tableflow uh basically does the same thing as confl and cloud tableflow. It materializes your uh Kafka topics as iceberg tables in object storage. Um Warpstream has the BYOC deployment model, uh, which Tim you described uh pretty well. Um you know the data plane runs in your in the customer's VPC uh with zero access delegated to us. And so there's no cross account I am privileges, so it all runs in your VPC. We don't have any access to the actual data. Um the you know, your workload egresses metadata to our control plane, which is basically the brain of the operation. And so what we're doing with TableFlow uh is we basically created a new cluster type that um, you know, the the warp stream agent, which is the replacement for the broker, the Kafka broker that runs in your environment, uh instead of doing the Kafka job, it does the iceberg job. It writes parquet files, takes input from a Kafka topic, writes parquet files and object storage, builds the metadata manifest thing, um, and ships the metadata about that back to our control plane. And just like Complent Cloud Tableflow, it's a fully managed table, meaning like table flow handles all of like the tableflow agents handle all the background jobs, like uh file compaction, compaction is the biggest, you know, cleaning up orphan files, deletes, things like that. Um, and so it it's a it's it's more of a managed service than, for example, using like the Kafka Iceberg connector, because you still have to figure out like how to manage the table and how to manage the data in object storage and like how to do all these operations, which TableFull automates for you. Yeah, um, so that's why we say it's sort of a managed service with a BYOC deployment model, um, because that you know it it all runs in your in your account, but it automates a ton of the tedious work that you need to do if you if you were trying to build an iceberg you know data lake yourself.
SPEAKER_01Uh no, I and you know I'm not here to to shill, but table flow really is pretty great. It's it's a it's a it's a really good idea. Um easy button.
SPEAKER_00Yeah. I mean the thing that people need to understand about iceberg is that it is a it's a table format, it's a building block. Like it's not a data lake in itself, it's a way to express how to store data in object storage such that it can be used as a data lake queryable with a SQL query engine. Yes uh that's like that's not that's like a that's like a uh a small part of building a database. Um what we've done, you know, Richie and Ryan like to say that like we've table flow, warp stream tableflow is kind of the bottom half of a database. Uh it's like the back end like primitives, and we're gonna build more on top of it, but like that's um that's how to think about it. It's it's sort of uh it makes the building blocks a lot more useful and a lot more easy, like easy to uh to adopt.
SPEAKER_01Right. If you had to reflect, and it's early. I mean, warp stream is is young and and growing and and and still having its impact, but so far, what do you think the impact of it has been?
SPEAKER_00Um Warpstream was if not the first uh one of the first Diskless Kafka iterations. Um and it I think it's not an exaggeration to say that that we we pioneered um sort of this idea. We didn't know it at the time that it was gonna be, you know, that it was gonna have the effect that it would have on the industry, but we pioneered this idea that you can sacrifice latency in order to gain huge amounts of cost savings, specifically infrastructure cost savings. Um and you know, now that you sort of poke your head up and look around, there's there's you know, there's an open source Kafka, there's three open source Kafka proposals currently in debate. There's um you know multiple vendors in the space. Confluent has a managed service that uh you know basically is built on a lot of the primitives that Wordstream is built on, um, which is freight clusters and confluent cloud. It's like the fully hosted version of what we do. Um and so, you know, the the I think that the industry, I don't know, I don't know if it was like I don't know what the direction of causality there is, um, but it definitely seems like we influenced the direction of the industry by just taking the first step to just be like, there's a possibility that latency doesn't matter as much as you think it does. Let's just test that out and see. Um and I I think that you know we gained a lot of a lot of certain traction um you know early on with that idea. And I think that you know it caught on very quickly. Um it's been super interesting to be a part of it, yeah.
SPEAKER_01My guest today has been Caleb Grillow. Caleb, thanks for being a part of the Confluent Developer Podcast.
SPEAKER_00Thank you, Tim.