Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov

From “This May Never Work” to WarpStream with Richie Artoul | Ep. 17

Confluent Season 2 Episode 17

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 30:20

Tim Berglund talks to Richie Artoul (WarpStream/Confluent) about his career in data infrastructure. Richie’s first job: working at Howie’s Game Shack, a walk‑in LAN gaming cafe. His challenge: working at Datadog on a new log storage system.

SEASON 2
Hosted by Tim Berglund, Adi Polak and Viktor Gamov
Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
Music by Coastal Kites 
Artwork by Phil Vo 

  •  🎧 Subscribe to Confluent Developer wherever you listen to podcasts. 
  • ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
  • 👍 If you enjoyed this, please leave us a rating. 
  • 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
SPEAKER_00

Today, from walk-in land parties to discless data infrastructure pioneer, this is Confluent Developer.

SPEAKER_01

And as we were getting ready to kind of like migrate our first product, we realized it just like kind of didn't work. Someone's day is going to be ruined every day at that level of nights. When it was not working and we were grinding on it trying to make it work, I was like, man, I might have to eat some crow. Like, this is not working.

SPEAKER_00

Hey there, everybody. I'm Tim Berglund, and welcome to Confluent Developer, the podcast where we explore the journeys of software developers tackling really hard problems. In this episode, I'm interviewing Richie Artoul, who's famous for being the founder of the pioneering Diskless Kafka implementation warp stream. We talk about his first job at a gaming cafe and some really pivotal work he did at Datadog redesigning their log storage engine and indexing system. Richie tells a great story. I always really enjoy talking to him, so let's get to it. Welcome to another episode of the Confluent Developer Podcast. I am your host today, Tim Bergman, and I'm joined by Richie Artool. Richie, welcome to the show.

SPEAKER_01

Hey Tim, thanks for having me on.

SPEAKER_00

You got what do you uh what do you do right now? What's your uh current job, current title? What do you what are you all about?

SPEAKER_01

Uh my current job is director of engineering uh for warp stream at Confluent. Uh so I lead uh essentially what is like the Warp Stream group at Confluent. Um so we're like a product line um um so we have kind of like our uh our own software and um product yeah um and it is a fascinating one.

SPEAKER_00

Maybe we'll talk about it today. Maybe we won't. What was before you did that, long before that, your first job? What was your first job ever?

SPEAKER_01

Um you want like my real first job, right?

SPEAKER_00

Real first job. Like I don't want your first, you know, uh interned at Google writing that's like your actual Okay, cool. Yeah.

SPEAKER_01

Um yeah, so I I mean I actually didn't study like computer science in college or whatever, so my my first job out of college also wasn't tech, but okay. Um my first real job was working at a place called uh Howie's Game Shack, uh which I think I was 16, and it was like uh I forget what they're called, like one of those places you can go, like an internet cafe. Yeah, but it was all games, so they had like 50 Xboxes and like a hundred PCs, uh and people would go and play like Dota, like on Warcraft. I don't know if you ever heard of that game, and uh Counter-Strike and World of Warcraft there. Um, and I think this is when like the iPhone had just come out and I like really wanted one because I just wanted to have access to the internet all the time.

SPEAKER_00

Sure.

SPEAKER_01

I was a nerd. Right. Uh and I was like, okay, so I need a job. Uh so I got a summer job there. Um and I think I made just about enough money to pay for an iPhone. So that is amazing.

SPEAKER_00

So it was like a retail land party was the the the business.

SPEAKER_01

Yeah, that's exactly what it is. Yeah, it was an interesting, interesting group of people. That is wrong. Yes.

SPEAKER_00

I mean look, it's it's kind of our people. Yeah, yeah. We can say that, and yes, it is a very interesting group of people. Yeah, yeah.

SPEAKER_01

Uh well, especially just because it was a lot of like um, you know, high school kids and college kids, and I was like 16. So uh it was fun. It was a good it was a good first job.

SPEAKER_00

That that is an amazing first job. I love that. Uh and I and I want to go there. I I mean I don't know if I see uh even a business that exists anymore, but just sounds like a fun place to hang out.

unknown

Yeah.

SPEAKER_00

Well, I you mentioned that you weren't a computer science major. Uh this isn't scripted, but what uh what was your major? What'd you what'd you do in college?

SPEAKER_01

Uh I studied biochemistry in pharmacology. Um so uh you know, I don't know. My dad's a doctor. I was told to be a doctor from a young age. Uh that was that was the plan. Um and then uh in college I kind of realized like I didn't like hospitals, so I was like, I probably shouldn't be a doctor. Um but um you know I was you know kind of like top of my class, pre-med, that sort of thing. Uh and then just decided to not take the MCAT. Um so I uh just ended up working some kind of weird um office job doing like you know, editing word documents and type of thing. Finding finding your way, yeah. Yeah. Uh and then obviously hated that. Um and I think like about 10 months into that job, I just was kind of like losing my mind and just kind of like rage quit and went to a coding boot camp. Um and so that's kind of how I got into tech.

SPEAKER_00

I love it. I I did not know that part of your story. Um that uh that is fantastic. Uh and you know, well, I think we're all kind of glad you did. Thanks. Um so the big question of the show What's the most interesting problem you've ever solved? Uh I mean I I gotta tell you, I'm expecting you to say Warp Stream, but I I I don't maybe you're not gonna. So you get to you get to surprise me. What is it?

SPEAKER_01

Yeah, I I think uh yeah, I mean it's lots of stuff obviously I could talk about uh related to Warp Stream.

SPEAKER_00

Uh we might we might have you on the show again, just so you know. Yeah, exactly.

SPEAKER_01

Um but one of the things I thought uh might be kind of interesting was um I was actually my job before Warp Stream. Yeah. Um which I was I was working at Datadog. Um and you know, we basic basically um you know Datadog ingests and stores all your logs and events and and time series data. Um and I was part of the team that was building the new storage system, basically, to replace the existing one um for a lot of that data. Um and you know, building that system was like super fun and we we we did a bunch of cool stuff, but like the the part that was actually like really hard and kind of a grind was like, okay, how do we get this actually out into production now for like a hundred percent of products and customers? You know, Datadog has like tens of thousands of customers and dozens of products that are all powered by like one database.

SPEAKER_00

Um you gotta imagine Let me actually I want to ask you about that. Like, what what was wrong with the world of Datadog that they wanted to build the new system? What what did it and and I'll take like the business perspective, I'll take the developer perspective, hands-on. Like, what was it?

SPEAKER_01

Yeah, so the the business perspective is I think super easy. Uh one, it was like just old, like insanely expensive, what they were doing. There you go. Um, which is for all the same reasons, really, that um I'll slip warp stream in. Uh, which is that uh, you know, running kind of open source Kafka yourself can be really expensive. Networking, storage, auto-scaling, all that type of stuff. So a lot of the same impetus for Warp Stream, uh, but this was before before Warp Stream. Um, but similar problems. Uh and so that was kind of like the cost was an obvious one, but actually a lot of it too was like there was a lot of stuff that customers wanted in terms of like features uh that we just couldn't deliver with the existing system. Um so like one of those, um I don't know how obvious this will be to someone who's not like a datadog user, but like if you imagine a logging UI, right? Like there's a little bar at the top, and you're like you filter on things. And in the older versions of Datadog, you used to have to say beforehand, these are fields that are interesting to me and I would like to be able to filter on.

SPEAKER_00

Before ingest.

SPEAKER_01

Uh correct. Yeah. Oh, okay. Good.

SPEAKER_00

As long as you know everything ahead of time, it's fine.

SPEAKER_01

Yeah, which is great. So it's like you'd be in an incident and you're like, oh, I need to filter on this field, and we're like, uh no, you didn't tell us beforehand. And then you could start indexing it in the moment, but it would only apply to new data coming in, you know, several minutes later and not to anything in the past. So that's obviously extremely frustrating. Yeah. Um but if you think about it, like about without going to too many details into how the old system works, if you imagine like a traditional search system, usually there's like an index, you like a schema for the index, right? You're like, well, these are the fields that are important. But like every customer's logs look different, so that doesn't really work for like uh an observability use case where the customer doesn't like give you the schema. Um so that was just like one example of a feature um that we couldn't really implement in the old system that the new one could. Um and the other was like like uh just query performance really like you know, the old system worked fine if you just need to query a couple hours of data, but what happened is someone wants to query like three months of data. Nope. Um and they're a relatively low volume customer, so they they were on like one shard over here. And a shard is like physically tied to like you know a machine with a fixed number of cores. And so if we only had your data on this machine, but you wanted to run some massive query, there was no way for us to parallelize it. Um beyond what we had decided three months ago how many machines yet you deserved, basically. Gotcha. Uh so kind of um that sort of thing.

SPEAKER_00

So which is a very much first generation distributed data architecture kind of choice. That's common.

SPEAKER_01

Yeah, yeah. Um so those are the kind of the reasons that we were building something new. We wanted something cheaper, something more cloud native, easier to scale, easier to manage, but also something that would allow us to kind of build some of these new features and something where like, hey, even if you're a tiny customer, if I want to throw a thousand computers at your query for a second, you know, I can.

SPEAKER_00

Okay, so thank you. That's that's that's the what was wrong, and that makes a lot of sense. And you were starting to dive into like a particular storage problem. Uh keep going. Now a quick word from our sponsor. Confluent Developer the Podcast is brought to you by Confluent Developer, the website, which has everything you need as a developer of data streaming systems. And it's completely free. We've got curriculum, hands-on exercises, executable tutorials, the online data streaming engineer certification, also free. A way to find a meetup near you, those are free. Everything is there. I really want you to be successful in your journey as a data streaming engineer, and this is the site that has what you need. Check it out at developer.confluent.io. That's developer.confluent.io. Now back to the show.

SPEAKER_01

Oh, yeah, it was just the general problem I would say of like getting out to production. Because it like, you know, about a year in, we were like, okay, we've built a pretty good system. Like it's not perfect, but it's decent. Like it can do some pretty cool stuff. Uh, but like the reality of like hot swapping the database for like 30,000 customers and 13 distinct products with different features, different query patterns, like the reality of doing that um without breaking everything was was hard. Um and um so there were kind of like a number, and we what we ended up doing is we just tackled it product by product. So we pick like the simplest product we could, and we'd be like, okay, like let's just focus on this one. And and because you gotta imagine too, if you if you mess up, like let's say filtering because Nog language query language is super weird because like so imagine you're like um duration larger than 200 is like a filter you put into your logs, right? Okay. Well, some customers may be emitting the duration field as a string, and some other people may be emitting it as a float, and someone else might be implementing it as an integer, and some customers may be mixing and matching those from like the same service. And so the query semantics get funky. And if you mess up the query semantics, someone might get paged in the middle of the night for something they shouldn't have been. Um, or worse.

SPEAKER_00

Or not.

SPEAKER_01

Yeah, the opposite, yeah, exactly. Um so you have to be very careful and do all this kind of query shadowing stuff. Uh and I I remember one of the things that like really killed me was um you know, we were kind of getting ready to migrate our first project, and this requirement that I like had been vaguely familiar with but didn't really understand the the the real implications of, which is that in Datadog, they'll only they they had the storage system basically ensures that you only ingest each piece of data once. Like every log, everything that we ingest, like a JSON record, has an ID. Okay. And the storage system is capable of making sure that it only ingests it one time. Okay. Um, which is also super important because like you can imagine a lot of people have monitors that are like, if this happens more than three times in five minutes, page me. And it's like, well, if some intermediary service in Didadog's processing pipeline restarts and replays a couple of messages, and then those end up in storage twice instead of once, that will like ruin someone's entire day, basically, right? Okay. Okay. And so it's extremely important that like if you send us something once, we store it once, and it shows up once in the query. Um, and that, you know, if you imagine kind of like a database uh that's on SSDs and sharded, uh it's relatively straightforward to implement that feature. It's not easy, but like you can do it. You're like, okay, well, this record goes to this node on this shard, and that thing can kind of check because it's you know it's a KV store or whatever it is. Um but when you've built like a kind of columnar store on top of object storage, um like hey, does this ID already exist in the system? Becomes like a very hard problem um to answer. Um and so we we built a thing to do this, and as we were getting ready to kind of like migrate our first product, we realized it just like kind of didn't work. Like it worked 99.9% of the time, but that wasn't like good enough.

SPEAKER_00

There were some cases that that were there, yeah.

SPEAKER_01

Yeah, I think at one point we had it up to like we measured it. We're like, okay, we're at like five or six nines of deduplication, which like seems like a lot, but is still like not enough. Sure, sure.

SPEAKER_00

When you're looking at volumes like you're ingesting, there's still a significant number of duplicates.

SPEAKER_01

Yeah, like someone's day is going to be ruined every day, uh, if we at that at that level of nines. Um and so it was like me and another engineer, we were basically like all of the work we did building the system is for nothing until we solve this like to like, you know, yeah, 29s or whatever, basically. Um and so it took us about three months of basically just grinding on that feature um to get it working. Um yeah.

SPEAKER_00

How how did you find that you know you started by writing something that I'm gonna guess you believed did deduplication and then you find out you're wrong. How did you find out you were wrong at the beginning of that three months?

SPEAKER_01

Yeah, I'm I'm trying to remember how we noticed at first. Um well in the beginning it didn't work well enough that like people would like notice and complain. Like our internal c like the datadoc things, like not before we migrated customers, we'd migrate datadog internal products. Okay. And there were like parts of the product that would just like break, basically, if this happened. Okay. Um and then when we realized the problem was more complicated than we had kind of anticipated, I think what we ended up doing was writing a bunch of tooling to detect it. Um so we had a couple of tricks. Um, one, we could detect like for certain types of queries, it was easy to detect that like, hey, there's two duplicate events in this result. So we did that. And then um we had this compaction system that was constantly merging data in the background. Um and so we and what that would the that system was designed to bring similar data closer together, basically, to make queries faster. And so we added a bunch of code in there too to detect um essentially like, hey, while you're bringing similar data closer together, kind of look around yourself to see, hey, are there duplicate kind of do like a sliding window scan during the compaction to see if you have duplicate events in there. Okay. And and I think we then built some other tools that would essentially go and download a bunch of files, analyze them, look for duplicate events, and then emit logs and stuff. Um so we did we built a bunch of tooling basically to kind of passively scan in the background to see if we messed anything up. Um and then we would go and manually cause this to happen to see if our tooling detected it. Right. Um so we we just had to build a bunch of tooling. Um and we probably should have done that. Um this probably should have been the first thing we did. I think we just were kind of like in a rush and didn't build enough tooling uh in the beginning.

SPEAKER_00

Uh you usually usually believe you're doing it right. You know, you think you have a solution that works until you bump up against reality.

SPEAKER_01

Yeah, exactly.

SPEAKER_00

Uh what was that tooling the final answer uh that these these tools run and and detect duplicates, or did that cause then you to go back into the deduplicate? I mean, did you did you effectively solve deduplication on the input?

SPEAKER_01

The tooling gave us like a benchmark, basically like, okay, until these go away, like we're not done. And even then we may still not be done, but we'll be like much closer to done. Uh because I mean it would like once we started building smarter tooling, we could there were tons of operations we could do, and we'd be like, oh shit, it fired because you know we scaled this up, scaled that down, killed this node at the wrong time, or we would inject faults and see that was happening. So I would say that's actually when the work began, um, is when we started having those tools. And then it was like literally like, okay, um well, and part of what made this hard too is it wasn't just localized to our system. Like our system, in order for our thing to function correctly, it depended on services sitting upstream of us. Because essentially what the system did was like, for example, if two duplicate events were ever sent to um we had these things called shards, and a shard was essentially a group of Kafka partitions. So if anyone ever sent um you sent an event with this ID to this shard, and then later you sent a duplicate event to another shard, our deduplication system just did not work. So all of our upstream things also had to be functioning as well, which ended up being cool because eventually we got our thing working perfectly, and like a couple months later it fired, and it turned out we had caught a bug in something upstream of us, basically. Um so our thing, it was our scanners, because they essentially just ran passively in production, were able to catch issues that arose even outside of our own system, um, which was cool. Uh, but that's what kind of what made the problem so hard in the beginning is like there's so many moving pieces, and you like, you know, it's not working, but where did it go wrong? That's very hard, right? Um so I think we we um someone ended up writing a uh uh TLA verification of the protocol we were trying to implement just to make sure, like, is there just like a logical fallacy in the algorithm we're implementing?

SPEAKER_00

And if uh anybody listening doesn't know TN TLA, tell us briefly about that.

SPEAKER_01

Uh man, I'm not the right person to explain these things, but it's a it's a uh don't quote me on my explanation, but it's a it's essentially a um uh under the category of formal methods, and what it does is allow you to express distributed systems and concurrent algorithms in like a formal language, um, and then uh essentially run it through a simulator. Um, and the simulator will essentially like explore the solution space and tell you if any of the like invariants you said must always be true or ever violated.

SPEAKER_00

Right. And it's pretty common, and I am blanking on what TLA stands for. We'll look it up, we'll put it in the show notes.

SPEAKER_01

Yeah.

SPEAKER_00

Uh but I I think it's common experience when you know you've built your system, you're like, hey, this works, let's do TLA. Oh, it doesn't work. It's it's it's usually humbling, right?

SPEAKER_01

Yeah, and the the thing that's really cool about TLA plus, I haven't done it a ton. I've done it like two or three times in my career. Um, but it's like you know what people say like writing, like you don't underst you don't understand something or you haven't thought something through until you've written it down because writing is thinking. Yeah, I think TLA plus takes that to like the logical extreme. There you go. For computer algorithms, it's like you don't really understand a distributed algorithm until you've implemented it in TLA plus, and then you really know like it's almost like I often will find bugs, not because TLA plus like the simulator found it for me. It's because while in the process of trying to translate my thought into a TLA plus spec, you're just like, oh, that's that's just like wrong. Like you'll see it while you're trying to implement it.

SPEAKER_00

The discipline of translating it into the formal specification forces you to grind through. Yep, just like writing or or uh even explaining verbally, you know, if you if you can't do that in a simple way, um it's yeah, I totally agree. That's uh without had having used TLA to verify any distributed algorithms I've w I've written, the whole the whole concept uh that makes sense. I mean you were kind of you you had uh deduplication solved and you said you had caught a bug. It had been running okay for a few months, and your monitors caught a bug in some other upstream service, which is really cool. What was um this could rabbit hole, and we don't need to do that. So kind of in brief, the you described the the problem at the beginning. When you you write something to storage, you need to make sure you're only writing it once. And uh uh there are two things, I suppose, two categories of things that could be hard about that. One is the definition of something, like the identity identity of the thing. Uh I'm gonna guess that was solved. There were some unique IDs somewhere, that was all okay. Uh, and the hard part is race conditions on the writing or the reading. Um what was the storage? Uh was it a blob store? What was it disks? Uh was it something you guys had built?

SPEAKER_01

So um this was another thing that was interesting, which was like we decided, okay, so we'd built this completely stateless ingestion thing, like you know, our compaction service stateless, our ingestion service was stateless, our query services were stateless. Go and delete nuke or whatever. Yeah.

SPEAKER_00

In your architectural preferences here. Yeah.

SPEAKER_01

I I like to sleep at night, you know? Yeah. Um and we um so we built this completely stateless system, and you could go delete anything at any time and you would never lose any data. And then when we were designing this deduplication system, we were like, we can make this work with no disks, too, and no state and any of this stuff. And I got a lot Not just me. We got a lot of flack for that because I think there were some other people at the company that wanted us to just like essentially stick Rocksdb on some nodes and do the deduplication there. As as one does. And have it be pretty stateful. And that would have made the problem very maybe not very easy, but let's say much easier. Um but we really didn't want to do that because we were like, man, we got like we literally built this whole thing, and now we're at the final step, and we really don't want to just shove disks in here at the very last moment. And then when it was not working and we were grinding on it, trying to make it work, I was like, man, I might have to eat some crow. Like this is not working. It should be able to work, but it's not working. And what was interesting is that we actually were using the system we built to store if if you think about it, like the deduplication thing essentially boiled down to um it was essentially essentially another database um that needed to track IDs and be able to commit atomically with the other data file ingestion. Yeah. And we were like, well, if we start creating files that just have IDs, we can organize them the way we want so we can query them quickly. But then we'll have a lot of files and we'll need to compact them, and then when the files get old, we'll need to expire them. And we're like, we're gonna have to rebuild half of this database. And so what we ended up doing was storing them in the system itself that they were also serving as the deduplication layer. So they were essentially special tables in Husky, and they would go through their own compaction, which sounds like meta and weird and like it wouldn't work, but it did end up working and saving us a lot of work. Um, because then it's like, well, you get compaction for free, and you get data experien for free, and you get a file format for free. Um so that's how we did it. And the the IDs did have they were very specifically shaped, like they had time in them and some like tenant information and whatever, which allowed us to compress them like crazy in memory. Because what what we didn't end up any doing any discs, but we did end up doing is like having essentially a a hot set that was in memory that we're that we could verify against very quickly. Okay. Um and you know, we would essentially page them in and out of memory um very quickly. Um sorry, I feel like I forgot what the original question you were asking was.

SPEAKER_00

Um what what was hard about the like I basically like uh I think the most recent question was were there disks or not? And you're saying ultimately no?

SPEAKER_01

No, there were no. We did managed to make it work with no disks at all, um built on top of the system that it was supposed to be essentially deduplicating for. Um which ended up being um and it was nice because it ended up being a really simple once we ironed out the kinks, it was very stable, and you just kind of like you could add new nodes and they could start processing data, and they would load whatever they needed for deduplication into memory. And if they died, work would get transferred over and you didn't really have to think about it too much. But there was definitely a window where I was like, this may not work, like maybe we do more than we can chew. Um and we just kind of eventually we ground through it. But I I was definitely worried for that. Was probably the most in my career of like, mm, this may never work type of thing. Um because we were working, you know, it was like three months of like probably six to seven days a week, 10 to 12 hour days, just trying to get this thing. Because it was blocking everything, you know what I mean? Like until this was until like this graph we had essentially went to zero. Uh you couldn't get any value from any of the system we'd built, yeah.

SPEAKER_00

Right, right. Of course. That makes sense. And that value after you shipped it, um, you know, beforehand, you had to define indexes in advance. Uh it was really expensive, it was slow, uh, those things got better. I mean, that is that that's yeah, yeah, way better.

SPEAKER_01

So, like, you know, product owners could be like, you know, I would basically tell them the difference between one day of data retention and two weeks of data retention is almost nothing. Um, and so cust uh products could have higher retention by default. Um, you didn't have to pre-index your field. So if you're in an incident, you saw something weird, you needed to go look at some historical data during the incident, you just could. Uh, and that started working. We were able to add more powerful auto-complete. Uh and obviously the the company saved just you know a ton of money not running this super expensive um you know, kind of SSD-based system for what is essentially the lowest value per byte data on the planet, right? Which is like you know, random metrics and logs coming out of your software, right?

SPEAKER_00

Yep, the exhaust. And uh what year was this?

SPEAKER_01

Uh that's a good question. What year uh was how old is Warp Stream? Warp Stream's like two and a half years old. Uh so I guess this would have been like five years ago.

SPEAKER_00

And and yeah, like Warp Stream's, I guess that's really the the thing. Warp stream's two and a half. Um and now this is late 2025 when we're recording this. Um Diskless whatever is the cool thing. You know, if if there's a kind of data infrastructure, someone is working, and probably two or three someones are working on a diskless version of it, yeah. Um search various kinds of databases, uh certainly, certainly uh distributed logs. Um that wasn't true then. That was a you know, you try to do this all diskless. That that was a much more much riskier choice for you and direction for you. And it's it's funny, your story kind of pivots on that that all is lost moment of oh crap, maybe I did the wrong thing and we have to go back to doing this the lame way. Um and you didn't, you pulled it out, but that's uh these days, like if that's happening in in 2025, early 2026, oh you want to do some discourse, big deal. You know, everybody does that. It's hard, but it's it's popular. It just wasn't then, and I I just I'm pointing out that that was uh um an admirable and slightly revolutionary approach.

SPEAKER_01

Well that and that was I think less uh I don't know what word to use, unhinged maybe, than uh than when Warfream actually, because like I think at that time, you know, we were kind of following in the footsteps of like like you know, the Snowflake paper was published. People knew that building analytical databases and columnar stores on top of object storage was possible. Um people thought maybe, oh, it'll have to be like a cold store. Because that's how the project started. People were like, well, you're gonna build like the cold tier or some special really slow version of the logs product that's that's cheaper or whatever. And then like six months into the project, like our engineering, one of our engineering leaders was like, no, you need to make this work for every product and to get rid of the old system to power all the real-time data as well. And I was like, Ah man, that's gonna be hard. Um, because there's you know, we'd already set some latency expectations with the the the old product. Um and I think warp stream was even more extreme that like people, you know, the the way people used to describe it was an aggressive architectural decision or whatever.

SPEAKER_00

I've probably used those words myself.

SPEAKER_01

Yeah, because it's you think of it being as a much more you can kind of imagine, okay, like if a query takes a second to run to answer a human, that seems more reasonable than like a computer waiting 500 milliseconds to make sure something is durable or whatever. Right. Um but I don't know. I think if you're working on stuff that's like really hard or bleeding edge, at some point you will find yourself questioning, like you will have a at some point you'll have a come to Jesus moment where you're like, I don't know if this is ever going to work. Exactly.

SPEAKER_00

Am I the stupid one?

unknown

Yeah.

SPEAKER_01

Like I I've had that moment in pretty much every major product I've project I've ever worked on. Yeah. Just because it's like if you're doing something ambitious, you will eventually reach a point where you're like, I don't know if it's gonna work. Like it the the napkin math said it would, but I'm not getting the results yet.

SPEAKER_00

You know, uh as I like to say, it's hard to make things, and if you're not struggling a little bit while you're making them, you're you're you know, maybe not making something that is is all that interesting. Yeah, I agree with that. My guest today has been Richie Artoul. Richie, thanks so much for being a part of the Confluent Developer Podcast.

SPEAKER_01

Thanks for having me, too.