Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov
Hi, we’re Tim Berglund, Adi Polak, and Viktor Gamov and we’re excited to bring you the Confluent Developer podcast (formerly “Streaming Audio.”) Our hand-crafted weekly episodes feature in-depth interviews with our community of software developers (actual human beings - not AI) talking about some of the most interesting challenges they’ve faced in their careers. We aim to explore the conditions that gave rise to each person’s technical hurdles, as well as how their experiences transformed their understanding and approach to building systems.
Whether you’re a seasoned open source data streaming engineer, or just someone who’s interested in learning more about Apache Kafka®, Apache Flink® and real-time data, we hope you’ll appreciate the stories, the discussion, and our effort to bring you a high-quality show worth your time.
Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov
155 Contributors Later: Where Kafka Is Headed ft. Andrew Schofield | Ep. 34
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
Tim Berglund talks to Andrew Schofield (Confluent) about his career in Apache Kafka. Andrew’s first job: working on queuing systems in 1991. His challenge: working at Confluent and in the Kafka community to bring queue semantics into Kafka while also helping shape major efforts like diskless Kafka, disaster recovery, and the project’s broader evolution.
► Queues for Kafka Explained (KIP-932): https://youtu.be/Wb0xyqgaIqw
SEASON 2
Hosted by Tim Berglund, Adi Polak and Viktor Gamov
Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
Music by Coastal Kites
Artwork by Phil Vo
- 🎧 Subscribe to Confluent Developer wherever you listen to podcasts.
- ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
- 👍 If you enjoyed this, please leave us a rating.
- 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
In Kafka's early days, nobody cared because it was all done as a data center, so you didn't pay for the network charges, but you didn't do with the cloud. There were 155 contributors to AK 4.2, which is a lot, yes. So if you imagine trying to convince a company to have 155 engineers in an engineering team, they you just can't do it, right? So it's really good you've got this kind of collaboration.
SPEAKER_02Welcome to the Confluent Developer Podcast. I am your host, Tim Berglund. Today we had a good old-fashioned open source Apache Kafka episode. I got to speak with Andrew Schofield, who is a principal engineer at Confluent and a committer and PMC member on the Kafka project. He's been a key figure in the Kafka space for like a decade at least, and he's really got a passion for the community. We talked about cues for Kafka, and if you don't know that feature, check out the show notes for a link to a great explanation, as well as Diskless Kafka and more. Kafka really is a vibrant project, and people like Andrew are a big reason why. So let's listen in. Hi, Tim. Good to see you. Yeah, good to have you on. Um I figured, guy like you, it would be good to talk about what's going on in the world of Apache Kafka. Uh there are a lot of really interesting things in flight right now. You got your hand in some of them, and some of them are are uh just things happening in the community. What um I feel like having Andrew Schofield on the podcast and not talking about cues would be weird. So do you want to start with cues?
SPEAKER_01Sure. Let's let's talk about cues. So um I've been working on cues for a long time. So, you know, I start I started working on queuing systems in 1991 because I'm that old. So I'm with you.
SPEAKER_02I've not that I was working on queuing systems then, but maybe like there was probably like a real-time operating system that had a queue in between tasks that around that age, that that year I was I was maybe working with. Uh, but yeah, queuing systems pre-Kafka.
SPEAKER_01Yeah, that's that that's right. So, you know, I I kind of I I know my stuff when it comes to queues. And the uh, you know, I started working in Kafka in 2015, I think, but not not hands-on on the code at that point, right? More kind of uh leading a team, leading a team, productising it, that sort of thing, right? And then over time I thought, no, no, I actually want to get my hands on the code and start changing it. Yeah. And I observed when I first started working with uh with customers on Kafka that they they struggled with uh trying to use it like a queue, right? They would they would they would be porting queuing applications over to Kafka and they would go, well, it's it's really difficult. And you know, the the the initial discussion you have with them is well, don't use it like a queue then. And that works to an extent, but then sometimes people, you know, they they just say, Well, no, I've got a kind of a queuing shape of problem. So I saw that a very long time ago, but I wasn't sure I wanted to solve it in Kafka to start with.
SPEAKER_02Right. And I I I've certainly taken that question on a number of occasions the you know, how do you make it act like a cue? And do you think it was because they actually had queuing use cases, or it was just habituation? You know, cues are all they've had for a career and and they're porting things and they they just can't free their minds. I'm like, what I mean, obviously, we're building cues for ka, you're building cues for Kafka now, so there's something there. But um what do you think about that?
SPEAKER_01I think I think it's it's a mixture, I think. So so quite often, yes, they were just used to it, right? So they just tried to port something verbatim, and then they ended up having, you know, thousands of partitions for no good reason and that kind of thing. So they kind of got themselves into a mess. And sometimes they just needed to uh, you know, kind of rethink how they did things because it was quite common in the you know in 2015, 2017 kind of era that somebody would come in with an architecture hat on and saying you're gonna use Kafka and then everyone would salute and and and rush off after it, right? Right. So they didn't really have the skills yet, right? And that's not true these days because you know lots of there's lots of Kafka skills around. But I think people just tried to replicate what they previously done, not necessarily for a good reason. Sure, but just because it was and and you you have to you know maybe think quite deeply if you want to change to you know from a message queue model to an event streaming one, and not everything kind of translates across, you know, easily, I think.
SPEAKER_02Could I ask you to dig into that? I'm sorry, you're you're you're trying to get warmed up on cues for Kafka and I'm like, well, let me ask you this. Fundamentally, um, you know, what what's the insight? I I think a lot of people listening to this have an intuition about what a cue is and what a log is and the difference. Uh, but if pressed, I don't know if everybody could give a really crisp answer. So um what's a cue and what's the difference uh between Kafka and a Q on the most fundamental level you can describe?
SPEAKER_01Yeah, so um I think it'd say between Kafka and an event stream, probably, right? So an event stream, you you you really are getting a uh you know a bound an unbounded sequence of events, they're in order, right? You're kind of you may well depend upon the order, and uh you intend to consume all of them, you know, all of the ones that you're given. And I know that your consumer group divides your topic up and and gives separate partitions to the different members of the group, but you still, the subset of the topic you'll be given, you're gonna get all of them, and you have to be able to keep up, you have to process them all yourself. Okay, you can't easily share them. And if you wanted to uh you know have have more consumers, you need more partitions. That's just the way that Kafka works. It slices everything up and then gives you the whole partition. Whereas message queue is typically much more individual items, right? You're processing them one at a time, right? And and there aren't necessarily links between them. Now, that isn't true in all cases, but quite often it's some of like a command, go and do this, right? And you don't necessarily need to know what the previous three were. You just kind of process it by itself. And if you've got a really kind of granular workload like that, being given the consumer group semantics is kind of inconvenient. You don't really want it like that because you have to, you know, with the consumer group, you commit your offset, essentially your current position that you're going to be consuming from, and you have to do it for every single record if you're gonna do it, you know, in a in a kind of a queuing way. And and it wasn't really designed for that. So uh, you know, it's it it's it's all about the kind of the granularity of the model that you have in mind, I think.
SPEAKER_02Okay. Okay. The um and consumer groups being very much baked into Kafka traditionally, uh, they did get in the way, right? Uh of Q behavior.
SPEAKER_01They did. Yeah. They did, they did. If some if somebody really wanted a Q, then uh then it was it was kind of uh you know a bit of a nuisance. So um so then I thought, well, okay, well what's that what's the best way to to uh represent this? Because you know, I kicked it around for quite a long time before we ended up with uh with with the kip uh in the community, and that was I think the kip was published nearly three years ago now, something like that. So it was just before a Kafka Summit London, I think in something like May in you know three years ago. I think that was about right. Uh and you know, we tried various things and tried different kinds of APIs like JMS and that sort of thing. So I did quite a few experiments with it, and then eventually thought, well, actually, maybe what you want to do is just make it a new kind of consumer in Kafka. That would be it, because people were building things, including me, were building things on top of Kafka to make it behave like a queue. And they just kind of you you you had an impedance mismatch there, and we were kind of you know solving it in in code on top of Kafka, but then you have to work out, you know, what's the life cycle of it, who manages it, and all this kind of stuff. So it kind of gets complicated doing like that. Uh, and then I thought, well, all right, well, what if we were actually to put it in Kafka as a new kind of group? What would it actually need to do? And you know, one of the big decisions was don't invent a queue in Kafka, right? With the share groups, you're still on topics, your data is on topics, you're just consuming the topic in a different way, because we could have made the thing called the queue. But you know, nearly everything works with topics these days. You know, if you think about uh the way schemas are configured, the way uh Kafka Connect works, all of this stuff, it's all topic-based. So it really needs to be using topics. That's right. Yeah. That's right.
SPEAKER_02And AI is about a new thing, you know, that's there's a big runway there, right? Like everybody knows topics, it's fine.
SPEAKER_01Yeah, yeah, that's that that's exactly right. And then it was all right, well, if we're if we're not actually going to create cues, what's kind of the essence of a queue? What do you actually need to make it behave like one, even though you haven't got a queue? So, you know, you need an API which has kind of record as a time behavior. You don't need to, you know, you don't need to make it laborious. You can still do things in batches if you want to, but you need you need that, and you need to be able to uh kind of give a you know a disposition, an opinion on every record that you're given, right? Did I process it, did I not process it, is it a bad one, that kind of thing. And all the different systems, they do this stuff in different ways, but they've fundamentally got the same kind of idea, which is essentially a kind of a state machine per record or per message, and and the system is you know coordinating all of these and giving the behavior of a queue. So that's kind of what we did.
SPEAKER_02Over uh your history of queuing systems, our concept of scale, uh the the popular conception of scale has changed quite a bit from 1991 to the present day. Um and I the there were to be doing things at sort of global scale in 1991, you had to be like a utility, a bank, a government, something like that, airline maybe. Um and and now post-internet, post-mobile devices, that's it's more commonplace. Not all of us build huge systems like that. But is there a difference in the scaling characteristics? So all that to say it's a little bit unfair to ask if the legacy message cues could scale like Kafka, because I don't know if they were built to, but is there something fundamental about the way they were built that just made them expect to be a little slower because of the message at a time acknowledgement? I mean, is there anything in there that has been a challenge scale-wise in building cues under Kafka?
SPEAKER_01I think that they have they have a pretty different model typically, right? Uh in terms of the way they handle storage. So they don't really expect to have a lot of backlog. Whereas in a Kafka system, one of the things about it is it's effectively unlimited, right? You don't care. Right. Uh so they're quite different in terms of that. Um, and another thing, you know, I I think the probably the area where Qs of the Kafka is kind of uh, you know, least comparable to some of the messaging systems, other messaging systems at the moment is in the area of transactions, right? Because we we can uh we can process data that was that was written using transactions, but you can't acknowledge it using a transaction at this point. You know, there's a kit for that, but it's not done yet. And it's it's much more difficult to do that when you have your kind of your state distributed across multiple logs. Whereas in a in a traditional messaging, message queuing system, you've got one log per server, right? And everything goes in there. So you can do coordination of all of your operations across your queues or whatever, and then to the right commit once, and then it happens, right? But then you've got all of the data going through a single log. So it's kind of swings and roundabouts. Um, and then what you you tend to find is that the uh the message queuing systems come up with some kind of clustering. Okay, so they they have maybe some kind of directory which says a queue of this name is actually hosted by multiple servers, but they're kind of still independent things, and each one of the servers still has its own singular log. So, you know, you can slice it and dice it in several ways, I suppose, would be the thing.
SPEAKER_02Um yeah. Now a quick word from our sponsor. Confluent Developer the Podcast is brought to you by Confluent Developer the website, which has everything you need as a developer of data streaming systems. And it's completely free. We've got curriculum, hands-on exercises, executable tutorials, the online data streaming engineer certification, also free, a way to find a meetup near you, those are free. Everything is there. I really want you to be successful in your journey as a data streaming engineer, and this is the site that has what you need. Check it out at developer.confluent.io. That's developer.confluent.io. Now back to the show. They're all they're all single leader systems clustered like you might shard Postgres or Mongo or something like that.
SPEAKER_01Um something like that. And you're you're finding that uh that they sometimes they're taking uh other consensus out, you know, protocols, algorithms from the world of you know, software engineering, uh, like craft and that kind of thing, and building something a bit more Kafka-ish, but still for a single server, just so that because the high availability, the availability characteristics of these things were you know, they they kind of looked a bit 1991, right? Really, where you get a crash of the server, you start it up again on some storage, which you copied, and then whereas there wasn't really any, you know, there wasn't any consensus, whereas there are a few other systems which are now picking up that as well. So they look like um, I don't know, little islands of consensus, I suppose, but you still have to, you've still got individual cues on each of these servers. So Kafka is a bit more, a bit more kind of distributed than that, right? Yes. Um, you know, you can have a big cluster that has all of these availability characteristics for each of the partitions within it, right? Without having to manage them all kind of independently and have you know lots of separate quora and that kind of thing. Right.
SPEAKER_02Um and so where are we right now? We're recording this in um the middle of March. Is it the odds? No, that's tomorrow. Middle of March uh 2026. And um what's the current state of cues for Kafka? Uh what's shipped, what hasn't? Yeah, bring us bring us up to date.
SPEAKER_01Yeah, so um we we did it as a three-stage delivery. So there was there was um you know an early access in 4.0, and then a preview in 4.1, and then uh and then um uh general availability in 4.2. So uh Apache Kafka 4.2 has GA of Q's. Only just a little bit ago. And so right. Oh, fairly recent. Yeah, absolutely fairly recent. Um and then 4.3, which is coming along pretty soon. So the code freeze for 4.3 is only two days away. Oh, goodness. So we we we've just got a little feature in that to um give you a bit more kind of configuration options. But the next big thing is is dead letter cues, that's 4.4, I expect. So engineering of that started, you know, we haven't approved kip, so we're working on that, but um landing it in 4.3 wasn't gonna happen. So that was the that was kind of the biggest thing that people said when we talked to them about it was okay, well, what do we do if something doesn't work? You know, uh in um in the original cues for Kafka KIP, it basically said we will perform a limited number of retries and then it comes off the queue because it will get in your way. Because that's always been one of the problems with Kafka. You get a poison message, how to jump over it.
SPEAKER_02This is from the side that I am reaching some computational failure processing the thing.
SPEAKER_01That's right, that's right. So a well-written application can handle a you know a badly formed message, it can, you know, exception handling and it can, you know, but but if you don't, then it's in your way, right? So we took the decision with use for Kafka that wouldn't be the case, and the message comes out of your way, but then there's well, what do I do with it? How do I find it? You know, you're you're kind of you've potentially lost something which you should have caught. So uh dead letter cues give you an you know another topic, and if you get a message which is not passed, it's you know not delivered successfully, it gets copied onto the dead letter cue topic, and you can do, you know, use it like a black box recorder or whatever. Yeah. So that's kind of the next thing. And you know, that'll uh so 4.4 towards the end of this year, but still this year. So I think that'll be the the most kind of requested um thing that we didn't have in the original delivery. So that's good.
SPEAKER_02Okay. All right, awesome. Do you see much? I mean, in in the past, you know, you said you worked early on with in Kafka with uh people wanting queuing behavior, either because they really needed it or because you know they just they just couldn't couldn't break free of old habits, whatever that is. Um do you have much touch these days with like what folks are doing, use cases and and things going on in the wild? I I realized as a as an engineer and PMC member and everything, sometimes you know, you're not hands-on with those folks, but yeah, I I think the honest answer is not as much as I used to.
SPEAKER_01Um, I think, you know, I because you know uh different different people working in the community have kind of different uh you know things they work on. So I tend to be quite community focused these days, right? So, you know, I do a lot of kind of kit reviews, kit writing, uh, you know, looking at other people's PRs, all this kind of stuff, right? But it's more technology focused and less kind of customer focused. But I work with people who do more of that. But I I used to get more customer visibility in my previous job anyway, because I spent a lot of time talking to customers directly. Um, you know, I would I would get flown over to places and have to give a presentation on Kafka and I get asked all difficult questions, but I don't get that directly these days, I suppose.
SPEAKER_02I well, yeah, that's I know I know that life. Um speaking of you being community focused, which you are, and it is absolutely wonderful. You do a lot of great work. Um, what else is going on apart from cues? What are some other things that are that are happening in Kafka? I mean, I I think it seems like it's as vibrant a project as ever. And uh I love that. So what what do you see?
SPEAKER_01Yeah, I I think it is it is very vibrant. Um, and you know, I I think there are there are there are lots of vendors who are who are contributing significant things to each days, which is which is good to see, right? Uh so there were, I think, 155 contributors to AK 4.2, which is a lot.
SPEAKER_00Yes.
SPEAKER_01So if you imagine trying to convince a company to have 155 engineers and an engineering team, they you just can't do it, right? So it's really good you've got this kind of collaboration. Um, and there's some there are some significant kips. One of them is more than more than just one kip. So one of the uh significant things is a project called Disclus, uh, which is being led by a company called Ivan. Uh, and you know, they that the the the principle behind that is running Kafka on cloud storage is expensive because you replicate data between the brokers and you pay for the network charges between the brokers. So you don't want to do it like that, right? Uh and you know, in Kafka's early days, nobody cared because it was all done inside a data center. So you didn't pay for the network charges, but you didn't do with the cloud. So the idea is essentially we can get rid of that by using cloud storage rather than local disks in the cloud, and then you don't need to replicate it in the same way, right? You'll just find that it's effectively free. Um, now that is a that's a pretty significant thing to put into Kafka. Um, and it's going to be a family of KIPS, right? And the first one, um, so um 11 1150 has been uh voted through. So that was kind of a foundational one. So, you know, you probably remember when craft was done, but there was kip 500, and it was basically we're gonna put craft in. Great, isn't it? And then it didn't say anything much more than that, right? And then there were, you know, eight kips.
SPEAKER_02I was doing podcasts about that for years. Like yes, a season of my life. I'm sorry, but it's not that it took, it was just it's I mean, you know, uh I if if I don't like that, I should go make my own consensus protocol. It's hard to do, you know, it takes a long time.
SPEAKER_01It is, it is. So um so 1150 is voted, right? Uh, and and that was really good, right? Because uh it so when you when you write a kip, you need to get at least three committers to vote for it, and that's basically a you know, go ahead, you can start working it, right? Gotcha. Whereas there were nine voters on on the disclus, which is a lot. So it was a broad, broad spread across a lot of voters. That's right. So the the the community basically said we want this, right? Um now, now you get down to the nitty-gritty, right? Because now there are the real kips behind it, uh, and they are significant in their own right. Um, they are still under active discussion. Uh, I think they've probably got a fair way to go. But when, you know, giving them this level of scrutiny is good because it means that when they land, you know, this community of people kind of collaborating in the open, they can still achieve great things, right? But it means that you need to give things scrutiny early on in order to make sure that things are going to be sensible. You know, it's it's more difficult to manage things like this than it is inside a company. So, um, yeah, there there are there are two fairly significant kips which are under investigation or under uh under discussion. Um, I think they've they've probably got a fair way to go yet, but they're in they're in pretty good shape in terms of you can you can kind of read and understand what they mean. Whereas with some some big kips, it takes a while before anyone outside the author can actually really understand them, but they're they're they're getting there. Um and until we get some of those voted through, then the code won't actually start entering Kafka. We know we need approval that this technical change can be made, and then they'll start getting merged. But they're still they're still under discussion, those ones.
SPEAKER_02You know, it's funny as a uh as a as an outsider, um, you know, not a committer, I find that most KIPS, if I'm just reading the description, you know, I'm not really trying to follow the detail of the design, but if I'm just reading the description and trying to understand, okay, what is this doing to the system? What are the implications for you know the surface area on the outside of the product, architecturally, what would be implications and how you'd use Kafka and all that? I I have to say. Most kips at that level, right? Um, I find to be pretty friendly. Sometimes there'll be a link to a longer doc that you have to go read or something like that. I imagine as an insider, you're like, yeah, no, no, Tim, they're just they're not that good. And all the other committers are are probably agreeing with you and each other, you know. But I the the they're I find them to be well written. I I just I just want to say that. Because I do have to go through those on a on a semi-regular basis, you know, make sure I understand what's going on.
SPEAKER_01Yeah, I think I think they typically are. Um when when you're reviewing them as a as a kind of somebody who will vote on them, then you basically are looking, you're looking for holes, right? What didn't you say? So they are typically well written, and you can kind of you know and understand them relatively quickly. But uh, you know, what what kind of things did you miss? So an example with the Disco stuff, and and I, you know, I I need to go and reread those ones now because I read them a while back and put some comments in, but I need to go deeply into it again. Um, one of the problems was uh how do you garbage collect um the cloud storage reliably? Because what you do in Discours is you effectively you you wait for a little bit, um, you chunk up a load of data, and then you you put it into object storage, and then you've got a handle to this object storage thing, and you this is basically a pointer, they call it batch coordinates off to the records. But you could crash halfway through this. So you then offered something. How do you find it? You know, that sort of thing. So um, you know, there are there are plenty of things in there because you know, you uh I imagine that when this lands, it'll be a multi, a multi-release thing again, you know, maybe maybe three releases like with use of Kafka, I expect. Oh, yeah. Maybe they don't entirely solve the garbage collection to start with. But before GA it has to be, and last time I looked, there they wasn't really, you know, there was discussion about it, but not a solution yet. So this, you know, it's that's that's how you kind of re review them when you're when you're when you're going to vote on them anyway, is what what could go wrong, what what hasn't been thought about, you know. Uh and and I'll I'll talk about another one in a minute because there's another kind of uh you know kind of wrinkle in that one as well. Um so another another significant one, which is actually this one has moved really fast. So there's one called cluster mirroring. So this is for disaster recovery, it's for linking clusters together and replicating data asynchronously between them in case you get a volcano go off on your data center or whatever, right? So it's it's a useful thing.
SPEAKER_02Version version three.
SPEAKER_01I mean, this this is a thing that has been done. I well, I think I think calling it Mirror Maker 3 is probably rude, even I don't know like that.
SPEAKER_02It occurred to me when I said that that we don't want to call it Mirror Maker 3. That's not very nice. But it's my point being it is the third uh at bat here uh to solve this problem. And tell us why it's the good one.
SPEAKER_01Yeah, so this is the good one because it's it's built into the brokers. Okay, it doesn't, it isn't Kafka Connect, right? It's not Kafka Connect. It's basically a pump and uh running actually inside the broker. And you configure it using a mirror command, there are mirror RPCs in the protocol, all this kind of stuff. So it becomes a mirror as a proper object now, rather than just a kind of a thing running on the side. Uh and this this is uh again an interesting one, right? It's uh you know, it's uh um I don't know, not exactly a collaboration between multiple vendors. So there are six authors on the kip, which is unusual, but it's but it's a really nicely written kit, right? It's it's not uh it's not totally finished again yet, but it is it's certainly, you know, it's definitely ready for proper scrutiny, right? So um, you know, the the stage of review of this kit, I think, is pretty mature, right? Um you can believe that if they implemented it as it said, then it would basically work with you know a few a few mistakes while we're sorting out the the review comments, but it's it's basically there. And it's only been a month since it was published, right? So this one has moved really fast.
SPEAKER_02Oh, right.
SPEAKER_01And I imagine that this'll this'll do another kind of multi-release thing, but you could if they manage to get the kit the discussion finished relatively quickly, you could imagine the first bits appear in 4.4, right? So it's it's moving pretty quickly, that one.
SPEAKER_02Okay.
SPEAKER_01Okay, very nice.
SPEAKER_02Yeah. And yeah, usually with with big multi-release or even multi-kip kips like that, because that's a significant feature, very complex. Uh, like the first release, you might get either a, you know, it'll it'll work. You know, you can use it locally to debug, or here are some APIs we put in place and started to deprecate some other APIs because, you know, laying the groundwork, but then you get a few releases later and you have the fully baked feature.
SPEAKER_01Yeah, yeah, yeah. Yeah. So that one, that one's got, you know, some interesting stuff as well, right? Because it's it's it's effectively puts the configuration for talking to another cluster in this cluster, right? Which we haven't kind of done before. So I I imagine that there'll have to be uh you know, maybe a follow-on kit for how to make sure that's done securely and that sort of thing. You know, so you know, um Kafka Connect has always had this sort of thing in config providers, and maybe we need something similar when mirroring is you know finally finished. But when we did when we did cues for Kafka, there was one big kip, 932. And uh so we delivered that in 404142, but 4.2 had another, I think, four or five kips. So things where we were maybe they would have got a 932, but we realized the error of our way if we kind of put some you know addenda in there essentially. I can well imagine that you do that custom mirroring as well, and just to you know close a few things off. Absolutely. But it's it's quite an effective way of doing it.
SPEAKER_02Yeah, no, I think it it is. And I mean, I think we all know it's very difficult to know uh what it is you are actually going to build, even when you know the direction you're headed. Uh, you learn along the way that there are things that you were just wrong about or ignorant of and didn't imagine. And uh, you know, those are those are the follow-ons. That's uh that's what it's like to build things.
SPEAKER_01It is, it it certainly is. It certainly is. Yeah. Because the uh you know the latest kit we put in for Kuse Kafka was to do with configurations, and it was just no, and if we've got them, we made some made some or left some gaps in Bine Street 2, we need to fill them in now, and it was it was quite cheap, but yeah, it was it had to be done. So boy, that's a lot.
SPEAKER_02Is there anything else that's that's big or even small on the uh so I think those are those are the big ones.
SPEAKER_01I suppose there's another one. Uh there's one to do with multi-tenancy as well, but that is that one. I I don't know. It's it's hard to kind of tell what's going on there. Uh in in that it was um there's there's quite a large document. Um I don't really feel that the discussion is moving forward at the moment, right? So I so my so my my my honest opinion is um we will end up with multi-tenancy in Apache Kafka, right? I don't really feel the kip which is out there at the moment will be it, but we will see, because it might it might migrate into it. But what people are doing is they are uh, you know, so the the the Kafka community is, you know, some people who use it and some people who sell it, essentially, right? Those are it's that combination because quite a few big companies use open source Apache Kafka internally, right? And they want to make it better in order to make their lives easier. So this is a good, you know, a good kind of uh contributor. And then there are others who uh they have a business selling Kafka and they want to make it better in order to have something better to sell. But at the end of the day, they're still trying to do the same thing. I mean, I which is provide something general purpose and useful.
SPEAKER_02So you and I work for a company like that.
SPEAKER_01So we do, absolutely, and it and it's perfectly fine because everyone knows it, right? But at the end of the day, you're trying to make something, you know, something more useful, something better. And and you know, so my my point of view on this is essentially we're trying to make Kafka more general purpose, right? We don't want to, you know, if you have to think you've got a problem to solve and you have to go, well, will it can I solve it with Kafka, or was it that doesn't quite fit here, you know, maybe I don't know, large messages or something like that, right? There are some things it doesn't quite support yet, which if it did, then you would then go, well, no, actually we can do we can do more stuff now. And building in DR is like that, right? Providing support for cloud storage is like that. Multi-tenancy is like that, because you see people um who've got one shared Kafka cluster and they kind of use name prefixing in order to make sure that different people get to access different topics. But if you could actually say everyone thinks they have a cluster and and they can't see the kind of divisions between it, but everyone, you know, everyone can have a topic called T1, but they're actually all on the same piece of physical hardware. That kind of would be good. And so I think multi-tenancy will land one day. I don't think I see the kip yet, but we will see. We will see.
unknownYeah.
SPEAKER_02My guest today has been Andrew Schofield. Andrew, thanks for being a part of the Conflict Developer Podcast.
SPEAKER_01You're welcome. Thanks.