Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov
Hi, we’re Tim Berglund, Adi Polak, and Viktor Gamov and we’re excited to bring you the Confluent Developer podcast (formerly “Streaming Audio.”) Our hand-crafted weekly episodes feature in-depth interviews with our community of software developers (actual human beings - not AI) talking about some of the most interesting challenges they’ve faced in their careers. We aim to explore the conditions that gave rise to each person’s technical hurdles, as well as how their experiences transformed their understanding and approach to building systems.
Whether you’re a seasoned open source data streaming engineer, or just someone who’s interested in learning more about Apache Kafka®, Apache Flink® and real-time data, we hope you’ll appreciate the stories, the discussion, and our effort to bring you a high-quality show worth your time.
Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov
Schema Registry Made Simple by Confluent Cloud ft. Magesh Nandakumar
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
Tim Berglund and Magesh Nandakumar (Software Engineer, Confluent) discuss why schemas matter for building systems on Apache Kafka®, and how Confluent Schema Registry helps with the problem. They talk about how Schema Registry works, how you can collaborate around schema change through `avsc` files, and what it means for this to be available in Confluent Cloud today.
EPISODE LINKS
- Schema Registry 101
- Schema Management
- Migrate Schemas to Confluent Cloud
- Schemas, Contracts, and Compatibility
- Fully managed Apache Kafka as a service! Try free.
SEASON 2
Hosted by Tim Berglund, Adi Polak and Viktor Gamov
Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
Music by Coastal Kites
Artwork by Phil Vo
- 🎧 Subscribe to Confluent Developer wherever you listen to podcasts.
- ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
- 👍 If you enjoyed this, please leave us a rating.
- 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
Apache Kafka is famously unconcerned about the type of the things you put in it. When you produce a message to a topic, it's just stored as raw bytes. Kafka doesn't care what's in there. Of course, you do, which is why Confluent made the schema registry. Schema Registry has recently become a standard feature in Confluent Cloud. So I invited lead developer Magesh Nanda Kumar to tell us all about it on today's episode of Streaming Audio, a podcast about Kafka, Confluent, and in this case, Confluent Cloud. Welcome back, everybody. My guest today on Streaming Audio is Magesh Nanta Gumar. Magesh, welcome to the show. Hey Tim. It's a pleasure to be on your show. Now, Magesh and I are coworkers, and uh but not all that closely. We don't work together uh too often. But what do you do here, Magesh?
SPEAKER_00Um I work as part of the engineering team. Uh I primarily work on schema registry and connect. Um so I've been working on schema registry when I started with Confluent and then I transitioned into the Connect team and started doing schema registry, connect, connectors.
SPEAKER_01Nice, nice. So we've talked about Connect on the show before, and I'm trying to remember. My my memory of this isn't very good. I don't think we've talked about schema registry directly before. I'm sure there will be somebody on the internet who tells me if I'm wrong, and we have, and that's good to get that kind of feedback. So um uh nominally, what I want to talk to you about today is schema registry in Confluent Cloud. Now, as of the time of this recording, it's a relatively recent feature that Confluent Cloud has has rolled out. Um by the time this goes live, it'll be uh, you know, only only a few weeks less recent. But um fairly new feature, but maybe let's just start by talking about schema registry itself, because uh you know, schema registry being in the cloud is sort of a thing that we can say, and that's that that subject is kind of done at that point. But we we should talk through uh schema registry itself. Uh what is it? Imagine if somebody in the audience just didn't know it, they knew Kafka, but they didn't know schema registry. What's it all about?
SPEAKER_00Yeah, yep. Uh that I think is a great start to uh get going. Um so if we look at actually Kafka, Kafka doesn't really care what we send. Uh the producers can just send anything. And uh the protocol itself is just carrying about uh the bytes that's being sent. Uh this makes Kafka really powerful just because you can send any kind of data into it. Um this also means that uh your architecture is completely decoupled, uh, be it being your pipelines or be it your event-driven architecture, uh, it's like completely decoupled architecture. Uh but this also leads to a few challenges. Uh let's say if you're a typical developer, uh you're building a REST API. Uh, you start with having an API interface, and you typically document it in Swagger or some kind of tool, and uh that's how you get going with it. Or let's say if you're building a back-end system, uh, you typically connect to a database and you have some kind of schemas that are defined against the database, and that's your contract. But when it comes to Kafka, since we are just dealing with bytes, there's no real contract here. So what really happens is like uh if let's say if there's a producer who is producing a data, the consumer would now need to understand what's in it to go like uh process the data. Um so if they really have to do this, uh then there's only two options. They either have to infer it themselves or go talk to the producer and understand what the data structure is. Uh, this would mean that there's a lot of coupling between the teams, uh, so which is not really desirable in a big enterprise or even if you're a small company, uh, because like a decoupled architecture should also mean it's decoupled teams. Uh that just doesn't mean that you cannot socialize. It just means that you don't want to like coordinate over a schema like every time there's a change to it.
SPEAKER_01That's a that's a really good point. Decoupled decouple doesn't mean that you don't like each other. Decoupled doesn't mean that it means that you don't absolutely have to talk in order to advance your separate projects. It's you know, yeah, you you get you get together, you have lunch, you enjoy each other's company, you send each other memes on Slack during the day. But when it comes time to actually make a commit, you don't have to ask permission.
SPEAKER_00Yep, yep, absolutely. Um so schemas are generally a great way to solve this problem. Um so that's why uh uh something like schema registry uh is pretty useful. Um so what we recommend in con uh from Confluent is like uh using uh the Confluent schema registry, use Avro as your uh data type. Uh there are a couple of reasons why we uh recommend using Avro. First of all, Avro is like a binary format, it's not just text you are sending. That's that gives you a lot of uh uh saving when you send data over the wire. And Avro is also rich in terms of its schema definition and also its compatibility rules. Uh so what essentially schema registry is doing is like uh it's a tool that helps you define what your schema is for the stream that you're creating. Uh it can be a pipeline or uh like event-driven microservice, like I mentioned before. And uh the consumer would now just have to like chuck to a schema registry to understand what the data is and process it. Uh that's how schema registry solves this problem of uh tight coupling between the teams.
SPEAKER_01Gotcha. So um at the code level, you've you know, when I when I'm when I'm writing code, there uh there's a schema registry out there. It's some piece of infrastructure uh that is running outside my brokers, right? It's it's a it's a Kafka client. Um and I am writing code in some language to produce and consume to and from Kafka, say it's Java. Uh now the producer and the consumer are going to be the ones talking to the schema registry uh sort of invisibly to me, right?
SPEAKER_00Yes, that's absolutely correct. Uh so the way this happens is like uh uh Kafka in general, uh both producers and consumers have this notion of serializers and deserializers. Um so we ship what is known as an Avro serializer, uh, which helps you talk to schema registry. Uh so every time uh a producer comes in, they produce uh data, let's say, which is an Avro object, uh, then the serializer magically talks to schema registry, uh, registers the schema for you, and then it also serializes into uh byte format, which Kafka understands, and sends the data. When the data is sent to the Kafka topic, the schema ID is encoded as part of the message that's being sent into Kafka. Now the consumer would have to now, all they have to do is like configure the Avro D-Serializer, and they could just get the schema. The desirializer would now just get the schema ID out of the message, go talk to schema registry, get the schema, and then des-erize using the actual writer schema. So there's generally two uh kind of schemas when it comes to like uh schema processing and even streaming. There's a writer schema and a reader schema. You would want both the writer schema and reader schema in general to be compatible to each other just so that you don't break uh your downstream consumers.
SPEAKER_01Right. Uh and I want to get back to compatibility in a minute. That sounds important. Um and also good, I I like uh the terms that you've introduced, writer schema and reader schema, because if we're decoupled, and again, these applications, whatever application is right is producing to a topic, and whatever applications are consuming from the topic, uh, because there may be many of those, uh, they may evolve at different pace. They may they may have different release cycles. If if they're event-driven microservices, microservices, I'll put that a bit more strongly. They had darn well better have different release cycles so that we can evolve them separately and not have to uh uh version all my microservices when I change one thing. You know, that's that's the worst in the world. So you do have different schemas that you're writing and you're reading. Yep.
SPEAKER_00That's absolutely correct.
SPEAKER_01Yeah, no. Um the so when I and and out it's I always default to Java because it's me, but um uh as at at the producer level, I remember the first time I saw schema registry code uh that was that was using schema registry, something was a little disquieting to me uh because it looked like there was magic. There's a little config that you give the producer or the consumer telling it the URL of the schema registry's uh interface, you know, the URL, the schema registry. Um and then you kind of don't do anything, you don't like register the schema. It just felt like there should be some API interaction. Like, dang it, I should have to pick up a bucket of Avro data and move it from there to here or something, you know, it just felt like there should be code for that, and there really isn't. You just have um an object in in terms of your language, there'll be like a Java object.
SPEAKER_02Yep.
SPEAKER_01Um, and you you know, you program in terms of that object, you you change the state of it, you create one, you alter it, and then you produce it, and as you said, it gets serialized and and uh written into the topic.
SPEAKER_00Yep, yep, yep. Uh so when a developer comes in, like like you mentioned, it's just an object, it's a simple Pojo. Uh if you want to go back to the page.
SPEAKER_01It looks like a Pojo if you if you open the code, it yeah, yeah. There's nothing nothing P about it, but Yep.
SPEAKER_00Uh so it looks more like a Pojo. It's a generated class from uh Arrow schema uh generally, and then uh this utilizer does all the magic to extract the schema out of it and then go uh uh register the schema for you. Um so this is uh typically uh how you want to get started with schema IST. It's super simple. Uh it's gonna register the schema for you, and then you can get an application up and running in like a few minutes actually, technically. If you just go to the docs confluent IO, you'll see the code snippet there. Uh you can just go paste it in your application and you'll get started. That's such good advice.
SPEAKER_01Just go to docs.confluent.io. I love that. Um so true. But and I I want to underscore another thing there is that you've got an Avro schema file, an AVSC file, which is a JSON format that is basically Avro's IDL, right? It's it's way of describing the interface of the type. And there's uh a step in the build process. In the case of Java, it would be a Maven or a Gradle plugin that turn that into this instrumented code that looks like a Pojo. Yes. Uh and so the the answer when I first saw the schema registry code and I was upset because it looked like more should be happening. The answer is that more that is happening is in two places. One, it's in that generated code, and two, it's in the producer or the consumer itself, uh, where um it's it's configured. Well, I guess it's in the serializer is really where it is. Uh there's the Confluent Avro serializer that is um looking inside that object, getting the schema out of that generated object, because one of the things that gets generated into that object that's built from the AVS C file is a description of the schema itself. Yep. And all that just gets kind of pulled out of there and the serializer communicates the schema registry and and you know, magic happens. But the good news is like the at least in my opinion, and I'll tell you this as you know, you're a developer who works on this, I think the surface area that the developer sees there is actually kind of nice. It's it's uh it's a well-built API.
SPEAKER_00Yep, yep, yep. Uh so just because that uh the fact that you brought in ABSC, I also wanted to add a few things there, actually, uh just for our audience here. Um so AVSC, like you mentioned, is like the schema definition. I often tend to compare that to DDLs in a traditional database world, uh if people want to have a comparison. And uh like I said, like the serializer automatically uh does all the magic and registers the schema for you. Uh, but there are also users who generally do not want to do such things in production. Uh so there's a small bit of advice for them. Um so you could always auto-register your schemas in your dev environment and then get your schema registered uh against schema registry, get it working and all of that. After that, you could always take the AVSC file that you uh uh used in your development environment. And there's an amazing Maven plugin that we have to go uh register your schemas and manage the lifecycle of your schemas against schema registry. Um so you can set up your CI CD pipeline to use that Maven plugin and then um uh register your schemas for, like, say staging or production, uh and then you can just turn off auto-register in your producer applications.
SPEAKER_01Nice, nice. So that way if uh somebody slips a change in uh and rebuilds the code and redeploys a service that's producing to a schema managed topic, then it doesn't just automatically work. You know, there's a little bit of discipline around schemas.
SPEAKER_00Yes. So you you you should just uh I mean it it's not really essential, but then you could just turn off the auto register and then uh set up security around schema registry and uh uh really specify who can register schemas and who can read schemas and do all of that.
SPEAKER_01Nice. Nice, yeah. So it there's uh I think the the default path is this extremely developer-friendly. Go ahead and migrate the schema, it's fine. Uh everything just works. But it when when more discipline than that is called for, you've you've got the option. Which is good. Yep, yep. It's always good to have it both ways. And it's good for the default to be the easy way, in my opinion. So I just I like the way that's set up.
SPEAKER_02Yep.
SPEAKER_01Um, let's talk about compatibility, because that's another thing. I mean, schema schema migration is always bad, right? It's bad in relational databases, it's never not a problem. Um, and it's not like there's anything magical in Kafka that makes it uh automatic or anything, but schema registry helps with that. So you and you had mentioned that early on. You said you want to make sure that the uh reader schema is compatible with the writer schema so that you know when I'm when I'm getting data out of a topic, uh, I have some guarantee that uh even if I'm of a different version, that things will work. So just kind of walk us through compatibility and schema registry.
SPEAKER_00Yeah, sure. Um so scheme, like I mentioned mentioned, like uh your consumers and producers would have to be compatible when you make a schema change. Uh if I just want to walk through with a simple example, uh let's say you have a like address event that you're publishing as part of uh some changes in your application, and uh you want to, let's say, like uh change the format of uh how your zip code is being sent. Um so it could just be you're extending the number of uh characters that are allowed. Uh that to me is a compatible change. Um whereas if let's say if you have a timestamp column and if you want to change it from time in milliseconds to like a completely uh different uh time format, uh that to me is not a compatible change because it could potentially break your consumers. Um what we really offer here in schema registry is like different kinds of compatibility modes. Uh we have something called as backward compatible mode, uh, which would mean that like uh uh your consumer would have to be upgraded before your producer. And there are also other compatibility modes like uh forward compatibility and uh full compatibility mode. Um so the default is backward compatibility, uh, but for a real uh good architecture where uh you want uh disparate schemas between uh multiple producers and multiple consumers uh producing it to the same topic, the recommended approach, uh in my opinion, should be like the full compatibility, even though it's uh a lot stricter. Uh, it would provide you a better decoupling. Um, because you really cannot assume that there's always only one producer and one consumer. Um so you could actually set these compatibility levels at uh what we call as subjects uh in schema registry. Um so subjects are nothing but a logical container to evolve your uh schemas. Um basically uh within a given subject, you could have a schema registered, and that schema corresponds to a topic and the subject. It's the subject actually that corresponds to a topic. And uh when you have changes to your schema, you register a new schema against this subject. And that's the first thing that we do is like go check if the new schema is compatible with the latest schema by default. Uh and then if it's not compatible, you're not going to be able to register the schema and like you're stopped right there. If it is compatible, then the schema gets registered. There's a new version of the schema that's created. That doesn't mean that all the producer applications would have to continue, uh would have to like start using the new schema to produce data. That could still be producer applications uh which are like uh using the older schema. This is very, very common, especially when you're doing a zero downtime release in production. Uh, you could have instances of your producer application which is still producing data using the older schema. You could bring up newer instances of your application or service that's using the newer schema. So your topic now contains uh data with both old schema and new schema. Um so even if you do not have multiple producers, there's a situation where you would have like uh data with both the schemas and where the writer schema is different. Now, since we are pinning the schema ID with uh the message itself, uh, the consumers uh should be able to get the writer schema from it. And uh they will be able to like they have their reader schema, they should be able to deserialize uh using that schema. Uh, if they have to be able to do uh that successfully, then the schema should have been compatible. Uh of course there's an option to uh set a non-compatibility in schema registry, but I would uh almost recommend uh not setting it at all. Uh there are certain scenarios where you might want to do a non-compatibility, uh, like when you're like uh migrating from one schema type to another or if you have a breaking change. Uh even in those cases, I generally recommend using a new topic than setting a non-compatibility.
SPEAKER_01Right, right. Um indeed, for big and breaking migrations, a new topic is always a good thing. But but we we have uh you made a bunch of good points in there, like zero downtime deploys. You're always gonna have two producer instances at some point.
SPEAKER_00Yep.
SPEAKER_01Um that's just that's just how it uh it works. Even if your producer is a single application and it doesn't need to scale out beyond an instance, you're still gonna do that when you deploy new versions if you do zero downtime deploys. Um and in general, with a with decoupled readers and writers, which is very important, um, and not everybody is gonna be on the same version. It's just it's just how it has to be. Now, um to clarify one thing, the compatibility support in schema registry is is really compatibility checking. I mean, it doesn't do anything to migrate message formats, right? It's just telling you you may register this uh because it's forward compatible uh or it's backward compatible or it's fully compatible, or you may not, based on how you've configured the compatibility for that subject, right?
SPEAKER_00Yeah, that that that's perfect. Yep.
SPEAKER_01And a subject is uh uh basically like uh topic value or topic key, right? Yeah, is that correct?
SPEAKER_00That that's uh default uh subject to uh topic mapping. Yep.
SPEAKER_01Okay.
SPEAKER_00There are other options available, but most users just use this.
SPEAKER_01Ah, what are the other options? Here's this one of the great things about uh about doing this podcast is sometimes I learn things. And I, ladies and gentlemen, I'm about to learn something about scheme registry, and I know many of you are too. So yeah, I thought I thought, honestly, I thought subjects were just like the pairing of topic and key or topic and value.
SPEAKER_00Yeah, yeah. Um so uh we have something called as a subject name strategy. Um uh this is something that you could configure on your producer and your consumer. Uh the only downside of uh using a subject name strategy is that like all your applications would have to use the exact subject name strategy. Uh, but then uh the subject name strategy itself is a configurable thing. Uh let's say uh like if you want to like uh evolve, let's say, a user schema, and uh there are different applications or services uh producing to different topics using this user schema, but you still, from a data governance standpoint, you still want to evolve the user schema under one subject. Uh then there's a way to do that. The way you do that is like we have something called as a record name strategy. Uh basically, it takes the Abreu record name as your uh subject name. And uh you would have to just configure your producer application to like uh use the record name strategy. And uh consumer applications should also use the same record name strategy. Uh can this be standardized across the enterprise? Yes, uh there can be uh some standardization uh that can happen. And currently we do not enforce this uh uh uh anywhere, but yeah, it'll be good to enforce it somewhere. Um but you could also plug in your own subject name strategy. Let's say if you want to like uh hold a mapping between your topic and subject in some kind of uh uh registry of your own, and uh you could do that, and uh your subject name strategy can then resolve the subject name based on that uh registry or whatever, and then uh like provide the subject topic to subject mapping.
SPEAKER_01Okay, okay. So that's that's pluggable. That's uh that's good to know. But the conventionally subject means um subject would just mean topic, but the reality is that there are two things that go into topics which can have schemas, and those are keys and values.
SPEAKER_00Yeah, and it's not necessarily you have a schema for both key and value. Sometimes key could just be a string or an H.
SPEAKER_01I think frequently uh key is well, I don't want to say about frequently, I know people do it a lot of ways, but um it's it's it's not at all unusual for key to just be a uh primitive data type like a string or a number. Uh and you don't you don't manage schemas for those because um you know those schemas don't change.
unknownYep.
SPEAKER_01Um however, you know, string is also a little bit of a trap. Just so just knowing that you can schema manage keys, uh sometimes strings uh uh get interesting, right?
SPEAKER_00But I've also seen a lot of users, uh, even for their primitives, they prefer to use uh like uh Avro and uh like they use schema registry for it just so that uh no one changes the key type. And that's compatible as well.
SPEAKER_01Ah, there you go. Okay, so you have set compatibility and determine that that no new nobody can just go producing a new a new type there.
SPEAKER_00That's not especially if it's like a let's say if it's uh like double and tomorrow you don't want somebody to change that schema type to integer.
unknownRight?
SPEAKER_01Right, that makes sense. So to stress again, you know, when you're migrating schemas, I I started this question by saying, really acknowledging that look, schema migration is kind of bad anywhere. Uh it's it's always a little bit of a pain. And so when you need to change a schema, you gave the example of say going from six digits to ten digits on zip code. Um and even even you know, zip code is a little bit of a uh it's it's not exactly a localized concept, but suppose you know your addresses were all US addresses and you had been six digits and you want to go to ten or characters. Um that's a schema migration. You you do have to ensure that the actual code that reads values, uh zip code values out of consumed messages in the new version of the code, you know, you need to make sure that that code doesn't freak out when it sees six-digit zip codes.
SPEAKER_00Yep. And as a producer, you don't want to be testing all of that. You want to know right away if it's compatible or not.
SPEAKER_01Exactly, exactly. You don't want to put that on the producer, but there isn't you you still have a responsibility to attend to schema changes and attend to compatibility, but you get to now define, say, the values in the messages that go in this topic, uh, the schemas must always be forward compatible. And you you set that in the subject. And anytime you try to register a new schema, whether that's auto-register because somebody changed the Avro file and rebuilt the code and redeployed it and it auto-registered the new schema, or it was an intentional deployment process, automated deployment process that registered the new schema. Uh, if you create a new schema that's not going to be forward compatible, you know then it fails before you produce offending data. Uh that's the key thing. It's it's you, you still, at the consumer level, you still have to manage the fact that things might look different in your in your code, and that could mean conditionals and all the things that we don't like writing. Um, but you get a guarantee now that based on the compatibility rules you have decided on for that value in that topic, uh, you're not going to produce messages that violate those compatibility rules.
SPEAKER_02Yep.
SPEAKER_01I like that. Here's a question I get a lot. Um, and I feel like I never have an adequate answer just because I wasn't there at the time. But why Avro and not, you know, fill in the blank?
SPEAKER_00That's a great question. I get questioned about it quite a bit too. Uh that's a lot of uh ask. I mean, just to go back with history, I mean, uh, schema registry existed before I started here, but then uh uh just from talking to people and getting some tribal knowledge out of people here. Um the key reason is like uh Abro by default has this notion of compatibility. Uh Abro is very rich in terms of how it defines its compatibility rules, and uh it's very natural for people to use something like that. Uh but there are also other users uh who have been using JSON and Protobuff recently. Uh people prefer protobuf because it's very simple and uh it's very easy to use. And uh it also has rich uh like uh uh cross-language and cross-platform support. Um so we have been looking at options to add support for other types. I just cannot promise when it's gonna come. Uh, but we are actively looking at it. Uh I've I've actually spoken to a few people in the community about uh um adding uh support for protobuf and what they think uh that should be. So if you're a user and if you're if you think like uh you want support for either JSON or protobuf, and if you have thoughts on how that compatibility rule should be defined, uh just uh send us a note. Uh you could even uh post a note in uh some GitHub issue, and then we can collaborate there, actually.
SPEAKER_01Nice. I like that a lot. So uh that GitHub repo, um I'm gonna make that make sure that GitHub repo is in the show notes. Um and that would be a great thing to do. If you've got thoughts on how that should work, um please uh please contribute, even even just by by uh discussing an issue, even if you don't want to contribute code. Uh but that's good to hear that that is you know potentially a future thing. Like I said, we don't we don't talk about release dates here on this show. And if we did, it wouldn't be you and me talking about them. Yep. Um but it's it's at least being contemplated, uh it could be a future thing.
SPEAKER_02Yep.
SPEAKER_01Um and uh you know, we all know that's the kind of thing that everybody is never gonna be happy because protobuf is certainly the one I get asked about the most. Uh, but there will be some other serialization format that still isn't included yet. And it's always going to be an opinionated thing.
SPEAKER_02Yep.
SPEAKER_01Yeah where there will be some subset of serialization formats that are supported, because you know, we have to do other things besides new serialization formats like like you know, pick up clothes from the cleaners and uh deploy schema registry in Confluent Cloud also, which is a segue. Um this, by the way, uh my gosh, this has been a great overview of schema registry. Um I think very clear, and I I learned a couple things, which is uh a bonus. That's great. That's good. But nominally, we're here to talk about this being in the cloud. And I I think um uh like I said uh earlier, starting this off, uh on the one hand, uh that's kind of a one-sentence announcement, right? You could say, okay, well, scheme confluent schema registry is now available in Confluent Cloud. And if you know what schema registry is, you know what that means. But uh as it turns out, there's still a little bit more to the story. Um, like, you know, I can have a Confluent Cloud account and have multiple things called clusters. Now, because it's in the cloud, cluster is a little bit of a looser concept than you know, a cluster I deploy on hardware I can see. Um, you know, I don't know how many brokers are in it or anything like that. It's it's a cloud cluster, it's a logical namespace of topics. Yep. Um and there are brokers certainly in that are in space, you know, and I don't need to know about them because it's a service. And you know, those brokers get managed for me and scaled for me, and it's it's just a great time to be alive. Um but because in my conflict cloud account I can have multiple clusters, well, um, I think that that raises an obvious question. How many schema registries do I have? So do I have to create them? Just kind of walk us through that. What's what's the there's sort of this data model uh in cloud. Um just walk us through that.
SPEAKER_00Yeah. Uh that's an interesting question. Um so just to give you some context here, right? Like uh even in a traditional non-cloud deployment model, uh, the general recommendation we give out to users is like uh you must just use one schema registry irrespective of the number of Kafka clusters that you have, uh, because schemas are like uh organization-wide, or at least uh business unit-wide metadata. And you should treat that as a centralized thing more than a cluster-specific thing. Um obviously, you could sometimes tend to have like uh collisions with your topic names. Uh, so you could namespace your topic names and then avoid those collisions. But the general recommendation is like uh have one schema registry because it's uh organization-wide metadata, and you want to have one single view of your governance uh schema as a governance thing, in my opinion. Um so with that in mind, when we started off with uh Confluent Cloud thing, uh we were like we were discussing quite a few options and then like uh we eventually we thought like uh there's this logical concept called environments in Confluent Cloud. Um so environment is just a logical thing, it's a logical grouping of your clusters. Uh yeah, the an environment could be your dev environment, prod environment, or an environment could be uh business unit, or it could be just be a combination of both. So what we thought like uh was that uh uh for a given environment, you should just have one schema registry. Um so we kind of uh uh thought that was the right thing to do, uh, because uh that's the same recommendation that we give for our on-prem users as well. Uh so that's why uh when you go into a confluent cloud uh user interface today, uh, which is a web interface, uh you would typically see uh the ability to just create one schema registry if you just have one environment. But if you have more than one Kafka cluster and if you think like uh there are different environments like Dev and prod, um you can actually uh work currently work with our uh support team uh to get them separated as two different environments. And uh you'll get two different schema registries if you want to treat them as two different environments. Um so that's on a very high level, like uh how the uh schema registry is available uh across clusters uh in Confluent Cloud.
SPEAKER_01It's really one schema registry per environment. And like you said, environment is a logical grouping, probably dev staging, prod. I mean, those are those are that's always how we introduce the concept of environments, and most people use them that way. And um key underlying point there, like uh you said, is that uh schemas are a thing to be shared. You would never well, I don't want to say never, but we would strongly recommend uh against a schema registry per cluster. Um it's it's uh in you know if you're building uh uh an event-driven system, um, you don't really have APIs and service discovery in the way that you do with uh an RPC integrated system, you know, classical REST integrated microservices or the bad old days of SOAP or whatever. You know, we we have had very we have various service discovery mechanisms there. Well, uh the API now is uh the format of the stuff in your topic. So schema is API. And as a result of that, you want the smallest number of schema registries available as you can possibly get away with.
SPEAKER_02Yep.
SPEAKER_01Um that's our opinion. As and you know, we're I say us as Confluent, we're a group of people who have a certain amount of experience building systems like this and have developed opinions that we think are darn good opinions, and we recommend you have them too, which is why in Confluent Cloud it's one per environment. So you can have lots of clusters in an environment and they're all gonna be able to share schemas, and that's really, really powerful.
SPEAKER_00And if you think that absolutely doesn't work for you, and there are different business units where you want to separate uh the clusters and the schemas, you can treat environments as business units.
SPEAKER_01Of course, of course. Uh environment doesn't have to be dev prod. I mean, you you can make lots of environments and name them whatever you want. So you know, you can have business unit one dev, business unit two, dev, business unit one, prod, that sort of thing. So you know you can you can separate them out how you want, but the the it is a hard limitation of the data model that you get one per environment, and that's um that's kind of embodying a strong architectural recommendation of ours.
SPEAKER_00Yes, yes. But that's also another thing, right? Even even though like it's one schema registry per environment, uh there's no restriction in terms of who uses that schema registry in a sense. Uh if you have uh like let's say an on-prem Kafka cluster and produces in your on-prem ice, uh, they could still talk to the schema registry. Uh there's nothing preventing you from doing that.
SPEAKER_01Oh, sure. Yeah, they still have access to it. External. Yeah, that makes perfect sense. Um now we don't this is uh we we try to make this a podcast of interest to developers, and so you know, pricing of enterprise software and things like that is not really a thing that we talk about. Um but you were involved as uh lead developer on this stuff. You were involved in discussions of the kind of the economics of schema registry, and you know, internally in a company like us, there's going to be teams that are trying to figure out how do we price this and what do we charge for it? And that's kind of a science in itself that is really interesting on the business side. But where did all that end up? What does schema registry cost me if I'm using Confluent Cloud?
SPEAKER_00It costs you nothing.
SPEAKER_01Is it that great? What a great answer. So and I know, you know, ladies and gentlemen, the of course the sausage making process behind that answer is complicated. You know, you're rolling out a new uh product. It's it's actually hard to figure that out. How do you present something that's valuable and capture enough of the value that the business is sustainable and blah, blah, blah. Anyway, it's free. Schema registry is just there. Yeah. Uh and it's it's a feature of Confluent Cloud.
SPEAKER_00Uh, and we think that schema uh sh registry is an integral part of uh even streaming architecture. And uh like we want all of our users uh uh in Confluent Cloud to be using it. Uh that's one of the good practices. And schema registry itself is a very lightweight service. Uh so we we went through the economics, we did quite a bit of analysis, we figured out the best deployment model, and eventually, like we thought like uh we figured out a way to give it out for free.
SPEAKER_01Awesome. Yeah, and it it that's also that that reflects an architectural opinion that we hold also, which is that you really should be using this. And we would like we just want people's clusters in Confluent Cloud to be more valuable to them, and uh sort of subsidizing uh schema management is one way of making the whole thing more pleasant because it it you don't have to, I mean, you you still have a cost in that you have to learn how it works and you have to kind of learn the APIs and Avro AVSC files. If you've never done it before, you know, there are startup costs for you to use it, but there are no direct operational costs. We just want to make it so that people use this because it's the right way to use Kafka.
SPEAKER_00So absolutely cool.
SPEAKER_01Um what else? Uh uh without uh you know, obviously talking about any future product plans or anything, what else are you excited about just in the future of Kafka, future of things in the cloud? Um what's on your mind there?
SPEAKER_00Um I'm really looking forward to like supporting multiple data types, like I mentioned earlier. Like uh we want to be able to support JSON, PhotoBuff. That's one exciting thing that I'm looking forward for. Uh and at some time in the future, actually, like it would also be great if we can have some kind of uh enforcement of the schema on the broker side. Uh today, like uh all that we do today is uh like the producers and the consumers talk to the schema registry, and there's no real enforcement uh of the data that's available in the topic. Uh so we want to do some kind of uh broker side enforcement. Uh that's been requested by a lot of users actually. Um so I'm looking forward to that. And then obviously, like uh being in the Kinect team, I'm also looking forward to a lot of connectors being available in cloud. I'm actually working on these and like uh we will look forward to some news there.
SPEAKER_01There is gonna be a lot of connect news there. Maybe uh maybe we can have you back on the show to talk about that when uh when things show.
SPEAKER_00I'll be glad to do that.
SPEAKER_01Excellent. Well, my guest today has been uh Magesh Nandakumar. Magesh, thanks for being a part of Streaming Audio.
SPEAKER_00Thanks, Tim. I think it was a great pleasure talking to you and sharing uh insights about schema registry and uh letting our great developers know about uh what they think uh the right Kafka architecture to use for Kafka should be.
SPEAKER_01And there you have it. I hope that was helpful to you. If you've got questions, you can ask me at TLberglund on Twitter. That's T-L-B-E-R-G-L-U-N-D. Or you can leave a comment on any of our YouTube videos. Your question might be featured on the next episode of Streaming Audio. And feel free to subscribe to our YouTube channel and this podcast wherever fine podcasts are sold. And if you subscribe through iTunes, be sure to leave us a review there. That helps other people discover the podcast and just generally helps us get the word out. We appreciate your support. See you next time.