Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov
Hi, we’re Tim Berglund, Adi Polak, and Viktor Gamov and we’re excited to bring you the Confluent Developer podcast (formerly “Streaming Audio.”) Our hand-crafted weekly episodes feature in-depth interviews with our community of software developers (actual human beings - not AI) talking about some of the most interesting challenges they’ve faced in their careers. We aim to explore the conditions that gave rise to each person’s technical hurdles, as well as how their experiences transformed their understanding and approach to building systems.
Whether you’re a seasoned open source data streaming engineer, or just someone who’s interested in learning more about Apache Kafka®, Apache Flink® and real-time data, we hope you’ll appreciate the stories, the discussion, and our effort to bring you a high-quality show worth your time.
Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov
How to Run Kafka Streams on Kubernetes ft. Viktor Gamov
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
There’s something about YAML and the word “Docker” that doesn’t quite sit well with Viktor Gamov (Developer Advocate, Confluent). But Kafka Streams on Kubernetes is a phrase that does.
Kubernetes is an open source platform that allows teams to deploy, manage, and automate containerized services and workloads. Running Kafka Streams on Kubernetes simplifies operations and gets your environment allocated faster.
Viktor describes what that process looks like and how Jib helps build, test, and deploy Kafka Streams applications on Kubernetes for an improved DevOps experience. He also shares about some exciting projects he’s currently working on.
EPISODE LINKS
- Installing Apache Kafka® with Ansible ft. Viktor Gamov and Justin Manchester
- Containerized Apache Kafka on Kubernetes
- Kubernetes 101 | Confluent Operator (1/3)
- Installation | Confluent Operator (2/3)
- Confluent Operator vs. Open Source Helm Charts (3/3)
- Streams Must Flow: Developing Fault-Tolerant Stream Processing Applications with Kafka Streams and Kubernetes
- Kafka Tutorials
- Join the Confluent Community Slack
- Learn about Kafka at Confluent Developer
SEASON 2
Hosted by Tim Berglund, Adi Polak and Viktor Gamov
Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
Music by Coastal Kites
Artwork by Phil Vo
- 🎧 Subscribe to Confluent Developer wherever you listen to podcasts.
- ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
- 👍 If you enjoyed this, please leave us a rating.
- 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
Are containers a significant part of your life? Well, how about Kafka streams? If they are, then you might want to run your streams apps in Kubernetes. Victor Gammoff tells us why on today's episode of Streaming Audio, a podcast about Kafka, Confluent, and the cloud. Hello and welcome back to another episode of Streaming Audio. I am your host Tim Berglund, and I'm joined in the virtual studio by my friend and coworker and colleague in the developer relations team, Victor Gamov. Victor, welcome back to the show.
SPEAKER_00Thank you. It's always good to be back to the show. I'm always enjoying to be here talking about some exciting stuff. Hey.
SPEAKER_01You are my uh my Mike Mike Munger. Uh you're you're Mike Munger, and I'm Russ Roberts, I guess. You know, you're uh a frequent contributor, and that's a good thing. Thank you. Uh excited, Russ Roberts. Let's let's let's be clear about that. Anyway, um we're talking today. I'd like to talk to you today about uh some things that are very near and dear to your heart, and that's Kafka Streams and Kubernetes. And before we do that, uh you've been on the show before, but uh there's a lot of new listeners, and there could be people who this I know this is unthinkable to you, certainly, but there may be people who don't know who you are. So just uh please introduce yourself.
SPEAKER_00Yes. So thank you for having me again. Uh my name is Victor, and I work in part of the organization called Developer Experience. Uh, in many different organizations, it's called Developer Relations. Um, and I am a developer advocate, meaning that I'm not only going and preaching about stuff that we do, but also talking to people, absorbing their pain, uh providing some of the feedback uh internally trying to reduce this pain. So this is how I see uh my my work. Plus, I am uh still quite technical. It's not only about it's not only about uh the keynote slides, podcasts, and uh YouTube videos, which we have plenty, but also um still typey typey, touching keyboard.
SPEAKER_01Typey typey typey typey. As they say, not yet post-useful. Um so absolutely, no, you are uh uh very technical, and being a developer advocate is uh is not a place for those who are not technical. So that's certainly true. So anyway, you uh I said as I said before, you are uh quite passionate about Kubernetes and about Kafka streams. These are things that always come up uh in our work together when we're planning out new activities, kind of looking ahead to the next quarter, saying, okay, what are we gonna do? Those are things that always come up for you. So uh there's a good chance, just like there's a good chance a lot of people listening know who you are, there's a good chance a lot of people listening know what Kubernetes and Kafka streams are, but I don't want to assume that. And so I want to start by uh giving us a good top-down description uh what's Kubernetes and why does it exist?
SPEAKER_00Um it's a good question. So the Kubernetes is uh what what is what's known as uh container orchestration platform. Um your application or any application runs inside a container. Um container is some sort of that medium that allows to unify the process of uh packaging things and shipping things and whatnot. So, regardless what kind of language you use, uh once you put this in container, you can run this everywhere. And by everywhere I mean uh Kubernetes. Uh there are different platforms that doing similar things. However, Kubernetes is a clear winner these days, and um a lot of things that uh we do internally or externally with our customers somehow involved uh Kubernetes as a platform. Um sometimes I can feel that the Kubernetes is like almost everywhere where I go, but uh I just keep reminding myself maybe it's just a confirmation bias, and I just you know see what my eyes want to see.
SPEAKER_01But hey, that's uh what I call the red minivan effect. You know, you never see a red minivan until you get a red minivan, and then all of a sudden there's all these red minivans everywhere.
SPEAKER_00Yeah, very data analogy. Yeah, to simplify this idea, like I think about this as a Kubernetes modern distributed uh operating system that runs not in your computer but runs in your data center. And when you deploy your application, Kubernetes will decide where to run your workload based on these certain specifics that you want to run, like how much memory, how much processor, um uh what kind of network uh things need to be done, and Kubernetes will make it happen, right? So we we usually uh use the reference uh to um Jean-Luc Picard, and where he said, you know, make it so Kubernetes go in and doing this type of thing and making things happen. Um again, medium of running your stuff, like you this universal thing that Kubernetes runs, it runs uh containers, um, actually the group of containers uh called pods. Uh we can talk about this in in uh in a few seconds. And Kubernetes provides you the way how to allocate workloads. So you as a developer needs to declare what you need.
SPEAKER_01Good deal. You declare that in YAML, and uh we have I think uh on a recent episode, you and our own senior vice president of YAML uh were on the show together. We'll have to put a link to that in the show notes.
SPEAKER_00Yeah, but that was different uh uh yeah, different YAML uh type of thing. But yeah, YAML is kind of everywhere. Um if you want to talk about this, uh I can, but I would rather not.
SPEAKER_01So no, no, it's in fact I I'll uh I'll make sure we beep out me even saying the word YAML beep uh later on. Yeah, so uh nobody has to hear that kind of language on a podcast that you should be able to listen to with young children around. That really is our goal. So my apologies for uh the harsh language there. Yeah, yeah. Um yeah, so Kubernetes, you know, you you declare in uh configuration uh what sorts of things you want to be running, and those things are containers, and Kubernetes makes them happen. So it's yeah, as I see it, you know, containers sprung up as this lighter weight way of encapsulating workloads, sort of co-evolving with this trend towards small shared nothing programs. Yeah. Cloud, it's a cloudy thing, but uh it's you know, this this tendency to building applications in many small pieces that themselves are horizontally scalable and shared nothing. And and once you've got all these little stateless programs all over the place that are you know stateless with an asterisk, and the footnote says not really stateless, but sometimes sometimes people you know say about this that changes the way how you think about your infrastructure.
SPEAKER_00So this is why I like to use this analogy of distributed operating system. Um then you don't really when you have a general operating system, you don't really think how it operates with all the chips in your computer. You know that there's some software, you know it will be running somehow. Vaguely aware there are device drivers. Well, you know, I didn't get like formal education as a um like as an engineer. I'm a mathematician by education, so you know how I would know that what is going on there. So I I only see software. So for me, it's just like very uh the fuzzy feeling that I have inside. So I don't need to worry about this electronic elements. If something will break, I will just bring this to uh to the people who know, who have a formal education as uh uh engineers.
SPEAKER_01Exactly. So um, you know, as I was saying, Kubernetes co-evolving with this trend towards deploying applications in in small clustered shared nothing or you know, close to shared nothing pieces and deploying them in containers. Once we started putting stuff into Docker, we'll just say Docker. It's not like it's all that generalized. It's Docker, Docker1. Um, once you start making Docker images your unit of deploy.
SPEAKER_00So let's let's um let's be uh less specific. Uh let's call it um OCI. Uh OCI is an open container initiative that actually defines a the format of images and Docker uh is one of the implementations of this OCI kind of like specification. Okay, got it. So because the Docker itself it's it's a program product because there are different uh runtimes that are available and Kubernetes uh supports um like swapping different runtimes. So if you if you say running um like some vendor-specific uh implementation like OpenShift, OpenShift runs with uh different uh runtime, which is like 100% different from Docker. It's not Docker, it's uh I guess it's a cryo or something like that. Um so we we can be like uh less specific about the programs, uh program products rather than saying, hey, it's like uh okay, it's OCI compatible image format or whatnot. Got it.
SPEAKER_01Okay, so to be precise, we'll say OCI compatible. Although really uh it's pretty much always Docker. So once there was Docker and you were continually deploying these small applications of which now there were many uh in containers, you needed a way to wrangle those. Kubernetes emerged as the way to wrangle those, and now we have this nice. By nice, I mean YAML uh declarative format where we can describe what we want our runtime workload to look like, and Kubernetes is the thing that makes it so.
SPEAKER_00Yeah. Um for those of you who want to go a little bit deeper on this topic, um there is a very popular YouTube channel if you go to uh youtube.com slash confluent. Um and we will put some notes to uh the links in show notes, but uh you can find a few videos where I'm breaking down um the Kubernetes and some of the things around running stateful workloads and with example of Confluent platform. Um so check this out. I think it's pretty good.
SPEAKER_01Those those are in the show notes. Um I'm I'm saying that as a statement about a future state of affairs, but there's no way this episode goes live without goes live without those links being in the show notes. Don't pause right now, wait till we're done. Stick with us, and then by all means watch those if you if you need more of a primer. Uh they're they're excellent ways to get started.
SPEAKER_00Yeah, I think I just just encourage some of the people's um what's the um attention deficit disorder uh saying, hey, go check this out links and people like okay, should I go is real.
SPEAKER_01Yeah. Okay, so uh what I want to get to is Kafka streams on Kubernetes. So very briefly, uh for anybody who doesn't know what's Kafka Streams and why does it exist?
SPEAKER_00So Kafka Streams is a native library that comes with Apache Kafka. Uh this is a native library written in Java and uh uses uh by any type of JVM applications to write stream processing applications. It can be used as a like a baseline framework for writing stream processing applications, or it can be used um as a like a general runtime, for example. If you're familiar with KSQLDB, KSQLDB is largely dependent on Kafka streams as a runtime. Um, but there are different uh the frameworks that integrate with Kafka streams uh that allows you provide certain opinions about how the things need to be configured, how things need to be done. But essentially, Kafka Streams is very uh mature um library, even if you never touched the Kafka streams before, let um the word library not discourage you. Uh it is actually quite a powerful uh system that allows you to focus on on things that your application actually needs to do. Um there are very various APIs that allow you to implement um stateless stream processing applications where you just need to simply like a filter, do some transformations, or stateful uh stream processing applications where you need to say have um a result of individual computation step will be dependent on some of the data that was available previously. So in this case, you can have aggregation, running average, uh some of the windowing functions when you need to say calculate certain result over sliding window of like five minutes or or whatnot. So it is it is quite powerful.
SPEAKER_01It is. And talking about that library, and this is going to connect us with the the Kubernetes topic. So everybody stick with me for just a second. Kafka Streams has this very interesting opinion about how it is deployed. And um, this is in contrast to a stream processing framework like, I don't know, Spark Streaming or Flink that do this, where there is a cluster. You actually stand up a cluster of machines that have a process on them that's doing the compute, and they have a way of getting at that at the streaming data, usually ingesting it from Kafka. Um but that cluster of machines is where you deploy your stream processing program. You code against the framework's APIs, uh, you do your work and you deploy that code to that cluster and it runs there the way that cluster wants to run it. Of course, you have some application that's interested in the output of that, uh, which you then integrate. So Kafka Streams has this opinion where it's it's not like that at all. This is a library, you know, it's it's a Maven or Gradle dependency. You pull it in and you you code directly against this stream processing API, and that computation is done in the context of your application. So your microservice that's doing stream processing, in addition to whatever else, whatever other work it does, it's where the stream processing runs, it's where Kafka Streams is. Uh there is, and if this were a pure Kafka streams episode, we would dive deeply at this point into the scaling story for that. There's a really interesting scaling story. All I want to say about that for now is that when I've got a Kafka Streams application, I can scale that out horizontally. I can have as many, deploy as many instances of that application as I have partitions in my input topics. Um so that that thing scales. All of this should kind of start to sound scary if we want to run that service in containers and it's stateful and it scales out, and and this all sounds complex from a deployment standpoint.
SPEAKER_00No, actually not very complex. Okay, let me let me talk, yes, let me talk uh through some of the um thought process that usually ping if people are going through. So you um you you brought a very good point um about traditional by traditional I mean like people run clusters of certain um systems that will be running whatever workloads. They did this in Hadoop Spark, they're doing this right now. When we said library, so you know exactly where logic is, you know exactly um what you're scaling, what you deploying, and um over, I would say less the 10 years, no, maybe not the 10, but like let's say five. Um people develop very good practices about deploying Java applications. So they're to the point of uh running this microservice, it's actually you're running Java application in container. Java application has library that is uh Kafka Streams is a dependency, and that's it. So this is what you do. Um, how you would do this? So, first thing is that you need to build a container. So, a couple things that simplify this uh this job. Um usually as a developer, you want to have a tool that will be um you know somehow connected to your workflow. So if you're using Maven and Gradle, it would be cool if you will be staying in the same workflow for producing this artifact. Say you're producing Java artifact, jar, right? Or not or we're just producing jars these days, not war.
SPEAKER_01Um no, we don't we do not make war.
SPEAKER_00Yes, and anymore. And uh yeah, and and uh and it would be cool if we can use the same tools uh to produce this. So for that matter, I guess like uh like two years ago, Google released uh uh a very cool project called Jib, J I B, James Iris, Boris. And this project is um a set of tools that allow you to build uh OCI um compatible image in um in a pinated way. So you don't need to specify your base image. You can if you need to, if you if you know what you want. If you're not, um Jib will figure out okay, looks like it's uh I don't know, Java application or looks like Spring Boot application or stuff like that. It will provide uh some of the more or less standard uh image that allows you to execute these kind of things. Um the couple couple interesting aspects in in regards to Kafka streams. Previously, uh over over the time when we were um going into this like container world. I I thought you were gonna say previously on streaming audio. Previously, on streaming audio. We talk about um we talk about infrastructure, but we will get there to to you know the Kafka and stuff. Um in the previous years, people kind of get um um confused, maybe, or they were trying to find this kind of image that will uh will be small enough so it will have a low overhead, but also um include all the things that their application needs. For example, in the past, I've seen the people using like Alpine image, but Alpine image um doesn't have certain dependencies. Why it is matter? It is matter because um Kafka streams depend on certain libraries specifically. Um some of the listeners know, uh some of them don't. Team definitely knows that uh Kafka streams uh default implementation of the state stores, this small piece of functionality that allows you to uh perform this kind of stateful uh computation is dependent on uh embeddable and quite powerful database called RocksDB. Rocksdb is written in C and uh has uh bindings for different languages. So we have a Java binding and I'm using this Java API. So because it was built in C, it requires some of the um runtime libraries available. So if like you were using Alpine, there was no like libc uh6 or whatnot, so it will require to install it separately. So nothing of this is happening anymore. So we do have a good tool that allows you to produce um pretty pretty small image, I would say. Like it's a it's it's a good balance between size and uh functionalities, and now Kafka streams are running this like successfully. If you want to see some of the examples, how it can be done, there's a very good uh website. Uh, probably um my second favorite uh website and internet um after streaming audio uh the website, of course, thank you, is uh the Kafka tutorials website. Uh if you go to cnfl.io slash tutorials, um you can find different examples. And uh if you look into Kafka streams examples, you will find the way how we can configure this uh image. And we're using Jib there all over the place. We generate these Kafka streams images fairly easy.
SPEAKER_01Could you I want to be a little more concrete about Jib. Jib is so the the builds in Kafka streams are Gradle builds.
SPEAKER_00Yeah.
SPEAKER_01And uh sorry, Kafka streams in um Kafka tutorials. We just mentioned this to this website of ours, which has these executable tested tutorials about how to do various things with Kafka and KSQL and Kafka streams. And so if you go to one of the Kafka Streams tutorials at that URL, of course, we'll link to it in the show notes, you'll see there's there's always three steps there's build, test, and deploy. Um build is really you know how to write the code and describing every part of the API and what we're doing. Test is how to write unit tests for the same, and then deploy is really how to make a Docker image, right? And that's using Jib. Yeah. And all these things are built with Gradles. It's a Gradle build for each one of these steps, and there is Jib takes the form specifically of a Gradle plugin and a Gradle task that takes our code and uh builds it and puts that jar into an image that has some sensible base image with an operating system and a JVM and all the things you'd need to run it. Is that is that am I yes?
SPEAKER_00That yes, yes, no, no, no.
SPEAKER_01You summarizing very well, thank you. My job here. My job here to start and end and then summarize. That's really folks. If you're a regular listener, you already know that is in fact what I did.
SPEAKER_00Best job ever. So they so once we in this like developer process, uh, okay. So we need to figure out a couple things here, right? Okay. Um, first of all, you build this somewhere. Usually your application uh is uh code of your application goes on the version control. Um there's rumor has that the team has his f favorite version control system as Git. So I mean at this point, yeah. Yeah. So the uh like we did you expect I would say, I don't know, subversion or or or what's the uh I thought you were gonna say perforce. Hg.
SPEAKER_01Yeah, perforce is actually not bad. I would like us to make sure we edit that comment out.
SPEAKER_00Anyway, go please go on. Yes. So the you comment your code, uh there's a process that will take this one. Build this image. Now this image needs to be placed somewhere. Usually container registry that will be also available from your um your Kubernetes cluster, right? So like if you build image locally, it doesn't make uh things good because your Kubernetes cluster would not know how to bring this image. And remember, um, you telling this Kubernetes that you want to deploy this image into Kubernetes cluster and Kubernetes will try to make it so. What does it mean? It will try to pull this image from container registry and try to run this image locally inside this Kubernetes cluster. So that's why you need to use things like Artifactory or Docker Hub to um if you're using Docker Hub if you're deploying some public uh stuff, um uh Artifactory maybe in your organization. Um good thing is that majority if you deploy your application in in cloud, for example, I use uh Kubernetes in GCP because it gives me um very nice experience in terms of developer. I don't need to you know think much about this, I just spin up these clusters and I deploy stuff. Um the Google Cloud, um Amazon, they also have um their implementation of container registry. So like do not forget that like things when you build it stuff, it's not stopped yet there. Now, once you do that, you need to define, as we already established here, right, um YAML configuration. And there's actually not so much to talk about this uh as as always uh devil in details, but let's let me uh okay walk you through. Now, there's uh few things that you can use from perspective of YAML. Now you have your image, right? You have your image in your uh um in your container registry, which now you let's say you've built that image with Jib and you've deployed to a container registry, which could be Docker Hub, if it's something that's public, or could be like JFR JFrog Artifactory does this, all kinds of things, right? Yeah, exactly. And uh the thing is that like JIP also understands different uh registry. You can specify like where you want to push it, and the image will be uh deployed, um deployed as a part of the build. Now, and only thing that you like only thing is left uh is defining your deployment. Deployment is the standard um standard resource in Kubernetes where you define um specification for your for your application that will be deployed, and that's it, you're done. So it's it's pretty easy. And uh using uh the same same tools and the same nature of if you want to scale your application, you should go say you know, scale my deployment by two, and it will start multiple multiple instances of application. Now, let's talk about details. Yeah, so the as always like devil and details. So essentially, you know, you you once you define this deployment, you you deploy this and you're done. So done. However, there's a couple of things that you need to be um aware about. So, first of all, the Kafka streams, good, even though Kafka streams depends on this like small uh database called ROXDB, it is actually backed by Bigger Brother, which in this case a Kafka, because all things that happen in your instance of your application by default will be uh replicated through uh through Kafka. And maybe you learned from the previous episodes. We also put the link to the previous episodes where we break down some of the challenges of uh stateful platforms. Um, you don't need to worry about this type of thing because you already got covered by by Kafka, right? Um and all the state stores uh they will be replicated through Kafka.
SPEAKER_01Let me let me summarize that before we go farther. I want to make sure everybody's got that. So in inside Kafka streams, uh sometimes explicitly and sometimes implicitly, uh a stream processing job or a what we call a stream processing topology needs to accumulate state. Now, sometimes what you're doing is maybe you're you're processing some input stream or some topics. Yeah, and you want to count how many uh times Victor's Victor has mentioned in uh Twitter feed based with some hashtag in it. So you're you know ingesting tweets and you just want to count. Uh well that's that's state. You have to remember the previous count. So you have to put that state somewhere. Uh sometimes you're you're you know building large lookup tables, uh, and that's state. So you you have you have state in your stream processing job. That's that's stuff that's in memory, and that stuff in memory is kept in RocksDB. And so RocksDB is uh, like Victor said, an embeddable log structured merge tree key value store. It's highly worth reading about and learning about. Super interesting little piece of technology. And that is it's a database, so it's got stuff in memory, it's got an image on disk, and that in-memory key value store is also flushed to an internal Kafka topic. So the the um persistence of last resort is Kafka itself for all of the state in a stream processing job. So a stream's application has you know, think of it as this big giant hash table occasionally uh that that it's using to do things, and the contents of that hash table are stored in Kafka and stored locally on disk. Right? Tell me if I'm wrong.
SPEAKER_00No, stay to this thought. Okay, you mentioned um stay stay to the disk, and some listeners might say, hey, but like what about um like if uh something will happen, you know, because in the world of Kubernetes uh things are changing, some of the um uh pods, which is group of uh containers, will be rescheduled to the different uh to different uh nodes, different physical nodes, what's gonna happen with this like file, right? So first thing uh as I already mentioned, this thing uh will be replicated through Kafka. So at this point, you don't need to worry about about this much. Another thing here is that if you still worry about uh this thing much, uh the Kubernetes provides you the ways how you can define persistent volume claims and you can have a persistent volume. Here's the one problem. Now, how you can tell the Kubernetes that this uh this persistent volume will belong to your container and not another. Um, again, in the previous episodes of uh streaming audio, we broke down uh the concept of stateful set, which is um a special type of resource that extends deployment, so you have the same features that you have in uh deployment, so you can scale it, uh scale up, scale down, very easy. But also you'll get two things it's a stateful network identity and stateful disk identity. So you will get um understandable and a human readable uh name for your container for your pod, and you will have a stateful disk identity. So if your pod with instance of your Kafka streams application will fail, it will be restarted and a state store uh this file uh the file systems will be reattached to the same pod.
SPEAKER_01Got it. Okay, and that is exactly where I was going because I I know Streams puts all this stuff on disk, and that is how we make sure that uh when a container that's running an instance of a stream processing app goes down and comes back up, then it it is attached to the correct storage. Yes, that's correct.
SPEAKER_00So and I I configure that in YAML. Um so yeah, so you essentially with uh Kafka Stream application, there is an environment variable or like there's some parameter that you can pass through environment variable that can specify which folder Kafka Streams will be using for all this state store information, and after that you define your PC system volume, attach this P system volume claim to uh your stateful set, and you go to go here. Now, this is also pretty we're pretty much done here. Now we are. Yeah, we are. Um the next thing is that let's talk about um uh scalability aspect, right? So the way how the the way how we we touched a bit a bit about like scalability, which is go there and say, hey, uh start another instance of your application. And the good thing about this, a good thing about Kafka streams in contrast from the different systems, the different stream processing systems, is that um Kafka streams, instances of application, they don't need to talk to each other directly, they never talk to each other directly. They talk in through uh Kafka uh through consumer group coordinator and all information about topology of application. So we started another instance, now we need to uh find a way how we can split the load. We need to find a way how this load will be distributed across this uh newly created instance in the previous version, uh previous instance. So all these things are happening through Kafka. Again, uh using tools like either Conflict Cloud or using uh like a conflict operator, we already figure out the ways how we can store data in Kafka. Kafka already deployed, we're already using it, so we don't need to worry about this. And now uh through Kafka, for coordination through Kafka, uh multiple instances of Kafka Streams application um you know start creating this small processing cluster. By small I mean just like because we can scale this up and scale this down as we need, depends on the law that is growing. All right.
SPEAKER_01Um and what in your opinion, what is the developer experience like there? This this maybe I should say the DevOps experience. But you know, you do qualitatively. We joke about YAML and like nobody frankly likes YAML. No, nobody left. It's it's it they don't, they don't. And it's and that's one of the kind of dark sides of Kubernetes for all the glory of it as the now standard container orchestration uh system. You know, you end up essentially developing software in this text file format that that is prickly with respect to indentation. And you know, there's all kinds of bad things about it, right? So, you know, we'll we'll we'll we'll keep making jokes about YAML, just like in 2006 we made jokes about XML when everyone realized what a terrible idea that was. Uh, but it's not it's not so much that. I don't want to focus on that, but uh qualitatively, when you look at how to do all this, are what is the developer experience like? It feels like you know, I have to mess with stateful sets and I kind of have to get all these things right and I have to have a lot of knowledge about how Kubernetes works. I don't just get to say, here's my jar, or um, you know, here's my here's my image at least. That that seems like that's not too much to ask. And go do it the right way. So how do you feel about where the developer experience is right now? And what do you think if we just kind of put on your your futurist hat, where do you think the platform needs to go to make the streams plus Kubernetes thing be better?
SPEAKER_00Yeah, I think we the things from uh perspective of day one type of things, you know, you need to deploy uh and up and keep this up and running, it's uh more or less solved. Again, the framework itself provides certain capabilities where you don't need to you know overthink much and how the things you know, how to make sure you're not losing your data and whatnot. However, there are many aspects that we didn't talk today, and usually those aspects are overlooked uh because of the I don't know why. Um, for example, monitoring, how you would perform monitoring. Um so essentially, you know, the Kafka streams provides uh the variety of metrics that can be exposed through GMX. One of the patterns that we recommend people to use is to have a sidecar container, and this sidecar container will be collecting these GMX metrics and we publish it to usually it's a Prometheus because it's the standard tool. Uh Control Center has integration um with uh with uh Kafka Shim's application once the monitoring interceptors will be um will be installed. So all these things as a developer you need to think about this, right? You need to think how I will be monitoring. It would be, you know, in my opinion, it would be great if the the things that we discussed on the previous episodes of streaming audio, like uh like operators, there would be operator for streams that will you just go and check marks saying, hey, yes, I want to state full. Um, I want to have monitoring enabled, like metrics enabled, and I want to have like a five instances of this at the end.
SPEAKER_01And yeah, five instances and go.
SPEAKER_00Yeah, so that would be pretty cool. Another thing that um I'm just start researching. I'm still kind of developer, and I was thinking or researching some of the platforms that allows you to kind of sort of go to this extra mile from point where you already have your CI CD pipeline, you're have your um developer environment set up, and like how this last mile from your CI CD pipeline to production. Um, and I think the tools like Hiroku, for example, give very good user experience in terms of like how the stuff will be running. Because you never know that you never knew that the actually your application when you do Hiroku push, this actually will be creating container and will be runs inside the container inside inside Hiroku. They did this way before it was in it was cool. Um, and because they didn't advertise it and just saying, hey, so this is your application in Git, um, just do push and we will be running this for you. So I think this is where we need to um look for inspiration. Or Cloud Foundry. Cloud Foundry has like a similar similar idea, and like bear with me, I will, you know, I will join these topics together at some point. Um very soon, not at some point, it's very soon. So so the Cloud Foundry provides the ways as a developer, you're also writing your code, you're doing CF push. And after that, Cloud Foundry also provides the way how your application will be built and how it will be deployed. Um, the way how the Cloud Foundry was doing this for a while, it's a fruit concept called bill packs. So essentially, build pack is um it's like a it's like a you go into your system administrator saying, Hey, this is my Java, and your system administrator opens his like a big book and looking, okay. Um I'm looking for index, okay, jar, okay, jar. I'm looking to the jar and I opened the page 55, and I know there's a set of instructions how I can deploy jar, how I like, okay, uh let me introspect this. Does it have some Spring Boot things inside this jar? Oh, yeah, okay, I'm going to the page 55, and now it's a jar that has a Spring Boot, and I know how to run um Spring Boot. So build back this is kind of this instruction. And um, I think right now there's an initiative to make this uh kind of like a cloud native uh standard, so that it would be the part of developer experience. Uh, when you uh build in your application, you can specify in order to speed up this thing that there's a build pack that you use and say it's a Java application, so it will use this build pack, um, provide the certain things, it builds image for you, pushes this image to whatever registry, and pushes this into um into running Kubernetes cluster. There are some tools that kind of aiming this kind of uh I don't want to be like very buzzwordy here, like a serverless um type of approach when you're running your Java application in the in the manner that you don't need to worry about you know CI CD thing. So I think this is something where like futuristically, hopefully, that would be not super futuristic, it's just a couple years of work, but essentially you can specify your application and the tool will be automatically build the image, push this image, and deploy this image in production and gives you all observability uh points, uh gives you access to logs, give you access to health checks, gives you access to maybe even um some extra knowledge because it's not just simple Java application, it's Kafka Streams application, so it also will have some contextual knowledge about like what is going on there, what's the stateful stores. Um I've seen I've seen like a few frameworks around uh that kind of trying to do similar things. Um we will have some of these uh things uh discussed in uh Kafka Summit London that come in in a couple um in a couple months. Um so we'll see how it goes. And uh we will probably um we can reconnect after and see over over a certain time and see if some of the predictions actually came to life or not. Does it make sense at all what I just said? Um it does.
SPEAKER_01It does. It actually sounds like a pretty uh a pretty nice future where I've got you know essentially a jar, which is the output of my that's what my Kafka streams app should be. Um and I hand that off and it gets containerized, and uh there's a controller that understands how to uh how to deploy it and properly attach storage, and like you said, bring in monitoring and all those kind of day two things. That sounds like a fantastic future.
SPEAKER_00Yeah, think about this. Like we have a Kafka um Kafka tutorials that uh we have this section, bring this to production, and it actually brings this to production. It just like deploys it and it feels like okay, what just happened? Like I have a give it your AWS credentials and it's it's there. Yeah, you're getting this like a conflict cloud uh the connection string to get your Kafka thing. You don't need to have an operator in your Kubernetes cluster if you don't want to, but you you always need to have a Kafka because Kafka Streams doesn't exist without Kafka. And it's totally fine. You can deploy your uh Kafka Streams application in Kubernetes cluster and still use uh the conflict cloud. Um, I'm again I'm not uh like trying to like oversold this idea, but hey, if you don't not in the business of writing Kafka, please don't.
SPEAKER_01So my guest today has been Victor Gamov. Victor, thanks for being a part of Streaming Audio.
SPEAKER_00Thank you. Thank you so much for listening to this episode, and uh as always, have a nice day.
SPEAKER_01And there you have it. I hope this podcast was helpful to you. If you want to discuss it or ask a question, you can always reach out to me at TL Burgland on Twitter. That's at T L B-E-R-G-L-U-N-D. Or you can leave a comment on a YouTube video or reach out in Community Slack. There's a Slack signup link in the show notes if you want to register there. And while you're at it, please subscribe to our YouTube channel and to this podcast wherever fine podcasts are sold. And if you subscribe through iTunes, be sure to leave us a review there. That helps other people discover the podcast, which we think is a good thing. So, thanks for your support, and we'll see you next time.