Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov
Hi, we’re Tim Berglund, Adi Polak, and Viktor Gamov and we’re excited to bring you the Confluent Developer podcast (formerly “Streaming Audio.”) Our hand-crafted weekly episodes feature in-depth interviews with our community of software developers (actual human beings - not AI) talking about some of the most interesting challenges they’ve faced in their careers. We aim to explore the conditions that gave rise to each person’s technical hurdles, as well as how their experiences transformed their understanding and approach to building systems.
Whether you’re a seasoned open source data streaming engineer, or just someone who’s interested in learning more about Apache Kafka®, Apache Flink® and real-time data, we hope you’ll appreciate the stories, the discussion, and our effort to bring you a high-quality show worth your time.
Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov
Installing Apache Kafka with Ansible ft. Viktor Gamov and Justin Manchester
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
“It’s one thing to get a distributed system up and running. It’s another thing to get a distributed system up and running well.” Ansible keeps your Apache Kafka® deployment, management, and installation consistent, and it enables you to implement best practices that make it easy to get started. Justin Manchester (Platform DevOps Engineer, Confluent) and Viktor Gamov (Developer Advocate, Confluent) discuss the problems that Ansible is trying to solve, enabling collaboration and optimizing all components for top performance.
EPISODE LINKS
- Learn more about Ansible
- Follow Viktor Gamov on Twitter
- Follow Justin Manchester on Twitter
- The Easiest Way to Install Apache Kafka and Confluent Platform – Using Ansible
- Join the Confluent Community Slack
- Fully managed Apache Kafka as a service! Try free.
SEASON 2
Hosted by Tim Berglund, Adi Polak and Viktor Gamov
Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
Music by Coastal Kites
Artwork by Phil Vo
- 🎧 Subscribe to Confluent Developer wherever you listen to podcasts.
- ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
- 👍 If you enjoyed this, please leave us a rating.
- 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
Ansible is today's leading deployment automation tool, turning everyday ordinary YAML files into fleets of running servers configured perfectly. At least, that's our goal. All of those files are organized into groups called Ansible Playbooks, which can represent years and years of collective deployment expertise in a few humble, white space significant files. Confluent platform has had Ansible Playbooks for a while, but as of version 5.3, Ansible Playbooks are production ready and fully supported. They cover all the components of the platform from Kafka itself to connect to Confluent Schema Registry and more, and include all of the security features like Kerberos that we know can sometimes be a little prickly to configure. I'm joined today by two guests, developer advocate Victor Gamov, and platform DevOps engineer and senior vice president of YAML, Justin Manchester, to walk us through all of this. It's all on today's episode of Streaming Audio, a podcast about Kafka, Confluent, and the cloud. And I'm very pleased to have with me two guests today. One returning guest, Victor Gamoff. Vic, say hello.
SPEAKER_00Hi, hi, Tim. Hi, our listeners.
SPEAKER_02Okay, go. And uh a first-time guest, Justin Manchester. Justin, welcome.
SPEAKER_01Hey Tim, thanks for having me. Uh longtime listener, first-time participant, happy to be here.
SPEAKER_02Awesome. We're happy to have you. So uh our subject today is Ansible. And I would like both of you to give just a brief introduction. Uh, we've got names, but uh what do you do? You're we're we're you're both co-workers of mine at Confluent, but what do you do here that uh would cause you to be here to talk about Ansible? Justin, what what would you say you do here?
SPEAKER_01Well, officially my title is Platform DevOps Engineer. Um and essentially what I'm responsible for is I'm the engineering lead on a project called CP Ansible, which is Confluence um set of playbooks to deploy and automate uh your deployment of Confluent platform. And I'm sort of responsible for the technical side of that.
SPEAKER_00Also, Justin known as uh senior vice president of YAML uh at Confluent. Because of that, because of things what he does here. I don't know if you guys have seen this, but it was actual uh like a screenshot of some of the like a job postings, like we're looking for like senior VP of YAMLs or something like that. So I think I think we're right, we're in the right uh audience here.
SPEAKER_02He's in the running, you know, we're interviewing for an executive global executive vice president of YAML. He's in the running for that. So uh Justin, we're pulling for you.
SPEAKER_00As you might understand, based on the my mood here, that I will be the person who will be throwing some of the YAML jokes because I know as a you know uh air quotes a fan of YAML here. So let's try to run with this somehow. And team, you had to manage and navigate this discussion because we have a very uh uh technical, accurate person and some kind of like a joke cracking guy. So let's let's let's do something with this, right?
SPEAKER_02Exactly, exactly. Well, this is helpful. I don't have to be as funny today because you're here. But uh Vic, what we don't I know you you've been on the show several times, even recently, but um what in case anybody's downloading for the first time, tell us what your role is here at Confluent.
SPEAKER_00Yes, uh for the first time, yes. My name is Victor. I work here at uh Confluent as a developer advocate. So um I talk a lot to people. Also, I listen a lot of people. So uh this is part of my job where I'm talking about some technologies, and I'm happy to be in this um world of uh DevOps, and we'll try to navigate some of the technical discussion around how Kafka and streaming platform feed in a like a technical aspect of DevOps and how organizations need to think about this. So that's why all sorts of configuration management and installation and provisioning and uh maintenance and day two responsibilities of uh or day two uh requirements of streaming platforms is also falls into my you know the field of interest and field of expertise, if I may say that. So um, so this is why we have this conversation here, because for me it's also the part that um I'm interested in talking and interesting to getting feedback from from our community as well.
SPEAKER_02Awesome. And Victor is also, I should point out, a member of my direct team. So I will be dealing with him after we record this podcast. All right. Um I want to start with uh, you know, I guess the obvious question, and again, we always try to take an approach of assuming that you're new to the things that we're talking about and you want to learn about them. The obvious question is what is Ansible? Um I want to back up a step behind that, and I want to ask what is the problem that Ansible is trying to solve? Um Victor and Justin, what what is that problem?
SPEAKER_00Um if I may quickly start and after that I will like uh uh pass the microphone to Justin. Uh because I, as always, I have opinions about everything. And uh one of the things that I like to talk about is when I um just talking about all sorts of conflict platforms, specifically right now on the cloud or or or or Kubernetes, I like to have this uh kind of introspection of evolution of uh Kafka DevOps or Dev Kafka ops uh personality or culture inside the organization. So the way how the people usually start with with uh um with Kafka and uh Justin, you can also tune in because of your past as a customer success uh engineer. Um the people usually just don't load either Confluent platform because it has all tools, or they start with the Apache Kafka open source, and they start with the shell scripts, like start Zookeeper, uh change the configuration, and after that they will start Kafka and things like that. And when they're running this, it's usually some person who's just like playing around. Because once you start bringing these things into the more serious environments where you need to do production installation, you usually don't have just one environment where you run your Kafka, right? Even though if it's just the only one um use case of your of Kafka in your organization, you still might be running at least two or three different environments, one for development, one for for QA or system integration tests, and one for production. And the task that you're trying to solve is to have a consistency between these environments. So when you're developing that, you developing close to the things that you have in production, and when you have some stuff in production that uh your development also some sort of somehow like a mocks it, so you would know that you know you're developing real stuff. So this is why it simple having the simple scripts is or some sort of like homegrown automation system may be okay for the first couple months of the project, as the project is growing, especially when you're trying to be you know agnostic from different platforms that you're running. Maybe you're running VMs in development, maybe you're running bare metal in uh in production, you're trying to find the ways how we have this consistency. This is what what uh like uh configuration management tools uh gives us.
SPEAKER_01I think I think I think your point about consistency is a great one, and I think that's really, really important. Um I think that's definitely one of the one of the major aspects that that Ansible and CP Ansible give us is that consistency. But I think it actually even goes further than that. Um it's not enough just to be consistent. Um I've been working with distributed systems for a very long time now. Um, and if anyone else listening has, they'll appreciate that it's one thing to get a distributed system up and running. It's another thing to get a distributed system up and running well to be performant.
SPEAKER_00And also keep it.
SPEAKER_01Keep the key to running and keep it performant. So we use Ansible not just as a way of keeping things consistent, but of also implementing um some of our best practices um both at an OS level as well as a CAFTA level to make things as performant as we can across most use cases. Obviously, there are always edge cases which are hard to capture and to control for. But it's it's also a case of keeping things uh not just consistent but performant. Um and with the right setup, auditable as well. Um as far as being able to understand who has changed what and when, and using it within conjunction of tools, say like Git, for example.
SPEAKER_00When you mentioned the best practices, quickly add this thing, uh, because I I really want to emphasize that when you mentioned this best practices thing, it's also have this um kind of call out to our some of the previous discussion. If you've been when we're talking about some Kubernetes and operator patterns where we have experience, and specifically like uh when the gentleman mentioned, we've been running distributed systems like for money, and uh we know how to do this, and this is where tool that meets um experience that will create this best practice, right? So it's not uh only tool that you know that allows to do stuff, but also have experience of running this these things. Uh is actually what is you know beneficial, that what creates this like essential like playbooks that we're gonna be talking about that will be you know better or worse than other things.
SPEAKER_02So to sum up, there's uh consistency among multiple environments. I'm I'm gonna wrap that up under the term automation, right? You've got you're using a computer to do this, and so you can do it in a consistent way without having to do a bunch of manual steps uh to you know to make real this configuration uh on some computers somewhere and get it to work the same way every time. So it's it's a it's a uh problem of consistency, a problem of best practices, assuming that the system you're setting up is complex and you might not know everything about it. Um we're we're trying to get trying to come up with a way of gathering all of the best expertise about how to configure that system and put it in one place. And Justin, you just threw in uh let's let's make this as collaboration friendly and as version controllable as code.
SPEAKER_01Correct.
SPEAKER_02So make it a lot of things.
SPEAKER_01So it's the whole it's the whole um infrastructure as code mindset, right?
SPEAKER_02Yeah, that's what we call that. Infrastructure as code. So used to be infrastructure was I edit a bunch of files in VI on a bunch of hosts that I know how to telnet to in the bad old days, you know, uh or SSH to in any time in the last 20 years. Um and now we turn that into a set of source files and a tool that takes those source files and goes and puts them out into the world. Exactly.
SPEAKER_00And uh the things with these um with these you know controllable version tools, uh Justin also mentioned audit. Sometimes it is important for for external people who perform like security audit or any type of like a penetration test and seeing like what the you know breaches that uh your system might have, having this like a single point where you know no one else has access to this machine except this like a CI server that performs this like deployment through SSH, and there is only one um you know, one entrance to this world. So people don't have access to these computers and things like that. So this is also um the part that um tools like uh the configuration management uh tools that that provide.
SPEAKER_01Something we're actually working on directly relating to this. It's on the on the backbone or a little sneak peek, maybe for some of the listeners, is I'm actually working on a guide right now for how to work with CP Ansible um in a structured, audible, auditable way. Um so the example I'm working with is Git, but you could be using any sort of version control system. Um so it's something we're working on right now to give them guidelines as far as you know, here's okay, great, you've downloaded this thing. Now how do you actually use it sort of effectively in a in a production environment?
SPEAKER_02Very nice, very nice.
SPEAKER_01Um so that's that's in the pipeline and coming. Um so we're not just looking at this from a technical perspective. We're trying to look at this from a usability perspective, really, like what's going to make the end user successful.
SPEAKER_02Cool. I like how you say that. You know, we're not saying you have to use Git. You can use any version control system you want. But you know, just for sake of example, we're using Git. Um, and I suppose there are some organizations that still use things that aren't Git, but it's uh it's a safe choice, I think. I think that that resource is going to be broadly applicable.
SPEAKER_00As long as you're using any version control system, you're good. Right. You know, as as as long it is something that allows you to have a centralized versioning and have, again, see for consistency of for these configurations, you're good. Yep.
SPEAKER_02Now, um I we are in some ways getting ahead of ourselves. So uh those are the problems Ansible's trying to solve. Ansible is obviously a solution to those things, but can you guys talk about? Well, you know, if you are answering the question, what is Ansible? I want to hear what we've missed so far. But really, since I think we all sort of understand the basic deployment automation infrastructure as code kind of problem, what sets Ansible apart? What are the other options that Ansible is not, and how does it differentiate itself from those?
SPEAKER_01Yeah, so the thing that I really like, and I know a lot of other people who work with Ansible really like about it, is that it's agentless, as Vic was mentioning previously. Um, we had a previous discussion, we were talking about this a bit. Um it's agentless, it just uses SSH. Um so essentially it uses SSH and keys and just goes into the remote box, um, deploys a bit of software that changes whatever configurations you want to have changed or installs whatever software you want to install, and then sort of cleans up after itself and disconnects. So there's no there's no post management, there's no infrastructure to manage with Ansible itself, um, unlike say Puppet or Chef, where there are agents and there's this whole system around that that you need to manage. So it's much cleaner, it's much simpler. Um yeah, that's that's that's the primary reason why I really like it as opposed to other automation solutions. Um the other aspect to it, which I know Victor's gonna give me a hard time about here, is um is is YAML. I am gonna say that. I know a lot of people treat it like a dirty word. Um the beauty of YAML is if you're new to automation, if you're maybe maybe you've got a bit of a development background or you're more of a sysadmin where maybe you haven't done too much development. Um pretty much anyone can read a YAML file and generally understand what's going on. Um it's generally pretty clean. It's in plain English, it's easy to understand. Now, before before Victor jumps all over me, which I I I will I will I will state that there are quirks to YAML around spacing and carriage return and things like that. Yeah. But for the most part, it's very easy to read if you're first sort of trying to figure things out.
SPEAKER_00It is it is uh perhaps it's okay. No, uh we still can be friends, um um even though you're enjoying doing this with with YAML. Uh what I'm trying to say here is that which brings us to a very important point, that um using YAML uh creates this interesting mentality. So you actually, when you're creating your playbook and you you're working in the different like beats of this framework, you actually um defining also state of the world, right? So it's a declarative uh language, so it's not procedural uh language where you're saying how to do it, but you rather say uh what uh you want uh this particular host, what kind of software it needs to have, and Ansible will try to uh uh execute that to make this into you know a certain state. Now and which which is like a very good point on um you know defining defining like infrastructure.
SPEAKER_01Yes.
SPEAKER_02Um and that that it is declarative, I think, is important. Um and Justin, you also mentioned agents. So just real quickly, um Chef and Puppet are the the the two other tools that mu uh Ansible is most often compared to.
SPEAKER_00Holy Holy Trinity of DevOps is uh Ansible, Chef, and Puppet.
SPEAKER_02Right. Now um Chef and Puppet are both they both rely on the target systems having agents installed, is that correct?
SPEAKER_01That's correct, yeah.
SPEAKER_02Um how does Ansible work without an agent? Can you it'd be interesting to talk about that a little bit?
SPEAKER_01I I I wasn't sure. Yeah, if you want to get into that level of detail, we can. So basically Ansible is all written in Python. Um so the underlying modules, while you're doing everything in YAML, all the underlying modules are written in Python. And so what Ansible does when it first connects to a host is it gathers facts on that host, and that's details, anything you can think of, you know, CPU spec, home out of memory, disk space. But it also looks for the version of Python installed, and either 2.x or now 3.x with 2.x going away as of January. Um and it uses that to ship its code over and to compile against the version of Python that's installed to then execute whatever tasks you've told it to execute. And then once it's executed those tasks, it then removes all those components and disconnects. That's that's basically how it works in a nutshell. Um the beauty of it is almost all major Linux releases these days have some form of Python installed. Um so pretty much out of the box, Ansible can work against almost any host.
SPEAKER_02And so it needs to communicate SSH essentially. Yep. Yeah, it's an SSH connection. And so you have a root password or or uh key key pair or something like that.
SPEAKER_00So typically password that allows to perform. We don't share the phone.
SPEAKER_01Typically, what you'll have, yes, you'll have a super user of some kind, um, the SSH keys for a super user, which you'll feed to Ansible, and it will do then it basically does a pseudo once it gets onto the box to install and configure what it needs to install and configure.
SPEAKER_02Got it, got it. Okay. Uh so yeah, the equivalent of root access. Um and uh it goes from there.
SPEAKER_00This is old it's it's all good days of telnet and the root access team.
SPEAKER_02Um don't do this anymore. We don't do those things anymore at all. But we need to understand how this works. Yeah.
SPEAKER_01So that's uh that's how it works at a high level. Whereas a pop-button shaft, you're generally dealing with an agent that's installed on the host before you can execute any of your code. Um so there's an extra level of management involved because you basically have to manage your management tool. Um and there are different approaches people take to that, for example, just having the agent included with the OS image or things like that.
SPEAKER_00Um Does anyone use Ansible to provision uh agents for Chef or Puppet?
SPEAKER_01I'm sure somebody has. It wouldn't surprise me because there's a lot of legacy environments out there, right? Still running Puppet and Chef, right? So it wouldn't surprise me if some admin had just decided I don't want to deal with this and is using Ansible for it. Yeah. Yeah.
SPEAKER_02And also Chef and Puppet, they're uh they're they're both they're also both infrastructure as code tools. I mean, they were kind of the beginning of that of that revolution from a tools perspective. And um their languages are more like uh they're less declarative and a lot more, I guess you call them procedural.
SPEAKER_01Correct. Correct. It's more like Ruby. Yeah. Ruby's the one that's Ruby DSL. Yeah.
SPEAKER_02Yeah. Which is is uh good in that it's super flexible, bad in that you can you get super flexible. Yeah. You can do all that.
SPEAKER_01It's a double-edged sword. Yeah, and and and there's and those agents as well. The other thing to remember about those agents is the way they typically run, and they're configurable, but the way they typically run is they're constantly running in the background looking for changes and updates. Um and one of the issues I've seen come up in many different environments, both when I've been consulting as well as for employees I've worked for, is where there's been sort of a rogue agent running that's overriding configs even after you've changed them. So you end up with production clusters in strange states because you've got a config in place that you were not expecting to have. Um with Ansible, it's pretty hard, it's basically impossible for that to happen. Um, because you actually have to kick off the whole job to run. Um so I think Ansible, in a way, is more is is safer in a way from that perspective, because of its simplicity.
SPEAKER_02Because there is no agent running looking for trouble. Exactly. Um Yeah, alright. So you've got the the YAML, which I don't think any of us needs to pretend that we love on its own merits, but you've got YAML, and because it's YAML and not a turn complete programming language, um your only option is for it to be declarative. Victor used that word before. So you're describing the state of the world that you'd like there to be. And to use Ansible in the operational sense is to take that YAML file and And run Ansible and have it go do the things to the computers, but there isn't any so uh apart from you pressing that button.
SPEAKER_00We need to we yeah, we need to just break down the overall structure, how this like Ansible thing, because there's uh multiple things involved, and once we will you know talk about this, it will make more sense um about the structure. So let's uh let's dive right into um this like an Ansible playbook structure, right?
SPEAKER_02Well, yeah, stop there for a second. Uh that is the place to go next, but I I wanna I wanna kind of put a chapter heading on that. You guys have used the word playbook a few times. Tell me what a playbook is.
SPEAKER_01Yeah, so so in Ansible you have the concept of a playbook, and a playbook contains a series of plays. Um that's why it's called a playbook. And a play could be, for example, um yum install package X. Um or a play could be set my open file limits for this user to X. Um so basically a playbook is is is is the unit of of a unit of work in a sense.
SPEAKER_02Um it is the equivalent to a file?
SPEAKER_01Yeah. Typically, it's usually a grouping of files, to be honest with you. Um typically in a typical playbook, um a simple playbook might only be one file. For example, I have a playbook that I use to set up my laptop. That is one file, because it just needs to install my apps and copy some things around. But if you look at something like CP Ansible, it's many files and it has a directory structure that you can traverse depending on what it is you're trying to automate. So if we look again, an example of Confluent platform and CP Ansible, what we've done is we have a playbook which consists of some subdirectories inside of it, um, each one named after a particular component of Confluent platform. For example, the Kafka broker or Zookeeper, um, Connect, etc. And then within those directories, you have a series of plays, which are files that contain various tasks that need to be run and implemented.
SPEAKER_02Got it. Um so a play is generally a file or always a file, and a play contains tasks.
SPEAKER_00Basically, essentially it's kind of like uh the scenario that will be executed. The certain order, I guess, also in for uh enforced. So you're defining all these tasks that will be part of one play. Um next, uh we do have a certain you know components that defy um like inventory, right? That this is something that we the fleet of machines that we want to manage. Um and we also define roles for uh different um different components that usually it's a different component of our you know thing that we're managing. In in in the case of simple uh web application, it can be Apache web server, Tomcat server, uh MySQL database, like three roles like Apache, HTTPD, Tomcat, and MySQL.
SPEAKER_01Um so and then and then within those roles, for example, so you can think of it from an application perspective, but another way to think about roles as well is from a configuration perspective. For example, we have roles within CP Ansible that configure certain aspects of security, right? That's so though, um, they won't install necessarily a specific application, but they will install some additional packages, say required for Kerberos, for example, and copy key tabs around and things like that. So a role doesn't have to be specifically tied to an application, but it can be.
SPEAKER_02Got it. So um and you've hinted at this a little bit too, but could you guys take us through what the Confluent Platform playbook is? And Justin, this is kind of a thing that you've written. Uh Victor, it's a thing that you've used and evangelized a whole bunch. You guys know it. What is this thing?
SPEAKER_00Let's let's do like a quick step back. And uh like why one of the things why do we decide we choose the Ansible? Um because there are certain so there's there's this history behind this project, right? So uh the history started as a need of these uh things that people were asking. So that's why we had for very very uh not to say a long time, um, some open source project they started as the you know the collaboration between uh the customer success teams, uh the cops, uh engineers. Uh Dustin. Dustin, yeah, what's what's his last name? He he started this open source project in order to um in order to automate uh certain things. Uh we have um Anthony Stubbs, he's uh one of the engineers in uh in our um professional services team. He also was uh doing some work around automating certain steps while he deploying some of the customer work. So we try to like we we have this like uh uh what's the word I'm looking for? Uh critical mass of uh expertise where it's about to explode. Uh and for some reasons we never were like huge fans of uh you know this other tools. We we did some of the some of our um uh system engineers, uh like Jeremy Kostenbrother, he has this uh puppet stuff that some people used, uh, but it never like take off the way how the people like to use Ansible for the reasons that we described uh earlier, with the you know the simplicity, uh declarativity, it's declarative nature, and uh the agentless thing. And over the time this is what happens with like pretty much all popular uh open source projects. So we we find this uh situation where the people were asking that yeah, there's open source thing, how we can use this for certain scenarios, and as a product organization, we realize that this is something that people are actually using, and uh we cannot leave them hanging. So we need to make this official. So I think starting from the conflict platform 5.3, uh the Ansible is officially full Ansible playbook for CP Ansible is officially fully supported, meaning that like if you're a customer, you can always like pick up the phone and start talking to our operations uh engineers if you have some things. We have a dedicated engineer who works on this like uh from 9 to 5. So that's why we have Justin here. And this is where we see that uh conversation with community, conversation with customers lead us to creating something that's you know become a product. So um I'm just passing the microphone to Justin where he was talking about like what kind of problems we're trying to solve and why it is we think the opinions that we're putting in this automation are actually good one.
SPEAKER_01Yeah, so so my background before I started doing automation work was in COPS, which is our customer operations team. I worked, me and Dustin worked hand in hand for a number of years. Um he's the one who started the project.
SPEAKER_02Um I've kind of picked up that's like a like a support role. People have things break and you're gonna be.
SPEAKER_01I'm the guy on the phone helping you, or WebEx trying to help you fix it, exactly. So we started this project. Um Dustin started this project and did a great job, but then got pulled into other things. And so what we found was that customers really love the project. Um I think the last time I spoke to one of our product management team, they were saying something like 30 to 40% of our customers are playing with Ansible. Um, they really love it. You know, they love what we're doing with this project. We should really make a go of this. So that's when I got involved to sort of pick the ball up here and keep running with it. And and the idea was with this project was to, first off, was to simplify and install a configuration of Conflict Platform, because there are a lot of pieces and a lot of configurations. And the second piece, um, which we alluded to previously, was you know, how do we not just get this thing up, but how do we get it up and running with our best practices? So, what we've been trying to do is to draw on my experience as well as the experience of our field team to go, okay, what are the pain points we consistently see from customers in installations? Um, give you an example, and I've referenced it previously, um, I think open file limits, for example. Um, a lot of administrators are not familiar with the concept of an open file limit and why they would need to tune that. So, one of the things we do in CP Ansible is we tune your open file limits to an appropriate value. Um, and that was something we used to see in the support organization all the time. Um another example, how do you tune heap? Um, a lot of admins, again, you know, they're not Java people. They're they're they're Unix guys, but they might not know Java that well. So they have no concept of what heap is and how the JVM works. So we try to set sane heap values to at least get you off the ground and running, running and off the ground in a healthy way. Um so again, it's it's about it's about making it easier to install, but also making things run in the best way possible for you. Um which ideally should should should lead to you know, hopefully higher adoption from our customers, but also uh less frustration as well, right? With how do I get this thing running? How do I get this off the ground? Um so yeah, so that's that's the kind of context of where this is coming from and how I've become become involved with it. Yeah.
SPEAKER_02And that's uh well before you even get to Kafka. I mean, you're talking about file image, talking about heap, it's it's not like there's a Kafka properties file that's wrong. It's it's OS and JVM level.
SPEAKER_01So we're doing OS, we're doing JVM level, and then obviously if we get into Kafka itself, for example, um one of our big focus areas has been security. Everyone struggles with security, um, both TLS and Kerberos. Um we also have Scram in our product, which is honestly a protocol I had to learn about when I first joined. I hadn't heard of it before. Um so it's about how do we configure these things in an easy, consistent way where you just give us the values and the certificates or the key tabs and we sort of take care of the rest of it for you. Um so you can be sure that everything's configured in a clean and elegant way.
SPEAKER_00Yeah, so think with security, it's not about uh like a complexity of the thing, or there's we know that some hidden things and things like that, right? So all this is well documented. However, the usually problem with security is doing right things in the right order and do all the things that outlined in documentation or tutorial or whatever uh tool that you're using to you know configure security. Where you have these automated steps um outlined as a part of this like a task and display book, um, you you don't have a chance to skip those. And we know the steps are important, so this is why we're putting this in place. Uh, usually the problem is is sometimes as customer trying to configure something, it misses one of the steps because you know we're all humans while we dedicated this task to like heartless machines. Um and uh yeah, this is uh this is where this is where the experience will be translated instead of just being um okay, go read this documentation. Oh, let's let's have this as a part of automation step.
SPEAKER_01Yeah, exactly. So it's it's about incorporating all the lessons we've learned as an organization. So hopefully the idea is that our our muscle memory, if you will, gets translated into this project and and customers can take advantage from that from day one. Um so avoid hopefully a lot of pitfalls right from the get-go uh by running these playbooks.
SPEAKER_02Kind of takes your experience and everybody else's experience as support engineers and productizes it. Basically, what it does is it takes what's important about you and puts it in a YAML file.
SPEAKER_01I'll tell my girlfriend that one. Honey, I've been distilled down to a YAML file. You don't need me. Literally, you don't need me anymore. Here's the YAML file. We're good.
SPEAKER_00And she can just it's not a joke.
SPEAKER_01Just YAML putting people out of work since what, 2000 We'll throw some we'll we'll throw some blockchain in there and some uh and some uh AI. And I mean I think we've got the next hot startup fellows. YAML.
SPEAKER_02Yeah, whoa, careful using all those words. We may just have accidentally raised a seed round while we're recording here. Yeah, exactly. Yeah, in fact, wait, there's a there's a guy, there's a guy outside the door in a puffy Patagonia vest holding a check. I don't know what that means, but um yeah, so uh right and and that that makes sense. So it's uh encoding best practices, which are extremely hard won. Um and you know, get getting those into now a thing that's that's push button. And you had mentioned another thing, and that was um yeah, yeah, security. So some of these things are just rough edges that it's not a Kafka problem, it's not a Confluent problem. It's you know, there are five or maybe six human beings who know how to configure Kerberos anywhere. Yeah, you know, and so kind of capturing that knowledge and and making it into YAML.
SPEAKER_01Exactly. And uh my my favorite line about Kerberos is from a colleague of mine was just can we just can't? Just can't. Just can't. That was the whole line of Kerberos. And I just can't just can't. And I just yeah, it's it's hard. Kerberos is hard, man. Um and uh anyone listening who's ever worked for Kerberos, I'm sure I'm sure is you know, having their coffee or whatever and chuckling to themselves. It's hard.
SPEAKER_02It just is. Chuckling, maybe weeping, tears. It is, and uh it is, but also relatively ubiquitous in the enterprise. So like you you can't you can hate it, but it's gonna happen to you. Exactly. So you better know how to do it. There's no escaping it. What are some other hard things um that have come up? Uh obviously, Kerberos is hard just because it's complex and and sort of prickly in terms of its own configuration. But what else is has been hard to do here that has felt like you're really glad you've captured lessons learned and put them into playbooks?
SPEAKER_01Yeah, that's a great question. Um the OS level configuration stuff is like I previously mentioned, is something that no one thinks of. Um so that's a good one. Um login configuration, and that's something I'm actively working on right now to even streamline it even further. Um Kafka, and uh as many open source projects do, you uses log4j. Um you know, unless you're unless you're a Java developer or you've worked extensively in open source, you probably don't know what log4j is, let alone how to configure it. It's it's it's not the friendliest thing in the world to configure.
SPEAKER_02Well, let me put it this way: even if you are a Java developer, you may also not really know how to configure Lava 4J.
SPEAKER_01Exactly.
SPEAKER_00It doesn't mean a thing.
SPEAKER_01Exactly. So it's it's it's so that's something where, again, just having a same template in place way that has the same defaults that you can configure with a few simple variable changes to meet your needs. Um that's another big one, for example. Um yeah.
SPEAKER_02Quick question for anyone listening who is a Java developer using Log4j. Uh when do your segments roll in production? Quick. Just pause right now and take five seconds and see if you can answer this. No, you can't. See, this is why uh this is why it's nice to do these things.
SPEAKER_01Yeah. So I mean I So what you what else? I mean, so Log4j, I mean, I've seen um uh heap tuning, uh JVM tuning. Um what else are we we capturing? Um uh OS limits and configs. Um we're also we implement with uh users and permissions so that all the directories that where we do install Kafka to um are running under a CPK user, and we have a minimal set of permissions in place. Um so only that user can access the appropriate system level files, um, that sort of thing. So sort of other security best practices that often get missed in a lot of organizations. Um so that's in there by default. Um yeah, um we do a lot of work with TLS, both OneWay and Mutual TLS are included, so you have the option to select which one you want depending on what your needs are across our component stack. Um geez, I I A lot of things. Many, a lot of things. Yeah, it's it's it's a long list. Um talking about actually maybe loop this back to YAML, it seems to be our favorite topic today. Um we talk about you know the challenges of YAML sometimes with spacing and carriage returns and things like that. But I mean, if you've ever manually set up a Kafka cluster and actually worked through those configs on your own, um let me let me preface this by saying that for Kafka in general has a really good set of configs out of the box. They're very sane, they're very reasonable. They will meet, say, 75-80% of all use cases, if not more. Um but if you ever do need to tune those configs, there are a lot of configs that can be tuned on a Kafka broker. A lot. Just I I challenge what we like to say.
SPEAKER_00Kafka, good thing about Kafka, there's a lot of configuration knobs. Yeah. Bad thing about Kafka, there are a lot of configuration knobs.
SPEAKER_01Exactly. And I would I would challenge anyone to go to our documentation or go to the Apache Kafka website and just look up the properties that are available. There's a lot. So so again, we try to break that out, again, using YAML in a way that is more English-based, that is cleaner, um, that allows you to set those things in a much simpler and cleaner way, and hopefully that's documented in a in a cleaner way that's a bit easier to understand what it's doing.
SPEAKER_00Um, essentially, with these configuration tools, um they might have a better precedence in terms of uh they might be describing or documenting the system much better than actual documentation. Because again, if it's uh if it's handled by installed version control, you always know what software you're running right now on uh on your system rather than maybe someone forget to update documentation to somehow the flag got changed and things like that.
SPEAKER_01Um relating to this, one of the other big areas we see a lot of issues with um uh both in the field and through support is upgrades. Um and one of the things, again, to give the listeners a bit of a sneak peek, we're working on um it's in its final stages now, um, going through code review and whatnot, is we're we're looking at adding upgrades, an upgrade playbook to CP Ansible. So you'll be able to upgrade in a fault-tolerant way going forward, your your Kafka installations. We've done the work whereby we do an analysis of the current state of your Kafka cluster, you know, how many brokers do you have, how many petitions are online, offline, and we and we do the upgrade and roll them in a safe way. So that's something that's coming out soon that I think is is something that's gonna be a huge boon to our customer base.
SPEAKER_02Ah, fantastic. And uh as always, when we talk about when we're talking about confluent things and we talk about future products, we don't say when and we don't promise much, but that is nice to know that that is um that's a thing to think about coming is what I would say to that.
SPEAKER_01Sooner than later. Excellent.
SPEAKER_02Okay, oh careful, but soon. I like it. Uh okay, final question for you each. Victor first, then Justin. Um, if somebody is looking to get started with Ansible, either where do they go first or what's the the most important concept they should get hammered into their head first?
SPEAKER_00Uh in terms of like Ansible in general or in terms of a confluent platform?
SPEAKER_02Uh well, it could be the CP Ansible playbook or Ansible in general. I'll let you take it either way.
SPEAKER_00Well, I would say that like uh the standard uh like the standard website, the Ansible website is pretty good because right now Ansible is uh uh is handled by Red Hat and usually Red Hat have all things there. Um so you can have documentation, it's open source. If you striving for consultancy and training, they also have plenty of those. Um in terms of like learning things, I would say it's a relatively easy to learn, but since like uh if you want to focus in on some of the things, you can just go and read existing um playbooks uh if you want to see how it works. It's not that difficult even for me, yeah. For a person who kind of like is struggling to understand YAML, but it's just because of me. Um the and uh yeah, so you need to install this uh small like a Python package called Ansible on your computer, uh, defining the set of machines that you will be um automating in your in inventory, and uh you can see whatever you know things you're running. Um in terms of yeah, I will I will pass this uh microphone to to Justin. Justin, you can tell about um like what we have from perspective of our documentation.
SPEAKER_01Yeah. So yeah, so I think you hit it. I I think I think yeah, there's two pieces to this. There's how do I get started with Ansible, Ansible.com, owned by Red Hat. It's a Red Hat tool. Highly recommend going there, reading through the documentation, it's really good. There's also some really good tutorials on there as well as on YouTube. And what I would recommend is you start with something very simple. Start with a single host if you're on a cloud provider or even on your laptop, and just see if you can get an application to install. Just you know, get the equivalent of Yum install working on a host. That's where I would start. Um from there, we have our repo, which is on GitHub, so github.com slash confluentink slash cp dash ansible. Um once you get your first playbook reckoning, I'd recommend going and cloning that and looking through the code and just seeing how we do things. Again, it's fully open source. So um and we have an active community on there, so you can ask questions. And yeah, and just you know, start small. Start with installing a single application um to get your feet wet, and then yeah, start looking at other projects, including ours, to to understand how it's actually built. That would be my recommendation.
SPEAKER_02My guests today have been Victor Gamov and Justin Manchester. Vic and Justin, thanks for being a part of Streaming Audio.
SPEAKER_00Thanks for having me too. Thank you. It's uh always great to be here, and uh, we'll hear you next time.
SPEAKER_02And there you have it. Before I go, I want to tell you that we have a pretty cool new offer to help you get started with Confluent Cloud without you having to pay for anything. If you're a new user and you go through the regular sign-up process and start using Confluent Cloud, your first $50 of usage per month are free. This will last for the first three months after you sign up. So that's $50 per month of serverless Kafka for three months at no cost to you. So go to the sign-up link in the show notes, I don't want to read you the URL, and sign up now. I think the only thing I could really do more is write your code for you. And I think we can both agree that's too much to ask. So check it out, and hey, let us know how you like it. Anyway, as always, I hope this podcast was helpful to you. If you want to discuss it or ask a question, you can reach out to us on Twitter at Confluent Inc. or reach out to me at TLberglund. That's T L B E R G L U N D. Or you can hit us up in Community Slack. There's a sign-up link for that in the show notes as well. And while you're at it, please subscribe to our YouTube channel and to this podcast wherever fine podcasts are sold. And if you subscribe through iTunes, be sure to leave us a review there. That helps other people discover the podcast, which is a good thing. Thanks a lot for your support, and we'll see you next time.