Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov

Ask Confluent #14: In Control of Kafka with Dan Norwood

Confluent, original creators of Apache Kafka® Season 1 Episode 43

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 23:50

Is Apache Kafka® actually a database? Can you install Confluent Control Center on Google Cloud Platform (GCP)? All this, plus some tips from Dan Norwood, the first user of Kafka Streams.

EPISODE LINKS

SEASON 2
Hosted by Tim Berglund, Adi Polak and Viktor Gamov
Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
Music by Coastal Kites 
Artwork by Phil Vo 

  •  🎧 Subscribe to Confluent Developer wherever you listen to podcasts. 
  • ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
  • 👍 If you enjoyed this, please leave us a rating. 
  • 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
SPEAKER_00

Has anyone ever asked you if Kafka is a database? Well, today my co-host Gwen Shapira and Confluent Tech Lead Dan Norwood discuss that very question. And they talk about Confluent Control Center too. It's all in today's episode of Streaming Audio, a podcast about Kafka, Confluent, and the cloud. Welcome back, everyone. Let's get to it.

SPEAKER_02

Welcome back to Us Confluent, where we answer questions from the internet. I'm your host, Gwen Shapira. And it's been a while, but with me today we have Dan Orwood. And Dan Orwood is really an amazing tech lead who joined Confluent pretty much four years ago.

SPEAKER_01

Yeah, just well.

SPEAKER_02

Yeah. And basically built Confluent Control Center from zero to the amazing product that it is today. So yeah, it's been basically we have a question about control center, and I brought the single most qualified person in the universe to answer that.

SPEAKER_01

Yeah, hopefully I can do it.

SPEAKER_02

So yeah, so we have a bunch of questions and not that much time. So let's get to it. Okay, so the first question is from Martin Klepman's keynote in Kafka Summit London 2019, where he talked about is Kafka a database. And the comment was from Mechanical EI is so Kafka is a database? And so I assume he meant Kafka. And I'm trying to figure out if you watch the keynote or not, because like is it a question you ask before or after you've watched that? But what do you think? Is Kafka a database?

SPEAKER_01

Um yeah, I think like the way I think about it, anyhow, is like obviously it's like a messaging platform, and that's how a lot of people use it. Um but it's also like it gives you like the basic building blocks to build your own databases. It's sort of an easier an easier way for me to think about it. Because like most databases are this like sequential log of changes that people are building up. Um so you like Kafka, like a compacted topic in Kafka is kind of a key value store. It's a key value store, stored in a weird way. Right? Like not normally the way you'd store it if you're storing it in memory for yourself to query. Um but it is a um it's another like view on that sort of data. And so you can also do other stuff on top of it.

SPEAKER_02

But it's you can build like infinite types of views on top, actually. Yeah, that's that's the really nice part.

SPEAKER_01

Right. So like if it's not if it's it's not an RDBMS, doesn't let you do all that sort of stuff. But conceivably you could be building RDBMSs on top of Kafka, which is some of the stuff that like LinkedIn has done with like espresso and and making uh Postgres uh go out.

SPEAKER_02

Yeah, that was kind of my my first thought. Like I've used relational databases before. Kafka is clearly nothing like them. Uh but I think the view of Kafka is a building block for a database is actually kind of a good way to think about it, and also 100% the way that Martin Kleppman really presented it. He said that hey, if we look at all the things you want out of a database, you want the ACD guarantees, you may want even stronger guarantees like referential integrity, you may want uh all kinds of uh easy-to-query tables that match different data formats. Kafka will not give you all of it out of the box, but he kind of shows how he can build them using Kafka. Right. Um but obviously you have to kind of still build them yourself, so it's not a database in the sense that you have to build a database out of Kafka.

SPEAKER_01

Right. Yeah, it's it's all the building blocks. So if you have a very like specific use case that you want to try to build a large scalable system around, like it gives you all the all the pieces to do that. Yes. Um but if you have something that fits into something else that is a database that already scales to what you're using, then that might also be a good way, uh good product to use.

SPEAKER_02

That's the thing, right? I mean, if you it gives you the building blocks to do everything yourself, but because you have it, there is a possibility you'll get it wrong. Right. And it's obviously not as cheap as taking something off the shelf if it's a good fit. And obviously Kafka is a good fit for the streaming use cases, which there are no really strong alternatives. Okay, the next question is around install and run uh Kafka uh with Confluent Control Center. And Tim Bergland basically went through how he installs and runs uh Confluent control center to monitor Kafka, but um he showed it using Docker, doing it in your own local laptop. It's a fairly different experience than doing the whole thing in the cloud. So our friend of YouTube, Tusif Zaki, asked any idea how to start Confluent Control Center with GCP? So let's maybe start by what is Confluent Control Center and why does Tusif Zaki so excited to use it?

SPEAKER_01

Um Yeah, sure. So uh Confluent Control Center is our on-premise solution for people to monitor and manage their Kafka deployments. And so it allows you to monitor and manage uh however many uh uh Kafka clusters that you have out in your system. Uh it also allows you to uh interact with Connect, uh use schema registry to manage all of your schemas across all of your clusters, uh as well as interacting with KSQL. So it gives you like a bunch of abilities to uh like mess around, like write KSQL queries, add connectors, remove connectors, um, all that sort of fun stuff, as well as giving you sort of insights into how your brokers are running.

SPEAKER_02

So you said that it's on-prem. So he's going to use it in GCP. Could he use it with Conflant Cloud, or does he have to kind of self-manage his own Kafka for that?

SPEAKER_01

Yeah, so you generally are only gonna want to run control center yourself if you're also managing your own Kafkas. Um so like I'm imagining maybe he's running his own Kafkas for some reason in GCP.

SPEAKER_02

Maybe he likes waking up by alerts at the middle of the night. Who knows?

SPEAKER_01

Uh, maybe. Um which is another thing Control Center can help you set up, actually. Uh there's a bunch of alerts on the Kafkas that you have uh that you're running. Um so if you have like your own deployment of Kafka and you aren't ready to move to our cloud product yet, um then control center can help you there. So maybe he has uh his own brokers running and he wants to get some more insights into them. Um so like uh but if you are willing to run on our cloud product instead, then you get uh quite a few of the features from Control Center um already built into the cloud interface.

SPEAKER_02

So you get a schema registry and you get KSQL and you get connect.

SPEAKER_01

Right. And the only things you don't really get are the things that you probably don't want to have to do. Like waking up in the middle of the night.

SPEAKER_02

That's the thing with confirm cloud. I wake up in the middle of the night. You don't have to.

SPEAKER_01

Right. Like we don't have um like some like the monitoring and the alerting pieces, right? Like the pieces that like we're paying somebody to go do, um and that you you don't so you don't have to deal with it. Aaron Ross Powell Yes.

SPEAKER_02

So we are now getting fairly close to a new uh control center release. Can you give us like a sneak peek of the coolest feature we are looking forward to seeing? Or is it still secret?

SPEAKER_01

Yeah, I think we have uh sort of a revamp of the entire UI. Um and not not really uh so there are some new features, but a lot of more uh trying to make it more natural how people interact with or how people's mental model, how they think about their Kafka deployments. Um trying to make it uh more natural flow for them to go through the app. Um so the idea is like when you open the open control center, it used to be uh there's it's a little bit difficult to move from cluster to cluster to understand how to go all the way back to the top. Right. And and you didn't really understand how they how different services were connected to one another. Um so you may have like three Kafka clusters and two KSQL clusters on each one of them. Um and it wasn't quite obvious in the old UI how those things were uh there's a bit of a hierarchy there. And now the new navigation system builds that hierarchy into the navigation. So you'll see like here's a list of all of my Kafka clusters, here's the multitude of services running on top of them. And you click on one of those Kafka clusters and then you can uh dig through it. So hopefully everybody has a um uh it's much easier for people to move through the UI.

SPEAKER_02

The thing I really love about having usability as kind of like the top goal for release is that suddenly like release blockers are things like people can't find how to do something, which is definitely not like a traditional release blocker, but it does kind of put our users front and center, which is something I really love. So yeah, I hope all of you will enjoy that. But now to the actual question. Uh how do we start confluent control center with GCP?

SPEAKER_01

Right. So there's a couple of different ways to do this, really. Um so it sort of depends on how you are launching your services on GCP. Um because you can just bring up machines, you can log into them, you could run whatever you want. Um, you can have a bunch of like there's a multitude of ways to do that. Um I think probably the easiest way is if you just want to use our uh Docker images. Um so in GCP you can just select uh the Docker images. You use Docker.io slash confluentink slash cp dash enterprise dash control. Anyhow, yeah. Um huge block of text right there. Um uh you can just use that and then you set uh very a couple of environment variables. We will also include those. Um and that'll let you uh point control center at whatever Kafka brokers you have available. Um so all of our uh Docker images are uh maintained. There's some open source repositories that we have that have um uh like actually the one that uh Tim uses in this demo uh that go through and show how to set up all the environment variables and how to get everything connected. Uh so basically it would just be reusing all of that but pointing it at your services in GCP.

SPEAKER_02

Yeah, I think the main things to remember is that control center actually takes memory, a good chunk of it. So pick at the you know one of those high mem instance types for your um deployment. Um otherwise you'll probably have a bad time. Uh it also by default runs, I think, a lot of threads. I'm remembering something like eight stream processing threads. So either get a lot of CPUs or tune down the number of threads. And the last thing, you will want to access control center from port 9021. Uh so make sure the port is open between your browser and the GCP. And you if you use things like uh private network environments and firewall-like thingies, then you have to be careful that you can actually do the communication that you need to do. That's uh usually that's how I fail. I spend the entire day troubleshooting, and it's like, oh, it's the port. Yep. In Docker, it's actually even harder, right? Because you have to do uh port mapping.

SPEAKER_01

Map ports from one to another and get your Docker networking set up, right? Yeah.

SPEAKER_02

Yep. I think that's how I end up spending days. Okay, the next question was actually probably the easiest to have ever done at USConfluent. Uh Cry R responded to transforming data part one, Apache Kafka streams API, and said it would be nice to know how to construct K-tables. And we 100% agree, it's kind of important knowledge when you use Kafka Streams to know how to construct K-tables, which means that you have to progress to the next video in the series, transforming data part two, which before the first minute is over, you will know exactly how to construct a K-table. So we've got you all taken care of. Just continue watching the series and all your questions will be answered. And then the next question is conveniently about transforming data part two. Uh too easy for Mike said, it is a nice video. We agree, it's a nice video. Uh, but I don't understand why the starters for value is long. Shouldn't it be a song surday because it's a song object? You are 100% correct, too easy for Mike, and thank you for your comment. Sometimes we kind of simplify things in a way that's not 100% correct when we copy code to fit into a slide. So I think one of those things happened. One thing you want to keep in mind is that the API actually changed a bit since the time the video was recorded. So I think the best thing is to use the new API, which makes it slightly easier.

SPEAKER_01

Yeah.

SPEAKER_02

So I'm going to publish a link to an end-to-end example that includes how to do a group I with the CERTS kind of the correct way. So that's a nice thing about streams these days. When we started X, there was literally so then here was literally the first streams user in the entire world, as far as we know.

SPEAKER_01

I was an alpha tester internally. Yeah. All of Control Center or all the monitoring in control center is built on top of streams. So yeah, this is uh this is actually like the third or fourth version of this API. It used to be even uh a little less readable.

SPEAKER_02

Yes. No, I think that that was also a problem when writing the book, like every new revision I have to rewrite the streams chapter basically from scratch. Uh because I keep getting complaints that hey, you got the your code is crap, you got every API wrong, it's nothing compiled. And I'm like, it's it was right back then, I swear.

SPEAKER_01

Yeah. Yeah, streams is moving very quickly. And a lot of their iterations have been around trying to actually sort back to the control center, like making making things more usable for people so like you don't run into issues like this.

SPEAKER_02

Yeah, I think it is getting uh way better over time. Although CERTs are still continual uh pain point, I think.

SPEAKER_01

But um You can just use KSQL so you don't have to think about Surties.

SPEAKER_02

That's true. That's true. Uh can you actually do everything that you could do in Kafka streams in KSQL these days?

SPEAKER_01

Uh I think that there's still some joins that aren't really possible. And then there's uh uh there's some work being done in K S right now, KSQL has uh sort of a bad keying story. Yeah, that's true. So you can't have sort of complex data types as keys, but it's also something that they're working on diligently.

SPEAKER_02

So cool. So what about interactive queries? I kind of heard that recently, or not so recently anymore, your team kind of started using those to speed things up.

SPEAKER_01

Um yeah, actually, so uh control center has always used interactive queries, like from the very start. So um and just to give sort of like a level set on what interactive queries are, I guess. Um so when you start up uh a streams application, you can build uh a table and it materializes in in memory and onto disk using RocksDB. And an interactive query allows you to interactively query that RocksDB instance. Um and it's constantly being updated in the in the background by the streams application. Um so like when we show you uh metrics data inside of control center, all that data is being live updated into a Rocks database, and then when the browser hits that uh endpoint, we read out of there to show that.

SPEAKER_02

Um, pretty soon we can expect a blog post into which you explain to the entire world how amazing interactive queries are and how useful they've been.

SPEAKER_01

Yeah, we can we can talk about doing a blog post. Um Yeah, there's there's some there's some really nice things. Uh right now stream sort of uh same same thing we we were talking about with uh is Kafka a database. Um right now, like you technically have interactive queries, but using them can be sort of difficult. I see for similar searches.

SPEAKER_02

Which is already a huge user usability improvement.

SPEAKER_01

Yeah. Yeah, part of the issue is that you you need to be able to query across all of your instances. Yeah. Um and then you sort of need to do a scatter gather and then like relay out all the data in the order you want to. Yeah, so all more stuff to make better in next release.

SPEAKER_02

And a question about the last AsConfluent. Um it's kind of a long question. I'll jump to the gist of it. Chirag Satasia asks, uh, we're using different machines for Kafka and Zookeeper. What's your recommendation? Should we continue to use Kafka and Zookeeper on different machines, or we should combine them? And that kind of goes to risk appetite versus uh cost kind of thing. You definitely decrease risk by separating them. Uh Zookeeper needs good access to CPU, it needs to be able to write to disk extremely fast. By having its own machines, you kind of guarantee that Zookeeper and Kafka will never interfere with each other. That's our support recommendation because support to try to make you successful and not take any unnecessary risks. Uh you, as a person paying for the machines, may choose to take more risks. If your Kafka is not super busy, if you have spare CPU, and especially if you have spare disks that can be dedicated just for the Zookeeper transaction log, you could combine them. I hope this description convinces you that you are taking some risk because anything goes wrong with with the CPU or the disk, uh Zokeeper will stop working and your Kafka will stop working. So decide how much you want to pay for extra machines versus uh pay for downtime. That's basically the calculation we all do all the time. Uh if it's any help in the cloud, it's uh different machines all the way. Like we have SLAs to meet, and uh our appetite for waking up in the middle of the night is incredibly low. We also have like literally hundred clusters. So the com probability that anything will go wrong on one of them is kind of high.

SPEAKER_01

Yeah, one one thing maybe to add here is if you have multiple Kafka clusters, you can point them to the same Zookeeper. That's a very good point. So if you have, say, five AK clusters, you still can possibly get away with only having one Zookeeper cluster to help manage all of them.

SPEAKER_02

Yeah, that's a good way to have basically the best of both worlds, as long as Zookeeper is not uh super busy.

SPEAKER_01

Right.

SPEAKER_02

Uh the thing to watch out for is like if you create and delete topics at rapid rates at specific points in your life cycle, that's where load on Zookeeper gets a bit tricky, and if you kind of share dev where everyone's always deleting and creating nodes with something that's more uh production, maybe that's not a good sharing scenario.

SPEAKER_01

Right.

SPEAKER_02

Ask me how I know okay. Now we get to the portion of the program where we get to pat ourselves in the back because you guys keep giving us amazing feedback. So Carlos Valencia commented on KSQL introduction, probably the videos that generated the most compliments ever in the history of our YouTube. He said, Whoa, thank you for the video. I've been working with Kafka for one year, I'm new, and I would like to learn a lot of it. Hello from Colombia, great job. So awesome that you started learning Kafka, uh Carlos from Colombia, and we hope that all our videos in our channel and we keep adding more and more of them will be incredibly useful for you. Patrick Nazar uh listened to Nan or Kiddy in the Kafka Summit 2018 keynote and said, I love it, I absolutely love it. The giant leap closer to the singularity, I tell you. Do you think that closer to the singularity is a good thing?

SPEAKER_01

Uh yeah, that's a debatable.

SPEAKER_02

What if the singularity is terrible? I mean, there's no guarantees that it will be a good. Yeah.

SPEAKER_01

Yeah, it's difficult to see beyond the event horizon there.

SPEAKER_02

Yeah. What if the robots turn all of us into paper clips? Like the whole thing is very scary to me. I don't know if I want to do a giant clip, but I kind of like that you think we did it. More near Narcid feedback. This one was Kafka Summit 2017 keynote. Uh Sayad Awesh Rahman said, fantastic presentation and overview, really worth spending time watching it, waiting for more thumbs up to confluent him. So I don't know, definitely worth time spending time watching this. You can also spend time watching other Kafka Summit presentations. They're all worth your time. Other keynotes are also worth your time. So yeah, don't just stop there. We have tons of content that is worth your time. Uh KSQL and other stream posting tools in Kafka by Nick Durden, which is probably our favorite SE. He died really does amazing presentations. And Siddharth Cotwal here says that he's a legend, which I tend to agree. I think he's like quite legendary. Uh his presentation skills are actually getting progressively better. I've seen him do kind of an intro to KSQL for our sales team. And like I usually don't enjoy sales presentations that much. I felt a bit weird, but he just did an amazing job.

SPEAKER_01

Yeah, it's terrifying at Nick's getting better at convincing people to do stuff.

SPEAKER_02

Yeah, the thing is that he now practices all the time, right?

SPEAKER_01

Yeah.

SPEAKER_02

It's like yeah. But uh yeah, I'm I'm I'm super happy. It's always good to see people getting better at something they were already ridiculously good at. Uh Nerkidi, Kafka summit, London 2019. I now I'm sorry that I did we had 2017, 2018, and 2019. I feel bad for not doing it in the right order. Uh to Steve Zacchi said, I follow these summits, and Confluent always surprises with new features, which is 100% true. Thank you for following the summit to Steve Zacchi and thank you for your qu good question earlier on. Um I think the features are good. I think this keynote is worthwhile because it also really shares real life lessons from what we learned running in the cloud. So definitely a worthwhile watch right here. And for the same presentation, Abel Morello said, I'm excited to try Confluent Platform. That's fantastic. Uh you may even be excited to try Confluent Platform in the cloud, in which you can get started in about 30 seconds and just create a cluster and use all the control center features and KSQL and Connect and everything in with a very usable interface. Okay, I think that's about it. Enough patting yourself in the back, right? That was lots of fun. Thank you so much for joining us.

SPEAKER_01

Absolutely. Thanks for having me.

SPEAKER_02

And I think it's time to go back and build more uh control center features, right?

SPEAKER_00

Yeah, something like that. Hey, you know what you get for listening to the end? A Kafka Summit discount code. Kafka Summit is coming up on September 30th and October 1st in downtown San Francisco. And you can get 30% off if you go to Kafka-summit.org and use the discount code AUDIONET during checkout. Just enter Audio 19 while registering at Kafka-summit.org, and that 30% off is all yours. I'd love to see you there. But hey, I hope this podcast was helpful to you. If you want to discuss it or ask a question, you can always reach out to me at T L Burgland on Twitter. That's T L B E R G L U N D. Or you can ask Gwen at at GwenShap, that's G W E N S H A P. Or you can leave a comment on a YouTube video or reach out in our community Slack. There's a Slack signup link in the show notes if you want to register there. And while you're at it, please subscribe to our YouTube channel and to this podcast wherever fine podcasts are sold. And if you subscribe through iTunes, be sure to leave us a review there. That helps other people discover the podcast, which is a good thing. Thanks for your support, and we'll see you next time.