Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov

How Kafka Expert Robin Moffat Tackles Open Source Problems | Ep. 6

Confluent Season 2 Episode 6

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 24:50

Today, Viktor Gamov talks to his colleague Robin Moffat (Confluent) about his career in data engineering. His first job: paperboy. His challenge: working at a retailer with Oracle materialized views as well as teaching others how to productively approach Kafka’s internal systems.

Blog posts mentioned in the podcast:
► Oracle Materialized Views troubleshooting: https://rnm1978.wordpress.com/2011/01/08/materialised-views-pct-partition-truncation/
► Kafka Listeners explained: https://rmoff.net/2018/08/02/kafka-listeners-explained/
► Kafka Connect converters: https://www.confluent.io/blog/kafka-connect-deep-dive-converters-serialization-explained/

Follow Robin:
► Blog: https://rmoff.net/
► X: https://twitter.com/rmoff/
► Bluesky: https://bsky.app/profile/rmoff.net

SEASON 2
Hosted by Tim Berglund, Adi Polak and Viktor Gamov
Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
Music by Coastal Kites 
Artwork by Phil Vo 

  •  🎧 Subscribe to Confluent Developer wherever you listen to podcasts. 
  • ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
  • 👍 If you enjoyed this, please leave us a rating. 
  • 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
SPEAKER_02

From delivering papers to building materialized use in Oracle database. This is Confluent Developer.

SPEAKER_01

One of these things that it's arguably more difficult than it has to be, but it is very, very flexible. That's the uh the flip side of it. A ton of people have that problem and benefited from that write-up. So that was that was a satisfying one to fix. And like people will be taking that and then like, well, let's just move that through into production because I found this thing on the internet that says it, which is like awful.

SPEAKER_02

In this episode of Confluent Developer, I interviewed my friend and colleague Robin Moffat about his past, present, and future. How Paperboy, in his younger years, became an expert in Oracle databases and Apache Kafka, Kafka Connect, and stream processing. Hello and welcome. My name is Victor Gamov with Confluent, and this is Confluent Developer. I'm your host today, and I'm thrilled and excited, and I cannot even express with my word because my vocabulary is not that rich with English language. But I do have my friend and colleague Robin Moffat today with me. Robin, welcome to Confluent Developer.

SPEAKER_01

Thanks for having me back.

SPEAKER_02

For those uh five people who are watching us, probably there would be two who never heard your name before. Can you just quickly introduce yourself, what do you do? And uh, you know, it's it's very rare that um people don't know who you are, but still, you know, maybe you know, when my mom will listen, she wouldn't like to learn. Uh, who are you and who where are you coming from?

SPEAKER_01

Yeah, so I'm I'm Robin, I'm based in the UK. Uh, as you can tell from my accent, uh the north of the UK. Um, I'll try and speak a bit more clearly than usual. Um, and I work at Confluence. Um, I've been working with kind of Kafka and Flink and stuff like that. Um, well, back in 2017, I originally joined Confluence, um, and I've been working with data in general for over 20 years now.

SPEAKER_02

Uh little unknown fact. Uh, everything that I learned about Kafka in the very beginning of uh Kafka, I learned from Robin. I was following Robin uh way before I joined Confluent and way before Robin started doing official uh developer relations, uh, when I think the Robin was with what, like partner engineering, and you were engaging with with uh customers and partners.

SPEAKER_01

Yeah, that's right. That's before I even knew that devrel was a thing. Like it was just I spoke at conferences, I wrote blogs, and then uh uh someone pointed me to this devrel track at a conference, and the rest is history.

SPEAKER_02

Yeah, we uh we were doing DevRel uh without it. Was a thing, uh before it was a thing. Uh one of the one of the famous demos that I I also basically built my reputation on top is a Robin's demo of processing uh streams of uh tweets when the Twitter was a thing and the fire hose were not expensive. Um so you could use the Twitter fire hose uh without robbing a bank. But uh yeah, it's it's great to have you on the show. Um uh many of you uh who watching us hopefully watch the Robin's videos. Uh Robin known to be a go-to person about the Kafka Connect and how to bring virtually any source from anywhere on the world into Kafka and process this with SQL. We always uh have this conversation. Uh Robin, you should start using Java. And he was like, No.

SPEAKER_01

You can do plenty enough with a config and some SQL. Enough for me anyway.

SPEAKER_02

Exactly. And um, I'm always wondering, and uh and now many of you already watched a few uh first episodes, you understand the premise. We would like to talk a little bit about the origins of um the excellence that you are today and how everything starts, like humble beginnings. So, could you tell us what was your first um first job? It not necessarily needs to be uh IT job because we uh we discussed uh different ways how the host and co-host and other guests uh were gaining first money. Can you uh uh talk us about what your first experience for making some bucks?

SPEAKER_01

Yeah, yeah, no, and I heard the episode that uh the Adi did and uh say again.

SPEAKER_02

Or pounds, I don't know.

SPEAKER_01

Pounds, yeah, pounds sterling. Um so I was thinking about this beforehand, and like there's my first job out of university, and blah blah blah. But my first actual job, and you're saying that was when I earned some money was as a uh a paper boy. So the I don't know if you even get them anymore, you're probably not allowed to, but like news agents, you kind of like you get 50 pence for delivering newspapers, and you'd be out there in the wind and rain, and um yeah, it was a it was a good job. Um but yeah, it's a shame that kids don't get to do that anymore. I think it's uh character building.

SPEAKER_02

We we don't have a paper boys here as well. There's a guy rides around in a in a van and throws paper. Yeah so there's no uh boys who you know riding bikes and and and bringing paper to you anymore. And uh but what was the um what was the moment where you kind of like bound yourself to computers in IT?

SPEAKER_01

Like what was your first IT job just for uh so that that was that was before university, actually. Um so I've always mucked around with computers, um, and I got a job for a local college doing work on their website. So this is going back a long time.

SPEAKER_02

So this is like year it was uh what year it was?

SPEAKER_01

1995? 96, like that.

SPEAKER_02

What was the uh technology stack on this website?

SPEAKER_01

So this was HTML, written by hand, and then a bit of Dreamweaver as well, if you remember. Oh Dreamweaver, yes, I remember. Which it wrote absolute junk, but it could do some quite clever effects. Um so a mixture of those two, um, and then like FTPing things to get them up onto servers, and like it was, yeah, it was you'd look back on it now, and it's like kind of ridiculous. But then almost with like static site hosting, it's like we're gonna be able to do that.

SPEAKER_02

We're doing this today, exactly like this. Uh the GitHub pages, uh, Vercel or Netfy, whatever. We're doing exactly this. Uh, we FTP, but like in this case, it's not FTP, it's a little bit more safer. Um uh but the idea, idea is the same. I also use Dreamweaver as my first HTML editor, my first uh web page. That um it was not a website. I was building a you know, remember the the CD uh CD ROMs they have this like auto-start uh inf and I was building like an auto-start script so when the CD-ROM would be inserted, uh it has a bunch of um pieces of software that I gain through different sources and that which would not be named. Um and I built with the HTML build this uh interest screen. Um so when when uh when you joined Conflant, you already was very established uh in a world of databases, you were part of uh the the group of uh luminaries that Oracle recognizes called Oracle Ace, right? Um and um can you talk? So I I think we're gonna talk about databases a lot. I I don't know, I I have this feeling, but um the the premise of this um of this podcast is that what was the most interesting problem you ever solved or you didn't solve, but it still kind of bothers you uh until uh until today. So I would be curious about uh this this uh problem um that you you know that you solved uh with with the technologies of uh that time.

SPEAKER_00

Now a quick word from our sponsor. Confluent developer the podcast is brought to you by Confluent Developer the website, which has everything you need as a developer of data streaming systems. And it's completely free. We've got curriculum, hands-on exercises, executable tutorials, the online data streaming engineer certification, also free, a way to find a meetup near you, those are free. Everything is there. I really want you to be successful in your journey as a data streaming engineer, and this is the site that has what you need. Check it out at developer.confluent.io. That's developer.confluent.io. Now back to the show.

SPEAKER_01

Uh it's a good question. And there's so before I was at Confluence, I was a consultant. And so kind of like by nature, you end up seeing lots and lots of different challenges like going around from different customers each week on week. Um I think the one that kind of like sticks in my mind is actually even before I was a consultant, I was working for a retailer and like building a data warehouse, and we were working with Oracle uh 11G, I think it was, um, and building out materialized views. Um, and it was one of these things where and like the Oracle community was kind of rich. Um and rich, yeah, you have to be rich the right, but like there's a there's quite a rich community around it, but not in the same way that like with open source, really. And so if you found someone who had the similar problem to you, that'd be kind of great and you would look at it together. But if not, you would be completely beholden to Oracle support portal. And so there was something weird that went on with the materialized views, and it was just one of those things where all you could do was like run this thing and poke the black box and see if you could get it to react differently to it, and it would be it was to do with like kind of partitions, you could you'd refresh the partition and it'd be really, really slow. Um, and it's that thing where you like you start going through the manual, and it's like, well, it's not behaving how it should do in the manual, but because of the volumes of data, you try and reproduce it, it's much more difficult to reproduce unless it's in production. And if it's in production, you can't just go mucking around with stuff. Um, but that was good, and that kind of actually got me started blogging because I like I found these things that they weren't in the manuals, they weren't on the support portal. I was like, well, I kind of found that useful, so I'm gonna write it up. And then it was really cool, actually. There's a guy called Doug, uh I think it was Doug Burns, I forgot his name, but he he got in touch and he was one of these guys, like everyone knew who he was in the community, and he blogged as well, so I'd read all of his stuff. And he dropped me a note, he's like, I've seen your blog. I was like, oh my word, like it's this guy. And and he shared his knowledge that he'd found from this problem as well. Um, in the end, I think I can't quite remember how we got around it. I think they probably tried to sell us Exadator or something. But it was it was just one of those things where it was kind of like a real tricky problem, you kind of you're working with different people on it, uh, you get this support from the community. Um, so that was a good one. I kind of kind of solved it, uh, but kind of didn't at the same time.

SPEAKER_02

And uh with with this type of experience, you kind of like um took this um mentality, I would say, right? So and uh started building community. What was the uh the most interesting problem you solved um with Kafka, Connect, and streaming? Maybe um something more you know the kids these days would can relate because people usually you know you don't don't know what Oracle is or they know uh what Oracle is because it turns out Oracle Cloud is probably the most popular cloud. Based on the earnings report from Oracle, looks like the Oracle Cloud is is killing Oracle's doing just fine.

SPEAKER_01

Yeah, um yeah, and with Kafka, I think probably the the one that most people will know me for based on blog views is around Kafka and listeners. So this wasn't like it wasn't so much solving a problem, it was just understanding how a thing worked. So it's the kind of thing that for a lot of people is a problem. So there's a setting, this is actually putting me to the test here because it's a while since I looked at it. But if I remember rightly, within within Kafka, you've got a listener, which is where the you've got the broker and the clients and uh talk to the broker through the listener. So like it's it's on this particular host uh or IP on this particular port. And because Kafka is distributed, you've got multiple brokers. So when the client initially connects to the broker, the broker tells it like you'll find me at this address. Um, which if you're just running it like just well, sorry, if it's been set up or configured, someone gives you the broker and you connect to the broker, and then like that's all great. But if you're running this for yourself, like you're just installing it on your laptop, um, and particularly, and this is why it became such an issue, if you're running it within Docker, networking suddenly comes into play. Um on your laptop, you'll probably get away with it, and you just like hard code everything to like a loot back address, and it's fine.

SPEAKER_02

I think it's uh also worth to mention that this problem also usually comes when you're trying to run like more than one Kafka broker. Um, and in probably even one, there would be some issue, especially inside the Docker. Just one will do it. Yeah, yeah.

SPEAKER_01

So it it comes about if you if you run your Kafka uh broker inside a Docker container and then try and connect a client to it either from the host machine or from another Docker container that's not networked correctly to it. Because it'll reach into it, it will say, like, go connect to that container, but then from within the container, the broker says, Oh, well, this is my address, which would just be like a little local Docker thing. So then you'd see all these awful solutions on the internet, which would, and this is the like the downside of people sharing it off of the internet, is that sometimes it's complete junk. Um, and they'd be like, Oh, well, just go in and like hard code your et cetera hosts file to like hard code this thing. It's like, no, that's like and like people would be taking that and then like, well, let's just move that through into production because I found this thing on the internet that says it, which is like awful.

SPEAKER_02

Um that's how we do in this these days, yeah.

SPEAKER_01

Just find the internet and yeah. So uh ChatGPT tell me how to configure it, configure it.

SPEAKER_02

Um now, now probably it's a good thing that the chat GPT know your blog post, so it's a good thing. Well, exactly. Yeah, that will know how to configure this. Uh yeah, yeah, good one. Yeah.

SPEAKER_01

That's where advertised listeners comes in. So you've got like the listener, which is like the broker will listen on this particular part. Advertised listener is like this is what the broker tells the client. This is where you can find me. Like, um, you're connected to this particular host, um, but the other nodes within this cluster at are at this address. Um, I hope I'm not butchering the explanation, but I wrote a good blog about it. If you uh actually hit this problem and need to find out more, but that that's kind of that was one of the most satisfying ones because it was it was one of these ones, like you could tell with the the materialized view thing, I kind of I understood it more, but I never quite nailed it because it wasn't like this is like this is a piece of Oracle software, like you're never gonna quite get to the depths of it. Whereas with the Kafka thing, like it made sense, it wasn't broken, it was just like an ungetting that understanding. And then once you get the understanding, it like it all falls into place and it's completely logical. Um, and a ton of people have that problem and benefited from that that write-up. So that was that was a satisfying one to fix or to figure out.

SPEAKER_02

And it and it was not uh like you said, it was not the problem. It's just like um uh the bar uh to to to configure this was a little bit higher. So we as Endeavours we have to you know bring the people higher in order to uh make sure they they know how to fix this problem and know how what's what's the difference. Um developers of this uh the project Apache Kafka they were smart in order to you know come up with this interesting solution and and kind of like uh uh they knew that it's gonna be running in some weird situations like inside containers and there would be some uh networking situation that's going on. It's always you know networking situation going on with distributed systems. Um have you um have you encouraged any uh or encounter something similar in something interesting in the when you were um doing a lot of um the Kafka Connect uh evangelism? And like specifically many people, at least I'm remember from the 12 um 12 days of uh Kafka SMTs. Um so it's another interesting thing because it's it's there, it's available, but not many people knew this. And people usually come with these type of questions. I have a problem, like a problem like this, how how I can solve this, and they try to solve uh they're trying to look for for the questions, right? So they're trying to say how I can do blah with Kafka Connect, or how I can do blah with Kafka. And I remember that's kind of like one of the discussions that we we had when we were building um Kafka tutorials as a matter of fact. So and we also have this idea that okay, so what if the people would be you know googling particular solution, right? How to do this. Um can you um can you recall some of the interesting things that you did uh during that time?

SPEAKER_01

Yeah, so I think with Kafka Connect, kind of similar to the listeners thing, it was like this thing works, it's just isn't particularly well understood in the community, and that was with the um the converters. So this idea that Kafka messages are just bytes. Um and particularly coming from a database background, I that just hurt my head. Like this idea that you just store bytes, because like I'm like, well, where's my schema? Where's my kind of columns? And um and so the way that Kafka connects, and it's it's a very I think it's a very clever way it's being built because it separates all out all of these different considerations and it means it's completely pluggable.

SPEAKER_02

Yeah.

SPEAKER_01

Um, but this idea that you could take data from a database and store it in a Kafka message, and something needs to work out like how are you going to take those columns and write them as a bunch of bytes? Um, so do you store your data in like Avro or JSON or Protobuf, whatever? But then how would you handle a schema within that? Um, particularly people would use JSON, still use JSON, and they'll look at that and they'll eyeball it. It's like, well, it's got a schema. It's like, well, yes, but no, it hasn't actually. It's not got like an explicitly declared schema. Um so under understanding and like writing a lot and like lots of stuff on Stack Overflow and on blogs and helping people on the forums and the uh conference uh community Slack group. Um just like if you put this converter in place, if you tell Kafka Connect you're using this converter like to get data in, it's not gonna like change it for you. So like if you've written um sorry, you're getting data in into Kafka Connect, um, it's gonna go and write it out as JSON if you tell it to, but it's gonna throw the schema away. So then when you try and like plug a Kafka Connect sync into that, it doesn't actually have an explicit schema. So if you then want to write it to some kind of column of storage uh or tabular storage, it needs that schema from somewhere. So just those kind of making the things line up and making people understand if you've got Avro going in and you want JSON going out, how to actually configure it to do that and how to handle the schemas, and then everyone's favorite magic byte problem. Um the way that Kafka Connect is built is I think is is very, very good. I think some of the error messages are esoteric, are kind of like less not not so user-friendly. So things like when Kafka is a good word, yeah. Yeah. If you if you give Kafka Connect and you're building a sync and you want to get data from Kafka into some place like iceberg or whatever, if you've got data in a Kafka topic that's serialized in Avro using the confluence schema registry, it puts on the front of the message like these little magic bytes that tell it what the schema ID is or whatever. And unless you configure Kafka Connect correctly, you get this like unknown magic bytes. Um it's you get that if it's not Avro data and you try and read it with an Avro converter. The fact that even I'm getting mixed up with it, and I've written lots about this, goes to show it it's one of these things that it's arguably more difficult than it has to be, but it is very, very flexible. That's the uh the flip side of it.

SPEAKER_02

I think the you you're right in terms of uh in terms of flexibility of the things. Uh, when you're trying to do things uh like generic enough or like cover multiple uh coordinate cases, which um that's I think that that's the beauty of Kafka ambiguity and how it can be used, but also it's a it's a little bit of curse. Um I think in I I I don't wanna I don't wanna sound as a um kind of like a self-proclaiming um the you know the the the solver all of this problem but in that's what we do in the devrel. You know, we there's there's a product that does incredible things and we're trying to identify the points where we can show some of the things and explain so the people will have less less problems with uh with serialization. Um thankfully right now we do have enough materials to to cover this, so even LLMs uh will give you right results. So I think our job, our job was done uh very well here. Um and I think one of the things that I want to ask you uh before we kind of like erupt this, um, do you have any kind of uh maybe recommendation for people who trying to you know solve what your kind of the way approaching solving the hard problems? Um I think when with the even you gain certain experience of um of the things, you're not looking to the problems, oh like I don't know how to do this. You approaching this with slightly okay, so let me see what I know, how much I know uh for for this in order to approach this, and I'm going to this uh F um um F A F I type of approach uh that we uh that we usually uh employ. Uh unfortunately I cannot say this out loud, but uh uh I uh I YK IK.

SPEAKER_01

Indeed, indeed. Yeah, and when it comes to working with with hard things, it's kind of I think taking a methodical approach is uh essential. Um the kind of the there's that quote from Linus Torvalds, isn't there, about the kind of the random unjiggling of things until no random jiggling of things until they unbreak. Um which is just brilliant. It's like you can like muck around and find out, but in the end, you really have to like start with what you know, and and sometimes there's a bit of like random jiggling, but it's doing to in a kind of a constrained manner, and you kind of, well, I'd change this thing, what impact did that have? And then move on to the next thing, and not starting with any assumptions. Uh, and that's why I find writing stuff up so useful because that's how I tend to think is like, well, I'll write this thing down, and if I don't know an explanation of why I've written that down, well then I need to go and understand that thing.

SPEAKER_02

And or you know, this like when this loop could uh fully fully close, or someone will read your material and maybe you will help with someone, but there would be someone who knows more about the subject and they will reach out and help you out to saying, hey, I want to point out it's a good blog post. Here's one thing that I would uh add or change because this is not like in the in the very beginning, how you how you start talking about your uh materialized view problems, and uh someone show up and and uh also either shared your blog or read this and give it a give it a go.

SPEAKER_01

No, exactly. It's uh yeah, and that's one of the great things about sharing stuff publicly, whether it's kind of as part of our jobs in DaveRel or just like writing stuff stuff up on the side, it's uh it benefits everyone.

SPEAKER_02

All right, folks. Uh thank you, uh Robin, for being a part of Confluent Developer. Um, we we will put some some links in uh in show notes. Uh don't forget to follow Robin in uh social media. I guess right now it's mostly LinkedIn.

SPEAKER_01

Um, Blue Sky's very or Blue Sky.

SPEAKER_02

Blue Sky is not bad. Like there's some good communities, like the data engineering communities in Blue Sky. So uh please uh follow uh Robin in social media. Please follow Confluent Developer. Check out our YouTube channel with the interviews uh from the eliminarist and around the world who are talking about their experiences of uh building stuff or maybe breaking stuff. Uh, my name is Victor Gammoff, and as always, have a nice day.