Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov
Hi, we’re Tim Berglund, Adi Polak, and Viktor Gamov and we’re excited to bring you the Confluent Developer podcast (formerly “Streaming Audio.”) Our hand-crafted weekly episodes feature in-depth interviews with our community of software developers (actual human beings - not AI) talking about some of the most interesting challenges they’ve faced in their careers. We aim to explore the conditions that gave rise to each person’s technical hurdles, as well as how their experiences transformed their understanding and approach to building systems.
Whether you’re a seasoned open source data streaming engineer, or just someone who’s interested in learning more about Apache Kafka®, Apache Flink® and real-time data, we hope you’ll appreciate the stories, the discussion, and our effort to bring you a high-quality show worth your time.
Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov
Contributing to Open Source with the Kafka Connect MongoDB Sink ft. Hans-Peter Grahsl
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
Sink and source connectors are important for getting data in and out of Apache Kafka®. Tim Berglund invites Hans-Peter Grahsl (Technical Trainer and Software Engineer, Netconomy Software & Consulting GmbH) to share about his involvement in the Apache Kafka project, spanning from several conference contributions all the way to his open source community sink connector for MongoDB, now part of the official MongoDB Kafka connector code base.
Join us in this episode to learn what it’s like to be the only maintainer of a side project that’s been deployed into production by several companies!
EPISODE LINKS
- MongoDB Connector for Apache Kafka
- Getting Started with the MongoDB Connector for Apache Kafka and MongoDB
- Kafka Connect MongoDB Sink Community Connector
- Kafka Connect MongoDB Sink Community Connector (GitHub)
- Adventures of Lucy the Havapoo
- Join the Confluent Community Slack
SEASON 2
Hosted by Tim Berglund, Adi Polak and Viktor Gamov
Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
Music by Coastal Kites
Artwork by Phil Vo
- 🎧 Subscribe to Confluent Developer wherever you listen to podcasts.
- ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
- 👍 If you enjoyed this, please leave us a rating.
- 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
Now, Kafka Connect is awesome, but in the end, it's nothing without its connectors. Hans-Peter Grossel, a wonderfully active member of the Kafka community, stumbled upon a need for a MongoDB connector while doing some consulting work. To no one's surprise, he solved it by writing the code for that connector and later open sourcing it. We talk about what it's like to be a sole open source maintainer and how to build community around a project and also how the Kafka MongoDB integration actually works on today's episode of Streaming Audio, a podcast about Kafka, Confluent, and the cloud. Hello and welcome back to Streaming Audio, everyone. I'm your host, Tim Berglund, and I have with me today in the worldwide virtual studio uh my friend Hans-Peter Grasel. Hans Peter, welcome to the show.
SPEAKER_00Hi, Tim. Nice to be here today uh and uh have this discussion. Awesome.
SPEAKER_01Uh so just to tell everybody, I'm I'm recording from my home studio in suburban Denver, Colorado. Uh, where are you located, Hans Peter?
SPEAKER_00I'm actually located in in Graz, which is in Austria. So pretty much at the center of Europe, if you like.
SPEAKER_01Awesome. Uh and a beautiful country. I love uh Austria. I have not been enough, I think is the right way to put it. Um now I think there are probably a fair number of people um among our listeners who know who you are. I think you're a pretty prominent person in the Kafka community. But uh tell us about yourself. Like what uh what do you do now? How did you get started with Kafka? Give us a little bit of that story.
SPEAKER_00Okay, uh it's my pleasure. So uh well, what do I do now? I think I think most of the things that I'm doing these days are are somehow related to to training and education, which is probably not what uh what people would uh know me about if they know me at all. So I actually I don't think that I'm that prominent, but thanks for for uh thinking. So uh yeah, that's good. So uh and then again, I I think uh so I I'm a part-time technical trainer actually at the company here uh called Netconomy, where we are building large e-commerce solutions for our global uh uh customers, basically, really huge brands are among that as well. So I'm working there as a technical trainer where I focus on on back-end web development, mostly related to Java and Spring. And besides that, I'm I'm working as an independent consulting where I try to help clients with um, let's say, designing and implementing sometimes at least POCs uh for modern data architectures. And unsurprisingly, uh very often there is uh Apache Kafka in one way or another there, and uh, we are most of the time combining that with um non-relational operational data stores, if you like. Um, apart from that, sometimes I'm I'm I I do I do some some teaching at a smaller local university as an associate lecturer, also in the field of of software engineering. And uh unlikely, but still you may have seen one or two conference talks of mine. Uh so time permitting, I try to uh find some way to spread the word about about uh some cool tech technologies that that I get to work with. And yeah, here and there I also try to write down some uh thoughts in in blog posts.
SPEAKER_01That sounds like a delightful professional cocktail. There's a good mix of uh a lot of things that I personally love in there. So that's a that's a great mix.
SPEAKER_00Yeah, thanks. I think I think yeah, what what's what's uh important is that that I really like to do that because it's a lot of uh diff actually different uh things in detail, but like I said, most of the stuff is is still somehow related to to training and education, which is I think that the the reason why why all of this works. So there are, let's say, enough synergies for that not to fall apart. I mean for me as a person.
SPEAKER_01Yeah, no, that makes sense. And I I I uh also I just on a personal note, I'm right there with you on that. I've always said if you you know you took me and you put me in a test tube and boiled away everything non-essential, which is I guess kind of a weird analogy. But imagine if you did that. What would be left, I think, at the bottom would be a teacher. So I get that. That's a good that's a good uh combination. Uh you left out, I think, uh open source contributor, and you mentioned non-relational operational data stores. I I know you've done a lot of work with Mongo. So tell us how you got involved with MongoDB and why it and not another thing. Uh before you answer that, I I remember in 2010, 2011, I was doing a lot of conference speaking, and I had this kind of NoSQL survey talk uh that was popular among a certain a certain segment of the the developer population at that time, just because it was still fairly new and people trying to get a handle on Cassandra and Mongo and Redis and Neo4j and all these crazy new databases. What's the world like? And people would always ask me, hey, should I use Mongo or should I use Cassandra? Which in retrospect is you know a weird question because they're so different. But yeah, tell us about your involvement with Mongo. How did that go?
SPEAKER_00Yeah, so so so actually, I actually uh I I think if if I'm not mistaken, I I started to to to look into MongoDB when it was version 2.6. So I cannot really remember when that was, time time, timing-wise, but uh but a few years. Uh and uh the the reason was uh I I I found it interesting uh because I wanted to learn uh about uh the document-oriented data stores, and it was uh one of the data stores that at least had had had some visibility already back then, uh got some traction already in the in the developer community. And uh this is how I actually started to look into it. So more so it really started more with out of curiosity and and trying to uh to leverage something other than an RDBMS system, which which definitely is something that that was dominant in everything I did prior to that. So it it it all started basically as an experiment, as a personal experiment. And then I I think I had the luck that I found a small local company who who wanted to leverage it, at least uh start they started with with with a POC, and I tried to help them. And at the same time, I was I was learning on along the way, of course. And and that was the original involvement. And then I I yeah, then then I had some uh after that uh we we had a POC there, I think they continued, uh, and then I personally had some uh some uh non-Mongo DP time again, uh just because it it was um uh the projects were set and structured differently, so I I couldn't really leverage it for I think one year, and then I came back uh uh when when uh 3.something uh was released. Uh yeah, and then I continued again with some other client that that uh tried to leverage uh first of all the document model, and then they also evaluated a few things and then they settled with with MongoDB. Yeah, and and so I I somehow stick to it because it was the the first uh document-oriented database that I that I uh learned about and I that I I could apply also in in practice. And this is uh yeah, I I had no reason, uh at least not in this uh uh in terms of document-oriented databases to switch away. So I I basically stick to it over the years. Uh yeah, and and and later then, uh yeah, I think we we talk uh about that anyway. Uh I I I then started to uh to build something that helped uh uh with with integrating MongoDB as a data store uh with Kafka as a streaming platform.
SPEAKER_01Right. And that um there's there's no point in pretending that we don't know what that is. That's fairly obviously going to be a Kafka Connect connector. Exactly. That's that's how you do that. So I I don't mean to pretend that we're we're building up to some dramatic reveal of Hans Peter wrote a connector. But I do want to talk about um I don't know if you could talk about like the project where that first became obvious. Because if you know a little bit about the platform, you know, okay, here I have a database, I have Kafka, I want them to talk. We use Connect for that because it keeps us from reinventing about you know 73 little wheels that you know you need to make that kind of thing work. Um but what was the you know, kind of going into that project where that need became necessary? Uh walk us through that. What was it like as you were working on the POC, providing architectural guidance? What were the things that presented themselves to you that made you think number one, Kafka and Mongo had to talk? Yeah, uh, and number two, that they had to talk in a specific way.
SPEAKER_00Yeah, so uh let's start with with the first part, which is uh why uh why was there a need at all uh to uh to have uh a connection between Kafka and MongoDB? So the the interesting part is that uh that the project was already uh they they were already settled with MongoDB. And and what they basically did uh at the moment is that at that moment is that they uh ingested uh their data over over a web service directly into MongoDB. And then they had um they started to look into different use cases where where it made sense that um the same data that got ingested uh can also be leveraged by completely different uh parts of their landscape. So um given that um the next thought was well, of course, you could now try to um point these other systems or services, uh, if you want to call them microservices, feel free. Um then I uh so they they were in need of of the data, and and and then they had basically two uh two integration parts uh using MongoDB alone, so so they could either do uh you know polling kind of queries against the database, which is obviously not the best idea. And the other part is that uh it was um it was, I think, at pretty much around the time where where MongoDB uh announced to uh to uh the feature of change streams so that you can subscribe to uh to changes that are happening in the database and do something with these changes. Um that would be uh the second option, and and then uh well at that time I I at least had some uh knowledge built up uh around uh Kafka as a streaming platform, and then we thought, well, why why why not uh reverse all of this in in the sense that the services uh don't uh ingest anymore directly into MongoDB, uh but uh start to write uh all the incoming data first into Kafka, and then from there we continue to uh to feed other systems. And at that point uh in time, well, like you said, it was if if you want, it was during a consulting session with with one of these clients that I then said, well, so the question was then how how to get data from topics into MongoDB collections, and then I said, Well, yeah, you know, there is this uh great kind of uh framework uh called Kafka Connect, and and you can actually use it and you can ship your data from topics into the data store, which is MongoDB in your case. Then, of course, the next question is okay, uh is is there something that can do that out of the box? And then yeah, I I I wasn't actually sure. Then I started to uh look around a bit, and of course, there was something available uh at the at the time, but it uh first of all it it it didn't uh it was nothing official. Uh so uh they said, well, yeah, okay, so so these are now now some some some projects by done by someone pushed to GitHub, which actually we don't know. Uh there was nothing official. Um, and then well, I said, yeah, so so if you have a problem with that, um I actually wouldn't have had a huge problem with that, but I said, well, then there is a framework. Uh how hard can that be, right? Uh you can write your own connector, uh, sync connector, that is, and and yeah, this you can imagine that these were the words that that somehow got me involved because they first of all, I think they didn't have um the knowledge, and uh, even if they would have had uh most of the knowledge that they would need, they I think they were like most of the time in in in such projects under high pressure, and they didn't have have the resources to do that. And then they said, Well, can you do that for us at least uh in a way uh for a POC that we see such an architecture could work like it was designed during during a workshop. Um and then, well, what can you say? You can say you have basically two options. You say no, or then probably it's it's it's the end of the consulting uh uh session, uh at least when you look at it from the perspective of working with somebody at least uh longer than than one or two uh gigs. Um and then I said yes, I I can try that. And this is actually how I personally got involved into uh this whole thing, where I then uh wrote the first uh version of the sync connector. I mean, I I wouldn't even call it first version, really, it was it was just enough uh to uh to see that uh what we discussed uh it works in principle and and uh can be uh one one uh a better approach than than the others that we that we dropped uh for for obvious reasons. Yeah, and then it was there. Uh they were to some degree happy. And then um it was not so much about them that uh that asked me to do uh a lot more. Uh it was then really my my own interest uh and I and I wanted to improve it because uh like I said, there were other connectors out there which uh could do more or less uh or at least the same, maybe even uh even a bit more in certain areas. But uh still I uh looking at at the code and everything, I I was not uh convinced that uh this is something that I I personally want to work with for some reasons of, for example, not having um a reasonable coverage, uh talking about tests or not having uh enough uh kind of documentation. I was also lacking uh looking at the code base, things like extension points, uh what how can you uh implement uh different behaviors, different features for the connector? And and and and that were the main reasons why I thought, well, I I think I can do better, and I had some some some good ideas. Uh and then yeah, I I actually by accident I I I continued to do that. Um yeah, mostly because of intrinsic motivation, and and and yeah, and then I I managed to stick to it um for some reason and it evolved over time, became better and better, and then uh went on to uh even even bigger things that we'll we'll talk about.
SPEAKER_01I want to say your your story of that meeting with that client reminds me of a UPS commercial that aired in the US maybe like 15 years ago or so, but there were these two consultants sitting in front of a CEO and you know giving him their conclusions about how he should modernize his supply chains and you know do all this stuff and using some nice abstract business language. And the CEO is very pleased with this. He goes, Great, do it. And they look at him like like all confused, they go, Wait, we don't do the things we recommend, we only recommend them.
SPEAKER_00Yeah, exactly.
SPEAKER_01And then at the end of the commercial, they're going down the elevator and one says, You believe that guy? So, right, as a consultant, which you have been and which you are, you know, that that's less of an option, right? You of course you you said yes, and I'm uh we're all glad you did. Um but I also I want to talk about you said that there were connectors out there, and like the specifics don't matter, but you weren't happy with them. You just it was your judgment that they were not a good way to proceed. And this is um this is an important thing about Connect. It's an important thing about open source, and the kind of ecosystem that Connect fosters uh brings this to light in an important way. So let me let me uh get my thoughts out. You can tell me what you think. Uh so Connect, uh, if you're listening and you're brand new to the Kafka ecosystem, it's a data integration framework. It's a process that runs external to Kafka brokers, uh, you know, from the standpoint of the Kafka cluster, it's a client, like a producer or consumer, and you have these pluggable connectors that are jar files at the most concrete level. Um and you configure them and they come to life and they read data from some other system or read from Kafka and write to some other system. So it's this integration framework. And so, obviously enough, that creates an ecosystem of connectors. And there are closed source connectors, there are commercial connectors, there are open source connectors, there are all kinds of connectors. And among the open source ones, like you can go search on GitHub. I want to connect to, you know, you wanted to connect to Mongo, and you found a few options. And I always make this point that it's the the whole point of using Connect is that you don't write that integration code yourself. You use a connector because it's the same integration code everybody else is gonna write, and it's not uh adding value to you or your company to do that yourself. It's not differentiated, right? But when there are all these open source options, this is you know, we used to talk a lot uh uh early in the 2000s, late in the 90s about uh open source being free as in speech or free as in beer. Yeah um well uh sometimes some open source projects, and I think long tail connectors in particular, are free as in puppy. Uh and speaking as a man whose family just got a puppy uh a couple months ago. Puppies are a lot of work, right? They're they're you gotta feed them more often, you gotta clean up after them, they wake up in the middle of the night, they bark, chew things you're not supposed to chew, you know, which is a lot like a long tail connector that somebody wrote and it did what they needed, and maybe it's still doing what they need, but they're not maintaining it, and it doesn't have any kind of community around it, it's just this repo out there. Yeah, and some of those are like you you do want to adopt the puppy. It's it's it's worthwhile to bring it home and housebreak it and give it obedience training, and it'll be you know a part of the family for a long time. But sometimes that's not what you want, and so you gotta be careful and realize that you're adopting a puppy um when you take one of those on.
SPEAKER_00You didn't know. Yeah, I feel with you. I feel with you totally, but I I I cannot really uh uh relate personally because luckily I I I have no no no such thing as a puppy at home. That's that's the good part of it.
SPEAKER_01Lucy is a good girl. I should be on record. Uh if uh maybe I'll I'll uh I'll tweet a picture of her or put a picture of her in the show notes when uh we publish this just so she can have absolutely yeah I I I I will give you at least one like, if not every tweet. Okay, thank you. Thank you, sir. So anyway, uh I think you encountered the Free As and Puppy phenomenon for open source uh Mongo connectors, and that's a thing to be aware of. And you went on the journey of writing your own, and then that in that engagement it actually got put into production. Is that correct?
SPEAKER_00Yes, exactly. So it it got put into production, uh not not uh by me directly. So of course I'm not running uh anything on that connector in production on my own, but there were uh there were two early adopters. Yeah, yeah, it there were two early adopters, actually, one uh coming from I I mean both both these early adopters, if you want, were uh are mentioned in the README, so they are in the talks. So it's actually two companies, one which is in Germany, um uh and another one which which is uh surprisingly from from the US. So uh they are listed there and they were quite happy. Uh, I think that again I'm not 100% sure, but I think it was somewhere uh at around one year after I started it or something like that. So yeah, so so then there were real customers, if you want, for for this uh connector. Um and then yeah, they they got in touch, they they discussed a few things, they they of course wanted to uh help improve it in the sense that they suggested um some features that that could get implemented, um actually some of which did. Um yeah, so so this is and then um starting from that point on, uh yeah, I think it gained at least some traction in its, if you want, in its own niche. And yeah, then I I I think I I then tried to get in touch with with you guys and and and other. asked if if if it could be um so back then it was not even the the hub where you can find this curated list uh really nice one uh about all the oh most of the uh popular connectors out there uh it was it was just at the connectors page back then and and there it got listed so at least people could easily find it from your uh from your list from your from from from your web page for that and then later yeah later you you uh released the hub and and then i i again got in touch and and and and wanted to have it uh listed on the hub also to yeah because i was convinced knowing from from the other two companies that it that it works out quite well in in production so um i i thought well everything that i can do to make it more visible uh might be of help to others uh and and and i think that putting it on the hub was actually the next uh step in terms of um getting more visibility for for for this still uh very very small uh side project of mine yeah yeah um so what was it like uh getting it listed on connect hub walk us through that process from the standpoint of open source maintainer um yeah so so so the first time i i i got it um uh to the hub uh I think it was uh yeah that there was first of all there was no um no build plugin or something which you provided later so um or it was in a very early version I got in touch with one of the maintainers or one of the responsibles from from from your side I think it was uh uh Chris uh and he basically um helped me also to uh uh to understand uh and and have an easy time to just modify uh the build a bit so that I get a package uh structure or actually an archive structure that I then send basically uh manually to you and and you put it on the hub so this was the the the very first uh time uh it got listed there and it still it it it required some manual effort but it was it was very straightforward after all and and and you were very uh first of all cooperative in in in in helping uh along uh this process um and yeah so it was pretty smooth uh to get it there uh and people could could find it and use it so um yeah I think from from from that perspective of of um getting it there it there's there's nothing really nothing to to complain about.
SPEAKER_01Nice nice good I'm I'm glad and I know you're on our podcast and you're gentlemen and you you might not complain if you if there were but it's still good to hear that from the outside that process is smooth and that's exactly how I understand it to be so yeah cool um and it is absolutely on connect hub um but what um and the the the story from there there there are two interest interesting things that happened or that I would like to dig into and one is um you know going back to this free as in puppy um at some point when you're maintaining uh as a one person show an open source project it's like a puppy as a service you know everybody gets to competit and um and enjoy its cuteness but they aren't the ones cleaning up after it or training it or having it bark in the middle of the night or anything like that. You are that the the sole maintainer of the thing and you know people have expectations.
SPEAKER_00So what what was that like as as it became popular um I'm assuming there was more adoption and uh what's it like what was it like for you to be a one person show maintaining a project like that um yeah I think you've got a very very important uh point here that that this is uh let's say that if you look at it from from over a longer perspective I think it's it's it's a lot of ups and downs honestly so it's of course there are let's say um some um rewarding aspects to it uh being being that the one man show the one let's say writer and and and contributor to to to the project code wise this is so so you know that that every time you do something you you were the one who who made something possible or or you were the one who could help someone um this is the good part of it that the then there are let's say not so nice parts of it like people um opening let's say issues in a in a in a non-polite way that that can happen or or people who are only yeah then then then people who are like only uh informing you what what what's missing like uh never really uh taking not at least uh two or three words to to thank you what's there and what's working but only you know um nitpicking about the things that they would like to see and the things that are probably not working as they think they should um that's that that's the hard part because that you you have no one literally no one uh who who you can let's say who who you can discuss that with who who you can share these feelings with uh so again emotionally I think there's a lot of ups and downs uh involved in that uh and I think that the it's I mean I I think I was lucky to one degree because um like I said that after some time I think again a year or so there were customers which is like like some very important let's say milestone from from that perspective of of of getting something back like a reward if you want uh it of course this has nothing to do with with with with money uh but still it's it's it's something where where you are happy about and where and and these are the moments I guess that that you need at least from time to time uh to be able to uh to stick to it and to be able to to work on it and and continue that project. I think if this is really vital uh if if you never get any kind of reward uh in in in what form uh in or in any form uh again I'm not talking about money here then uh you can get frustrated over time I think you might even suffer from I don't I don't know maybe even even burnout kind of kind of uh um symptoms I I don't know I like I I I think I was lucky that uh that I I didn't experience it in in in that um uh in that way but uh yeah I I think that can happen easily uh uh when you are doing all of that on your own. Uh another thing is that people even if documentation is there uh obviously don't like to read this documentation. So on the one hand we all know that people don't uh or developers sometimes are not particularly thankful if they should write documentation but even if you do and even if you do that in your free time and you try to have some kind of reasonable documentation people are sometimes not reading it. And then you it basically what happens is that you get issues which are more questions or invalid ones that could be easily so solved by themselves by spending five minutes in in in the README or something like that. This is also something that can that to some degree it's funny but again it's it's it's a lot of personal attitude how you see such uh events if if if you find it funny if you try to be helpful uh even in that situation or if you see that as as as a another frustrating moment. I I guess a lot of is a lot of or what's important is to not to uh to let it into your head too much uh what happens on on the issue tab in your project. That's that's important. I think you you should find a way to um keep some distance there.
SPEAKER_01Yeah yeah no all that is great and I really wanted to talk about this because um this is a side of open source that we don't always think about like if you haven't been a maintainer um this is not going to occur to you um and you know we have the the well-known problem of of sometimes people treat open source projects like public utilities you know like this is it's free therefore I have a right to it therefore I have a right to influence its backlog and demand fixes and you know communicate with the maintainer in a brusque matter of fact or even rude tone. Like that happens and it kind of depends on the maintainer's personality. You know there are people who just don't care and are going to be unaffected by that. You're not that kind of guy I happen to know I'm not that kind of guy. You know that kind of thing uh is very taxing on me. So um it it's it's just a good reminder that if you use a a small open source project that has one or few maintainers um it's a good idea to remember that like you are receiving the gift of that person's time. Like you said you had a motivation uh you were trying to solve a problem for a client and your continued investment in it, you know maybe you're motivated by this is uh marketing for future consulting or something. You know there's there's some angle that you have uh but uh still you know the the user who's not contributing is you you are giving that person a gift of your time and you should treat open source maintainers like people who are giving you a gift. Now big projects that have big companies behind them. I think you know everybody still has an obligation to to be polite and to generally love their neighbor um but the economics of that are different if you're if you're dealing with something large. But a small thing like the Mongo connector, you know, there's one guy look do it do a get blame on any file you want. You see you see one name behind all those commits uh and you should treat treat the person like somebody who gave you something and you know you can report bugs you can ask questions you can have opinions but do it in a polite way. And if you're so inclined, you know it would be nice to occasionally file an issue that just says hey this worked and it's wonderful and they know that issue will get closed but uh there are again just depends on who the maintainer is there are people who thrive on positive feedback and you might consider that if you're a beneficiary of a project like this to to uh do that.
SPEAKER_00Uh yeah uh be nice to open source contributors yes that's definitely a plus and it helps uh keeping motivation high.
SPEAKER_01Right right you'll keep getting the thing that you like and they won't go away and stop doing it because chances are they probably have if there's somebody who could write a MongoDB connector uh they probably have uh relatively high value opportunity costs that they could fall back to with their time and go do yeah instead of doing that thing.
SPEAKER_00Yeah and actually what what's also interesting is the fact that uh we we always said from from from like from the beginning that we and are now discussing this story that there was nothing official and actually that was also uh let's say a a a limiting factor to some degree for myself because I always thought is it even worth um putting some effort into implementing yet another feature because the day has to come uh where something official will get announced and will get released and then we all know that uh it's basically means uh the end of of a one-man show project most likely because you can never really uh be competitive as as as soon as some big company uh starts to to invest in that and and allocate resources for that and and that was uh something that that I was personally struggling with a bit um at least sometimes along the way does it make sense to continue because I'm sure that something will come in the next one or two months. And then for quite some time uh surprisingly for quite some time nothing came. So I I I was lucky that that yeah that let's say the project uh lived on uh even if I was still uh or still am today for this community project that if you want it the only uh one maintainer cool um what uh in the process of building it uh we I think we've been talking about you know frankly the personal side of of writing open source were there any interesting technical hurdles in the Kinect API or in Mongo and just making it work um from 10,000 feet database sync connector once you know what connect is and if you know what a database is it's very easy to understand what it does but the devil's in the details and were there any details that were particularly devilish? Actually it was it was relatively smooth but I I have to say that uh that was definitely what what what was very very helpful to to to get started was the um I I'm not sure if it still exists in the same form but back then you you published this kind of uh connect developer guide or guideline I think um this this was a PDF basically about I don't know if I remember correctly something like 10 to 15 pages which uh first of all explained how you should approach uh writing uh a new connector so that that that was very helpful to to understand um uh how to get started quickly then again I think um the talking about uh another aspect is that when when you really implement the the sync connector part so getting data out of Kafka into some some data store I think honestly in in in more or less all cases writing sync connectors is uh probably easier by a large part than writing uh source connectors uh for several reasons uh so it's it's just it's just much much easier to to to to to write the sync than a source uh and this was also I think one of the reasons why why this this uh one man side project of mine um never really uh tried to provide a source connector first of all there there is uh there has been a very good one uh out there already uh baked uh baked by by by Red Hat we all know this great project this family of connectors called Divisium uh and it it it it just didn't make any sense for me to uh to try uh and write uh the harder part which is definitely the source connector so I stick to the sync side and and then I mean talking about any any any devil uh is in the details kind of uh of moments I I I think there weren't many honestly uh at least nothing that that I can remember now uh from the top of my head where I was really frustrated so working with with with the connect API and the framework is is uh really straightforward what the challenges are again uh there in the sense that you need to find a way let's say do you want to support uh upsert semantics towards the sync do you want to how do you uh want to commit uh consumed offsets uh but but that's most of the time that's that's that's some decisions that you take um and you you decide to go this or that way and you can change that and enhance that uh later on as you go uh and and actually that is what I did. So first of all uh I started with just bringing data over there with upset semantics um and nothing else uh and and that what that was a good starting point. And later I I I added things like native support for Debezium's change data capture format so that you can uh leverage the sync connector together with Debesium source and basically replicate uh MongoDB instances over over Kafka at least uh from a data perspective yep uh that is something that I added later which which which again just needs some time but it's it's not so so challenging. So honestly most of the things are are absolutely not not rocket science. All the heavy lifting seriously is is is done by by Connect uh under the covers and that's a good thing because yeah you you have uh really really great uh you know distributed systems engineers and and and they know how how how to do that uh in a in a reliable and distributed way. So yeah again as a as a as a connect implementer um it's pretty pretty smooth.
SPEAKER_01Cool. That is good to hear because that again is um that's my view of connect and this is how I always make the case for connect when I'm explaining the platform to people. Because some some people will actually ask me hey why why wouldn't I just write that on my own? Why use a framework? And I'm like you haven't been doing this for very long have you um because it's it's it's a deceptively simple problem right I'm just gonna subscribe to a topic consume messages from it and write them to a Mongo collection. Come on how that takes me two hours to write that code um and and you think that but then scaling that connector fault tolerance every other of the 308 corner cases that are going to present themselves to you are actually a lot of work and having the framework handle all of well I mean like I said before that the purpose of a connector is that it is non-differentiated code. Your syncing of Kafka data into Mongo is exactly the same as everybody else's syncing of Kafka data into Mongo so you shouldn't write that. But then you zoom into you and here you have Hans Peter writing connector code so that it can be written once and everybody can use that undifferentiated code. Well from your standpoint as the connector writer fault tolerance scalability offset syncing all that stuff the the distributed systems and framework elements of Kafka Connect itself that's your undifferentiated code that would be foolish for you to write along with every other connector author. So connect is in place to be that and it's nice to hear you know the first person report of a of a connector writer who says yep that that stuff actually works the way it's supposed to yeah definitely. Yeah because we all think it does and it's nice to know that it really does. Cool. Now uh and and things happened from there it got some attention and uh eventually the folks at Mongo came to be aware of it.
SPEAKER_00Tell us that story yeah so that was that was funny. Like I like I said I had here and there I I always thought well uh something must be done uh in that regard from from from the official side like like I was really really surprised uh not to see that that addressed by them on the on the other hand uh the uh yeah you could say well there was something in that case my project which which was obviously good enough for some time even for for some of the customers to use most likely so some of which I I I definitely do not know about um because not everyone like I said uh got got got in touch with me um they some were just using it so and then uh they they approached me I think it was uh in in late 2018 so December 2018 or something and and and and then yeah they they they actually got in touch first by email then um then uh we set up uh a first uh uh webcast uh then we yeah and and and then I was really really happy that first of all they already knew uh a lot about uh the project they they they were also happy uh the about the way it it it was done I mean from a let's say from a yeah how how the code basically was was designed uh leveraging the framework so it was like like I said stuff that that's not there in some other connectors uh is uh were were considered when when when I wrote it like like extensibility um customizability and and stuff like that a lot of uh things regarding configuration was was there already and they knew their desk coverage was good documentation was there so basically yeah they were happy and they were then asking if we can cooperate somehow because they they were wondering how how they they they actually want to address that topic On their own. The time was there for them to start doing something about not having something official in their product and services list. And yeah, this is how it all started. Then we had a few more meetings with some engineers, and then we basically discussed how we can proceed in making the code base and changing it towards the official connector. Of course, a lot was done then from their side with regard to repackaging, making sure there are coding guidelines, static analysis checks, and stuff like that. So everything that a more professional project would need anyway. Like I said, they were happy with with how the code looked looked uh back then in the in the community repo of mine. And yeah, it was pretty smooth. And they were really, really um uh cooperative. Um and I think it was um honestly a clear win-win situation also for myself, because uh I knew, and I think this is not something that that happens uh often, or let's say um from my own perspective, I think this is this is something that happens probably once in a lifetime, that somebody comes along and and and really is happy about a side project of yours and and wants to uh or is convinced uh and confident that that it is um it makes sense to to leverage that code instead of uh starting to write something from scratch on their own. And this is uh, like I said, one of these really, really rewarding moments that you yeah, that that that I was just lucky to to have experience because I think a lot of people are doing much, much uh better work than I did in in different areas and and and and never were lucky enough that that that somebody came along and and and actually um wanted that that the code base lives on in one way or another. Uh so and yeah, it was it was uh just so a lot of emotions, but uh definitely everything was was was positive. So I was never I never felt uh bad about anything that that happened since we we got in touch and since we cooperated on that.
SPEAKER_01Wonderful. And that really is a great story. I mean that's uh that's uh you you did some work for consulting project, you invested a lot of yourself into it, and now uh it's it's the sync connector. You know, we had the source connector, um, like you said, that's that's a a different ball of wax, and that's Tibesium, is what people usually use for that. Uh and now this is the official way to get data from Kafka into Mongo. Um exactly. Which is super, and I think it's uh I think it's a great contribution. What's next for you? Is there anything you can talk about that you're working on now or excited about in the near future?
SPEAKER_00Oh yeah, so uh there actually there's a lot. So uh the problem is my my backlog is really huge, my personal one. So uh I I have to to find some time during my vacation to to to think about it, how how I want to to proceed. I mean, there's different things I have. On the one hand, I I'm thinking about uh new features for for the connector, obviously features that I would want to contribute to to the official one right now. Uh maybe trying uh how how they work and if they make sense in my community repo before, uh, and then maybe contributing this this stuff to the official connector. That is one one thing uh in that regard. On the other hand, I I I want to to uh have some some topics in mind with regard with regard to um uh leveraging uh Kafka streams together with with Connect for doing certain things, like implementing certain data integration patterns uh where you would probably need Kafka streams or or combine that with KSQL because it it's doing things that that cannot or should not be done in that way with Kafka Connect. Then I have some ideas about writing blog posts. Uh I I wrote a few of them over the last couple of months, but I have uh a few more ideas to that I want to to externalize, let's put it like that. And and maybe I I managed to find some time to write some blog posts. So yeah, so basically a lot of stuff that uh that needs priorization, uh but maybe uh to to your delight, some of the things, or most of the things actually are are related to Kafka.
SPEAKER_01Well, that does indeed make me happy. My guest today has been Hans-Peter Grassel. Hans Peter, thanks for being a part of Streaming Audio.
SPEAKER_00Really, team, thanks for having me, and yeah, hope to hear you soon.
SPEAKER_01Hey, you know what you get for listening to the end? A Kafka Summit discount code. Kafka Summit is coming up on September 30th and October 1st in downtown San Francisco, and you can get 30% off if you go to Kafka-summit.org and use the discount code AUDIONETINE during checkout. Just enter Audio19 while registering at Kafka-summit.org, and that 30% off is all yours. I'd love to see you there. But hey, I hope this podcast was helpful to you. If you want to discuss it or ask a question, you can always reach out to me at TLberglund on Twitter. That's T-L-B-E-R-G-L-U-N-D. Or you can leave a comment on a YouTube video or reach out in our community Slack. There's a Slack signup link in the show notes if you want to register there. And while you're at it, please subscribe to our YouTube channel and to this podcast wherever fine podcasts are sold. And if you subscribe through iTunes, be sure to leave us a review there. That helps other people discover the podcast, which is a good thing. Thanks for your support, and we'll see you next time.