Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov
Hi, we’re Tim Berglund, Adi Polak, and Viktor Gamov and we’re excited to bring you the Confluent Developer podcast (formerly “Streaming Audio.”) Our hand-crafted weekly episodes feature in-depth interviews with our community of software developers (actual human beings - not AI) talking about some of the most interesting challenges they’ve faced in their careers. We aim to explore the conditions that gave rise to each person’s technical hurdles, as well as how their experiences transformed their understanding and approach to building systems.
Whether you’re a seasoned open source data streaming engineer, or just someone who’s interested in learning more about Apache Kafka®, Apache Flink® and real-time data, we hope you’ll appreciate the stories, the discussion, and our effort to bring you a high-quality show worth your time.
Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov
Fail Fast & Ship It with Jeremy Custenborder | Ep. 18
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
Viktor Gamov talks to Jeremy Custenborder (Confluent) about his career in large-scale systems. Jeremy’s first job: paper boy. His challenge: keeping MySpace running at a massive pre-cloud scale while building the tools that didn’t exist yet and learning to fail fast.
SEASON 2
Hosted by Tim Berglund, Adi Polak and Viktor Gamov
Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
Music by Coastal Kites
Artwork by Phil Vo
- 🎧 Subscribe to Confluent Developer wherever you listen to podcasts.
- ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
- 👍 If you enjoyed this, please leave us a rating.
- 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
This episode takes us back in time into MySpace era. Times where if you need a tool, you had to build this yourself. This is Confluent Developer.
SPEAKER_02That was one of those spots where it didn't exist yet, so we built our own. The biggest thing that that taught me was you don't have to get everything right the first time. Fail fast, figure out what's uh what your problem is, and then keep moving on from there.
SPEAKER_01Welcome back to this episode of Confluent Developer Podcast. My name is Viktor Gamoff, and I will be your host today. And today I have a very special guest. Every guest in this show is special, but this one is truly something. And uh you will understand by the end of this episode why I'm saying this. Jeremy Custom Border, the legend of Kafka Connect. And uh we're gonna talk about this. Jeremy, welcome to Kampf and Development. Thanks for having me. Thanks for having me. It's great to uh finally uh reconnect after you know many years that I haven't seen you in person. I've seen you um, you know, all over like internet and LinkedIn and other places, but it's finally good to be with you on the same room and talk about some cool tech stuff.
SPEAKER_02Of course, my I'm Mr. Man.
SPEAKER_01Yeah. Uh for those of you, uh those of you who uh don't know uh Jeremy, can you tell us like in a few words what you do um and like how people you know can uh can know you?
SPEAKER_02Yeah, so um I'm an I'm an engineer like a lot of a lot of y'all. Um I worked at Confluent. I I joined uh originally joined Confluent in in 2016 and worked till uh 2022. Um I um I'm the type type of person that I'd say um I I worked in the field and that worked out pretty well for me because uh I have a short attention span. Yeah. So it's like I go from problem to problem to problem and and I enjoyed that. I had a I had a great time working here, and it's it's fun to be back chatting with y'all.
SPEAKER_01Yeah, and um Conflict Developer Podcast, this is the show where we explore the origins of this uh technical excellence and how um the luminaries or some like a visionaries and some some very cool people and some awesome nerds are came came to be, and like uh what um you know how they started and how they end up and the things. And one of the first questions that I like to ask people is uh what was your first job? Not necessarily in related to technology, not necessarily relating to software. What was your first time when you get this kind of like a crispy uh like a bill or maybe like a few coins and stuff like that? Depends when you started and what was the rate of inflation that time.
SPEAKER_02Well, I mean, my my absolute first job was um I was a paper boy from 10 to 16. Nice. And like I had that I had that down pr pretty well. And that was a great job. And I made um a hundred bucks a month. And like that money went in and then went immediately back out.
SPEAKER_01Yeah, you probably could afford you lots of uh candies and some some other interesting entertainment.
SPEAKER_02Exactly. But I mean like that was uh that would be the 90s.
SPEAKER_01So you know, you can go, yeah, you can you can do a lot with a hundred bucks a month in nineties, I guess.
SPEAKER_02And then and then my first uh my first real IT job, I I used to be the network administrator for the um the Green County Library System in uh in Ohio.
SPEAKER_01Okay, so tell us about this. What year it was, what was the network landscape at the time, like what kind of like hardware, software things do you're working with? Tell us everything.
SPEAKER_02Oh, it was it was a blast. It was so what I loved about this organization is um I was one of the first people I think that was that was nice. Okay. And it's like, you know, you're working for with for a bunch of grandmothers. Yes. And you know, they would be freaked out if they broke a computer. And you know, I'm in this phase in my life, and you know, this is like 98, so I'm I'm 18 years old. Yeah. And um, I didn't go to college, I'm self-taught. So, you know, I'm I'm I'm in this like like pre playground that had a decent budget. I had 400 machines that I that I uh administrated with you know T1 connections between buildings, and I had uh, you know, I had I had a decent budget and it was like a great place to learn because was it like a Windows machines, Linux machines? All Windows machines. Windows machines. Yeah, and so it it was it was great from the perspective of of that things weren't perfect, so that you could make a lot of improvements. So and then the other part with that is you're working for all these these grandmas. Yes. And you know, like so it was great for me because you know, I wasn't living on my own at 18, and you know, it's not like that job paid that great. Yeah. And you know, you so I would I would I would say, hey, I'm gonna be coming in to work on on y'all's computers on you know, Tuesday. I show up and there's all these baked goods. It was amazing. It was amazing.
SPEAKER_01So you're working for all these grandmothers. So you you have you have like incentives, maybe not uh with the monitor, but at least you will be well fed, and you had the your you know, the dose of sugar that was, you know, the the homemade cookies the best.
SPEAKER_02Yes. And the the cool part about that is like I could uh I had enough time where I could I could work and learn learn stuff, and that's where I I wrote my first programs.
SPEAKER_01What was the language you you used first time?
SPEAKER_02Uh first time was was ASP.net. Okay. So the the first portion of my career was was pretty heavily uh Windows-based. Yeah, yeah, me too.
SPEAKER_01I I also um after I get my first job in you know the Russian like bank called Zburbank, um I started learning. Yes. I I learned uh.NET and uh uh surprisingly not to not to um I didn't go uh PHP route because lots of uh folks from my school went to PHP route. I went actually SP.net as well. Yeah. It was version uh.NET uh.NET 2.0. So it's just it just came out.
SPEAKER_02I spent a lot of time with that later on. Yeah. So I I I did that for a while and uh I wrote um we basically had these CD ROMs with uh um uh like genealogy things on it. So you could come in and do research on genealogy in one group. Uh-huh. So I I put something together to help searching and help kind of find uh which CDs had what for folks.
SPEAKER_01Nice.
SPEAKER_02And that was my that was the first thing I ever I ever wrote.
SPEAKER_01Yeah. So basically you're uh the classical three-tier application, right? You have a you have uh like a front end, you have a like application server, which is probably what internet information server, right? And a database.
SPEAKER_02Yeah, well, I mean, man, you know, it's one of those deals in those days that didn't compile, where I would, I would say just it compiles, ship it.
SPEAKER_01That's that's what's pretty cool. My uh my first um uh.NET application was uh first complex application. It was attempt to implement kind of like a trading system. Kind of I thought like, okay, so I was working in the bank as a as my uh internship uh program uh during the summertime. I was like, okay, so what about kind of like I write the training application also compiled, uh and uh but my demo uh for uh when I was went to show this to to this like a committee of the professors and stuff like that, um it only ran on my laptop, obviously. So runs on my laptop, so that's that's how and uh also I and now we know where we got Docker from. Yeah, exactly, exactly. Um so so you started the uh you already mentioned that you self-taught, you learned uh these technologies um uh basically on your own uh time on your own pace, and looks like you also very was very interested. One of the one of the things what we like to discover uh discuss is uh the the biggest professional challenge. Like I um I know you from probably it was you know you can think about this. It's it's really was a complex challenge to find the ways to bring the virtually any data source in the world into Kafka. Uh and we will talk we will talk about this. I really want to talk about this, but on your opinion, something that you would be super proud or super embarrassed if it's like some of the complex tasks that you solved, like the the biggest challenge of your life like today, or something that you kind of like underestimated your your your your your power and it's like uh yeah it didn't work out, but still I learned a lot. So that's that's what I want to talk about.
SPEAKER_02Yeah, yeah. So if I had to go there, I will tell you probably probably the place I learned the most, and uh this is gonna age me. This is definitely gonna age me. Um so you know, I I I was around Ohio for a while, and then you know, just at some point I decided like I gotta get out of Ohio. And so like I I went to uh I I went out west and then I ended up in LA and I ended up working uh for for MySpace, the social network. And uh I I did that for five years. So um working.
SPEAKER_01For those of you kids who don't know what the MySpace is, is something like uh okay, Facebook also ages. Those kids are not using Facebook these days, it's probably using TikTok or something like that. But before Facebook was a thing, there was a MySpace. Exactly, exactly.
SPEAKER_02And I I uh worked on that. So and the the fun part about that is like you know, when you talk to folks, they would say, like, okay, we needed to have a caching tier. Yeah. They're like, well, why didn't you use Redis? Yeah. It's like go look at the first commits on Redis. That was like five years after. I think it was I I think Redis was was built in 2010. I I'd have to go look.
SPEAKER_01You know, but before that, there was a memcache.
SPEAKER_02We had one in yeah. Uh Memcache came out in like 2007 or 2000.
SPEAKER_01I think it's even came up for you know this type of like applications, they start popping up the like initial wave of web 2.0, kind of like what we know today as uh web scale things. Um and I think the um yeah, memcache was also around that time, around the 2000s.
SPEAKER_02Yeah, yeah, but that was one of those spots where it didn't exist yet. So we we built our own and we were we were uh running monstrous scale on Windows machines, which uh and and actually making great numbers. So uh you were talking about your ASP.net 2, you know, 2.0.
SPEAKER_01Yeah.
SPEAKER_02That was mine. That's what we were working on. It was the the whole web front end was ASP.net 2.0.
SPEAKER_01What was the uh I guess right now probably it would be like a very like small numbers, but the dead time, what was the profile of load? Like how many um daily users of the of the platform?
SPEAKER_00Now a quick word from our sponsor. Confluent Developer the Podcast is brought to you by Confluent Developer the website, which has everything you need as a developer of data streaming systems. And it's completely free. We've got curriculum, hands-on exercises, executable tutorials, the online data streaming engineer certification, also free, a way to find a meetup near you, those are free. Everything is there. I really want you to be successful in your journey as a data streaming engineer, and this is the site that has what you need. Check it out at developer.confluent.io. That's developer.confluent.io. Now back to the show.
SPEAKER_02Well, when I when I joined, uh when I joined, they had about uh uh 10 million users and were was pushing about uh I want to say at that time, about two or three uh gigabits a second uh from their data centers. And then when I left, um we had 320 million, if I remember right, users. And uh they were pushing a terabit from their data centers.
SPEAKER_01So what was the again about uh we leave in 2025, we have a great deal of um open source technologies available for building stuff because of this uh era of uh web scale. But uh if we you know back in time, what was the like biggest challenges apart from what we already discussed with the uh with caching, probably like a load balancing would be so much charting of databases and stuff?
SPEAKER_02So much has changed, and just so like you know, like for example, like everybody uses Kubernetes now. There was nothing like that. And then you know, Kubernetes can do like a slow rollout and handle all the health checks, and we had to build everything around that. So, like all of these were was stitching things like net scalers and other uh components together to get that type of functionality. And you know, today I look at what you have and it's like it's it's amazing. Yeah. And then, you know, like we were moving files around, and that's how you would do deployments. And you know, to you would you would image machines by setting up monstrous multicast groups and sending tons of image traffic down to to to machines, and that's how we like a way that's how we are uh vs or something like that, right?
SPEAKER_01Or not even that.
SPEAKER_02No, it's all all but all bare metal. Oh wow, yeah. And then like at the end, it was uh like right around when AWS was was starting to come up. We were also running Zen as well and and uh virtualizing it's instead so instead of patching, we would just swap out machines.
SPEAKER_01Yeah. Wow, that's uh that was pretty cool time. Um so what this uh this time uh taught you about data?
SPEAKER_02That's that's the question that I'm trying to slowly slide in to the I mean I'll tell you the the the the the biggest thing that that taught me was you don't have to get everything right the first time. It's like so you you need to obviously you need to to have good principles around around storing your data, but you can you can increment and you can you can add on and you can you can do more at a later time. It doesn't, you know, but for you know, I'd say that's right around the time agile development start has started kind of kind of adopting. Yeah. It's like fail fast, figure out what's uh what your problem is, and then and then keep keep moving on from there. So uh you said how how long you've been in the MySpace?
SPEAKER_01I was there for five years. So I think this is the one of the things that I heard from also former colleagues of us, maybe probably we need to get him on this podcast as well eventually, uh Sriram Suburbanian. He mentioned one of the things that these days you can put on resume whatever you want, and it's kind of like uh should be a little bit alarming for some of the hiring managers when you have like a distributed systems engineer in your resume and you only spend like a one year in each company. So there's not enough time to you to be like a fully solidify as a distributed systems engineer if you're not went through the multiple iterations of the same product. Oh yeah. Like there is a joke or maybe semi-joke when the people saying that if you're not embarrassed for the first version of your product, you're probably doing something wrong. A hundred percent. So that's that's that's uh aligned with what you just said, kind of like you don't have to do everything you know right at the first time, but you have a uh you need to know that you need to iterate and learn from this and how you can do this better next time. So that's why the idea of um relieving multiple major versions when you just constantly swapping not the constantly, but you know, you're swapping stack every like two years. You saw that something works, the technology of the time works um after two years, there would be totally different landscape that there might be some of the problems that you try to solve that someone else is will be solving. Like that's how we end up uh having Nginx instead of Apache for for for web front end. This is how we ended up having some other like a lighter like HTTP load balancers in in front of instead of like having some vendor uh based you know load balancers and things like that.
SPEAKER_02So it's funny, and you know, I go back and I look at like uh like how many uses of Kafka I would have had there. Yeah, it would have been uh it would have been insane. Yeah. And then you know, if you look at, hey, how would you do things now, you know, like like metrics we built our own. Yeah, you know, well Prometheus is pretty amazing. Why don't you use that? Yeah, but it was not there anymore. None of those things existed. And you know the other the other thing I got from that was um a good understanding of how to do uh production debugging. And you you kind of you kind of get a lot of uh of principles. So one of my jobs was kind of being an SRE before that. SRE was award, yeah.
SPEAKER_01Yeah, yeah. System administrator. That's how uh there was called that time. Like called Jeremy, you know. Yeah, Jeremy's just like showing up the for for like for cookies and like doing something with fixing the thing.
SPEAKER_02It was fun. I mean, you'd you get in these spots where like in all of these charts go red, and then you've got to go figure out, hey, what's the problem? And we learned a lot, and it was a great team. And you know, like I I had that similar experience and like uh working at Confluent. You just had a had a great team where you all have the same goal, yeah, and you're there to help each other and win.
SPEAKER_01Yeah. That's yeah. Yeah, and I think um interesting thing when you you have this like this mentality where those things were not existing, so you had to build yourself. And when I started um learning Kafka, it was around 2000 uh maybe 17, 16. Uh, we had the product that was similar to uh to do stream processing at Hazelcast that time. And I would start looking how to get data, and there was already availability of uh different ways how we can get data in. But you came a little bit earlier and you didn't have this luxury, and again, bring back this mentality that okay, we I didn't have these tools, I need to build those tools. So you went up and built some of the some of the very popular uh connectors.
SPEAKER_02Yeah, yeah. So you know when I when I joined uh Confluent, the only products that Confluent had at that point um that was an Apache Kafka was schema registry. And so and and I I believe we had the JDBC connector and the HDFS connector. And I think that was that was the entire products the product stack. Um and so like I I ended up, I was one of the first folks in the field. And so I would I'd end up talking to a customer and there, and a lot of it would be like, man, if I can get data from this and then do this to it and put it in that, you know, we'll we'll actually be able to get this into production. And so for me, you know, I was like, okay, you know, the short attention span can send, all right, I like this quick tactical deployment.
SPEAKER_01Waiting in the in the in the boarding in the airplane, and you have a time uh when you can have a snack and open laptop and start.
SPEAKER_02Yeah, I I actually used to do that a lot because I've I've lived in Austin for except for like like three months that I worked at Confluent. I've lived in Austin, and I used to have to come out here to the bay all the time. So what I would do is I'd meet with the customer, find out what they what you know they needed to do. Okay, pull down the Docker containers, yeah, and then work on it on the flight home.
SPEAKER_01Yeah.
SPEAKER_02And usually I could get, you know, I like to say it compiles, ship it. Yeah, I could get that quality usually by the time I got home.
SPEAKER_01Yeah.
SPEAKER_02And then, you know, hey, Mr. Customer, play with this. Does this kind of like what you're you're you're trying to do? Yeah. And then that's really where a lot of it came from is, you know, I wanted to, I wanted to help people, I wanted them to be successful. Yeah. And then, you know, at the same time, I wanted, I wanted Confluent to grow and be successful.
SPEAKER_01Yeah.
SPEAKER_02And the the easiest thing for me is I mean, if you've if you've ever talked to me before, I'm always pushing everything has to be Avro, everything, blah, blah, blah, blah. But in a lot of cases, people don't want to do that. And the um the connector ecosystem let me focus on the system. Yeah. So like I just need to get data out of the system and into connector format. Yep, and into a struct format. Yes, right. And then I I hand it off, and customer can say, Hey, I don't want your Avro. I want to use use JSON. And it's all transparent.
SPEAKER_01Yeah, because you can configure uh converter, and after that it will be handled uh by uh Kafka Connect uh ecosystem.
SPEAKER_02Yeah, and so that that worked out perfect for me. And so that so I just started building them and then putting them out on uh on my GitHub.
SPEAKER_01Yeah.
SPEAKER_02And then for a while I I used to directly publish them to uh Confluence uh a registry.
SPEAKER_01Yeah. But the registry was like years after, kind of like uh we we we realized that uh like having like App Store for connectors, it's it's not only uh beneficial for company uh to kind of like uh have a sort of taxonomy, but also beneficial for people to search them. And have a little bit of assurance that if it was in the part of some uh marketplace or some sort of like a connectors hub, uh there would be some assurance that it was at least tested. It's not only kind of like you know, compile shipped, but also it was tested and validated by by um uh kind of reputable vendor. So I would say that was huge um huge help for people to understand having this registry. And I think that time was like your connectors and there was like uh what the data mountaineer's uh connectors. They were they were building.
SPEAKER_02Yeah, they did a they did a great job.
SPEAKER_01Yeah, so there was a um huge ecosystem of this of different uh data sources um that would definitely benefit for bringing data from the sources into Kafka in order to enable some other um other use cases. Um so during this uh like what was your favorite connector to build? What was the most exciting or um something that maybe you know the closest? How close the biggest deal in that time, maybe.
SPEAKER_02Um oh man, you cut caught me off guard with that. So you didn't you didn't like send me a question?
SPEAKER_01You didn't ask uh which of your kids you love more, Victor. Um I uh like the My favorite was Twitter Connector. Before Twitter was kind of free social network. Uh Twitter Connector was great because it was able to enable such a great uh demos.
SPEAKER_02You but you had to be brave to do it uh uh like a live Twitter demo.
SPEAKER_01I I love this doing the trust y'all in that. I was doing this in uh in uh in Russia where the audience is uh wired to be kind of okay, let's see how we can break this stuff, how we can create some some machine malicious thing. So the Twitter connector was great. But it was for me, it was great because I was using specific hashtag, so I would capture some activities and they said, hey, you see, the people were having a blast during my talk. And um it's it's actually what great demonstration of um what's the streaming data, how this would look like. 100%. 100%.
SPEAKER_02Yeah, I mean that that one that one was a fun one. Um I uh the syslog one, that one, that one was a lot of fun. I I tend to like a lot of network, yeah, uh, um network admin type things. And so I always wanted to do like a I never had the opportunity to do it, but I always wanted to do like a large-scale network monitoring uh project using using Kafka. That was one thing I never never got the the opportunity to do. But I I did build some uh some of that and some like some Netflow collection stuff. Um I like I liked uh uh uh building and implementing uh protocols, doing doing some of those. Yeah. Um the file system connector, that one, that one opened a lot of.
SPEAKER_01Correct.
SPEAKER_02The the spoolder connector. Yeah, yeah. Yeah. So that one was that one was a lot of fun. And and um I because that one was fun for me because I always um I if if you ever talk to me as a customer, I was always everything's gotta be Avro. And I didn't like that, you know, the other connectors out there wouldn't let me get things in with a strong type. And so, you know, that one would if you would use my um kind of messed up idea of a of a schema, you could apply schema to data as you put as you put in. And that ended up being used pretty pretty heavily.
SPEAKER_01Yeah. Awesome. So, Jeremy, um, if you would summarize like everything that you learn over the time and already see kind of like there's uh uh one of the things that we already talk about is kind of like it's not has to be like a perfect and the first step. Um maybe some advice uh for people who listen to us. Like we I don't want to sound like it's like two old dudes sitting here just oh in my times, blah blah blah. Uh we still we still uh the teaching the technology of the future, there's still you know huge uh uh the cluster of the of the people who would love to learn things about Kafka connection processing and all those kind of things. So maybe some some advice for those type of listeners.
SPEAKER_02I would I would say design principles that I like to live by is if you can't explain your idea to another engineer in less than five minutes, it's too complex and needs to be broken up. And that's that's something I like to to to live by. The thing I will tell you throughout my career I've seen and in and some of the big environments I've worked in, simplicity scales. And the other thing to keep in mind, you don't have access to production. So make sure make sure you have proper logging.
SPEAKER_01Yes.
SPEAKER_02And then also logging that you can turn up.
SPEAKER_01Yeah.
SPEAKER_02So like sometimes info is not enough. Exactly. So like, you know, it's it's okay. Uh if if you look throughout my connectors, there'd be there I would use uh trace and trace you shouldn't turn on in in production because it could be a lot of it's gonna put a ton out there and it could put and it could uh dump some data.
SPEAKER_01Yeah.
SPEAKER_02Debug is gonna say, hey, I'm thinking about doing this. Yeah. And I did that. You know, but you you need to you need to uh um love yourself and make sure that you know if I have to uh uh get this data, there's a way for me to get it. So like pushing a config change to production is a lot easier than pushing an entire new release.
SPEAKER_01And that was Jeremy Custom Border, ladies and gentlemen. Uh amazing person to talk to. Uh and that was another great episode of Confluent Developer Podcast. I'm your host, Victor Gamoff, and as always, have a nice day.