Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov

Paving a Data Highway with Kafka Connect ft. Liz Bennett

Confluent, original creators of Apache Kafka® Season 1 Episode 83

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 46:01

The Stitch Fix team benefits from a centralized data integration platform at scale using Apache Kafka and Kafka Connect. Liz Bennett (Software Engineer, Confluent) got to play a key role building their real-time data streaming infrastructure. Liz explains how she implemented Apache Kafka® at Stitch Fix, her previous employer, where she successfully introduced Kafka first through a Kafka hackathon and then by pitching it to the management team. 

Her first piece of advice? Give it a cool name like The Data Highway. As part of the process, she prepared a detailed document proposing a Kafka roadmap, which eventually landed her in a meeting with management on how they would successfully integrate the product (spoiler: it worked!). 

If you’re curious about the pros and cons of Kafka Connect, the self-service aspect, how it does with scaling, metrics, helping data scientists, and more, this is your episode! You’ll also get to hear what Liz thinks her biggest win with Kafka has been.

EPISODE LINKS



SEASON 2
Hosted by Tim Berglund, Adi Polak and Viktor Gamov
Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
Music by Coastal Kites 
Artwork by Phil Vo 

  •  🎧 Subscribe to Confluent Developer wherever you listen to podcasts. 
  • ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
  • 👍 If you enjoyed this, please leave us a rating. 
  • 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
SPEAKER_03

When Liz Bennett was at Stitch Fix, she got the chance to build a company-wide stream processing platform from the ground up. She learned a lot of technical lessons and even more importantly, some technology leadership lessons. She shares both on today's episode of Streaming Audio, a podcast about Kafka, Confluent, and the cloud. Hello and welcome back to another episode of Streaming Audio. I am, as always, your host, Tim Bergland, and I'm joined today in the virtual studio by Liz Bennett. Liz is, of course, famously the second of the five Bennett sisters growing up on the Longburn estate in Hertfordshire, England.

SPEAKER_00

Uh uh, Tim?

SPEAKER_03

Yes, Liz Bennett.

SPEAKER_00

Yeah, that's a that's a different Liz Bennett. Yeah. Oh. Yeah, she she's from the 1800s.

SPEAKER_03

Oh which Liz Bennett are you?

SPEAKER_00

I am Liz Bennett, who is a software engineer at Confluent. Oh. I've seen you in Slack.

SPEAKER_03

Oh, this is awkward. Okay. Well, uh all right. Uh, everybody, this is Liz Bennett, and she's a coworker of mine. She's a software engineer who works at Confluent. Um Liz, tell us about yourself.

SPEAKER_00

Great. Yeah. So first of all, thank you so much for having me on the podcast. Uh I just I love the idea of having this streaming audio podcast, and it's just like such a fun medium for diving into these, you know, sometimes complicated technical subjects. So um okay, about me. I uh let's see, I have been working with Kafka for a long time. Uh a few years ago I worked at Logly, which is a logs as a service company. And um after that I switched to Stitchfix, where I built all of their streaming infrastructure, and I really got my hands uh I really got a lot of experience working with Kafka Connect, and we we invested heavily in it, and I'm a humongous fan of Kafka Connect. So much so that I decided to join the rest of the Kafka Connect fanatics here at Confluent uh recently, about four months ago.

SPEAKER_03

Aaron Ross Powell Nice. Yeah. So yeah, and you work on Kinect here.

SPEAKER_00

So I was originally going to join the Kinect team, and I may end up on the Kinect team eventually, but for the time being, I'm on the cloud team because I wanted to sort of learn some new skills and you know get better at Kubernetes and things like that.

SPEAKER_03

So being on the cloud team is definitely a way to get better at Kubernetes.

SPEAKER_00

Yes. It's been a crash course in Kubernetes, for sure.

SPEAKER_03

Yeah, I bet it has. And uh, to be fair, I mean it's uh I always it's always important to say when we're recording this. I mean, most people listen to them right after they come out, but we're recording this in the middle of January 2020. And the current state of affairs in Confluent Cloud is um kind of the rapid rapid cloudification of as many connectors as possible. So I expect having a connect expert on the cloud team isn't exactly a harmful thing.

SPEAKER_00

Right. Yeah.

SPEAKER_03

Cool. All right. Well, um awesome. And it is great to have you as coworker. I know you have uh uh kidding aside, you've joined just two months ago before this recording. It's great to have you. It's great to be here. And what I would love to talk about today is uh, you know, I a lot of what you did at Stitch Fix was data integration things, and data integration things are in Kafka are sort of definitionally Kafka Connect things. Right. To the extent that uh you're able to, I would love to talk about what you did there and just kind of it's been a while since we've had a Connect show. Um it's honestly everybody always loves talking about Connect, you're hearing us talk about Connect. So yeah. Um yeah, tell us about a project or what just take us, take us from the beginning. What did you begin by prototyping?

SPEAKER_00

Yeah, yeah, I would love to. So when I was when I joined StitchFix, I was leaving this logly, logs as a service company, and I I was hired at Stitchfix basically to build their logging infrastructure. So um the director of the data platform team there, which is the data platform team, is the kind of engineering team that supports all of the data scientists at StitchFix. So there's like a hundred data scientists there, and there's a 20% data platform team building all of the uh systems and tools and utilities to enable data science at Stitchfix. Uh for those of you who aren't familiar with Stitchfix, um it's uh it's basically uh uh personal shopping as a service. Aaron Powell Clothing as a service, right? Yeah, clothing as a service.

SPEAKER_03

Maybe it's we should describe it as dressing nicely as a service.

SPEAKER_00

That's a good way to put it, yeah. Uh so uh they're one of their main one of the main pillars of their business is data science. The the business runs on recommendation algorithms, and also there's tons of algorithms that run internal internal to the business that you know operate the warehouses and help with um the personal styling aspects, and just uh they really um invested so heavily in just making data uh the first class citizen of Stitchvix. And StitchVix is part of this kind of new crop of companies that is is shifting toward that. So they had a chief uh algorithms officer, a C-level executive in charge of the whole data org. So that's okay. So that's context about Stitchvix. So the data platform team is was building all the engineering systems to support that. And when I joined, they didn't really have logging infrastructure. Uh like um, you know, and most companies start out not having logging infrastructure, and then at some point they invest in it, and then they have logging infrastructure, and it changes everything and it and it unlocks a lot of capabilities. And so it's great. So I would I joined to basically make that happen for them. Uh so there was no there were no um guidelines, no really hard requirements. It was just like do what you can do with logs.

SPEAKER_03

Okay.

SPEAKER_00

So yeah.

SPEAKER_03

Um, they say we have log4j, now make it not bad.

unknown

Right. Yeah.

SPEAKER_00

Uh and I so I at first I was thinking, okay, um I can just build a system that gets log data and ships it into the data warehouse, like build a a um kind of a purpose, you know, purpose-built system that kind of just fills that one need. But as I was doing more research, I came across Kafka Connect, and I read uh JKreps' iHeartLogs book. Um I read a lot of things that J. Krebs wrote actually, and I started to see the possibility of having this centralized data platform and uh like data integration system. And the I started to think like, what if you instead of having a logging system, you had an event streaming system where you could ingest data from anywhere across the entire company and then write it to anywhere in the entire company. That's so much more powerful than just a logging system. I mean, it really is the central nervous system of the company. So I was tasked to fix the logs and I made it my mission to build that central nervous system. And the biggest piece of that was Kafka Connect. So very early on, I stood up a Kafka Connect cluster, I stood up a Kafka cluster, and uh then the hard part was making it useful for data scientists. So uh one of the major tenants of the data platform team at Stitchfix is that all the tools we build as engineers should be self-service for data scientists. So a data scientist should never file a Jira ticket and someone in the engineering team goes and makes it happen. We provide an interface and a UI so that the data scientists can do it themselves. Nice. Uh so yeah, definitely the the tricky part of this this whole project was how do I stand up uh a Kafka cluster, Kafka Connect, and also make it directly usable and accessible by data scientists.

SPEAKER_03

Now I wanna I want to talk about that. Um that self-serve thing sounds very important, but I want to back up a step because you said something actually a lot bigger than that, which was that you started with a mandate to fix logging, um, because they had you know probably the sort of the primitive default um logging kind of situation that most people have before they think about it, right? Okay, we want to make investments, let's make this centralized and get stuff into the data warehouse. And you said uh that would be nice. Also, uh a revolutionary new architectural paradigm would probably be a better solution. Those aren't your words, but that's my translation of your words. How did you go about that? Um, and I I ask because like when I'm occasionally like on a panel at a conference or kind of hallway track questions at conferences when I'm um traveling and speaking, probably like I don't know if it's the majority question, but the plurality question is how do I get my company to adopt this thing that you just got me excited about? And for the kind of engineer who is forward-looking and is keyed into the fact that, okay, streaming is is gonna unlock new capabilities, it's the thing I should do next in my world, it's it's still uh a new paradigm, right? It's still a scary new idea that nobody knows how quite to do. So, what was that like to like how did you do that? And in organizationally, yeah, it's really it's a leadership question more than a technology question.

SPEAKER_00

Yeah, that is such a good question because it you're right that it was it wasn't always easy, and uh it took a lot of convincing people and a lot of evangelizing. So yeah, um that was that was tough. I'd say the first thing I did was to set up a prototype. So just download Kafka, like set it up. I uh I scheduled a hackathon on the data platform team, a Kafka hackathon, and I invited everybody who was interested to come and join and just play with Kafka and play around with it. Most of the people had never used Kafka, they didn't know anything about it. And so uh we all just got in a room together and and and tried to see what was possible and what we could do with it. I think one one guy, he um created uh like an ASCII movie. Like he would pipe in ASCII characters from as a like in a Kafka producer, and then he had a consumer that would play it on another machine and it would like play a movie. Um so you know, your typical hack day stuff. So I think like that was really useful for getting the engineers on board and and just sort of having bottom-up uh momentum. And then to sell it further up the management chain, the one thing I really attribute to the success of the project was um picking a good name for it. And so we called it the Data Highway, which is actually a pretty common name for this project. I I've definitely heard uh several other companies with the same exact project also call it the Data Highway. Uh but it was really useful for explaining what I want it to what I wanted it to be. If you think about how hot what a highway is and how there are highways across the entire world, really, it it is just a way to get people from anywhere they want to, from anywhere they are to anywhere they want to go. And it's just such a perfect model for capturing what this infrastructure is. And it also makes you think like, wow, people used to take the train everywhere, and then we built all these highways, and now everybody gets you can get places so much faster, and you you it's so much more um granular, the the way transportation works, and and it it's just so perfect. It captures it so perfectly.

SPEAKER_03

Aaron Powell Well that analogy is uh painfully good. Trains trains are much more like batch, right? You have to load them up. And there are certain economies of scale there, right? There are some things that are good for that. True. Um but I don't wanna I mean I'm an American- I don't I'm an American. I don't want to have to get loaded up on a thing with a put me in my car.

SPEAKER_00

I know, yeah. And I think it resonates with Americans uh, you know, Sizu's American company. And also like in a different thing.

SPEAKER_03

We love trains. We also love you.

SPEAKER_00

I mean personally, I I like trains more than I like cars. So but it worked for the for the purposes of this uh, you know, uh as a metaphor for this project. Also, since uh it was a data org, the whole org was very batch-oriented, so uh it helps a lot to use that train metaphor. So that was one thing. Having a great name, having a great branding for it, uh people remember. They think, oh, the data highway, I remember what that is, I know what it's supposed to do, kind of. And then the second thing was I created a vision document. So this was um quite an investment of just sort of thought leadership for me. I spent at least a couple weeks working on it, and it was this like eight or ten-page document that was that laid out exactly what I wanted to do, exactly what like how I wanted to implement it, exactly what the user interface was going to look like and what capabilities it was going to unlock, what problems was it solving. And um when when people read that, they I think they really started to get it. And actually, when I first wrote this document, I sent it to my manager and I sent it to the director of the data platform team. And the next day he scheduled a three-hour meeting that just had Data Highway on it. And I was like, oh God. Like when the director schedules a three-hour meeting with you and your manager, uh that was that was a little bit of a scary moment in the project. A little bit.

SPEAKER_03

Uh, yeah. That uh that escalated quickly.

SPEAKER_00

Yeah, but it ended up being of really productive meaning. We just went through the document paragraph by paragraph, and he critiqued it and he he really shifted a lot of the alignment of it and and just sort of made sure that we were all of us were on exactly the same page and that what I was building worked with what his vision for the team was and so on and so forth. So I think it just helps so much to put in writing and to have a very clear vision for what you're going to do before you start doing it. So Yeah.

SPEAKER_03

So the steps were uh hackathon to get to get grassroots interest, cool name, compelling and precise metaphor, and then vision document that addressed business goals.

SPEAKER_00

Right.

SPEAKER_03

Uh and not just technology goals.

SPEAKER_00

Yeah. Yeah. And it did include a lot of detail around the technology also. Okay.

SPEAKER_03

Technically beefy vision document, including business goals. And shout out to your director who was a sufficiently open and forward-looking person to be able to respond to that.

SPEAKER_00

Yeah, yeah.

SPEAKER_03

So you you had to be clear. I mean, I'm I'm I'm trying to put a spotlight. And and by the way, listening audience, I actually didn't know this part of what Liz and I would be talking about. She just kind of said those things, and I'm like, eh, we're gonna spend a while talking about that because that's extra awesome.

SPEAKER_00

Yeah.

SPEAKER_03

Well, it was a great question, actually.

SPEAKER_00

Yeah.

SPEAKER_03

What's that?

SPEAKER_00

It was a great question. It's a really good thing to dive into because it was very tricky.

SPEAKER_03

Because everybody stumbles on this.

SPEAKER_00

Yeah.

SPEAKER_03

Because this this podcast is about a technology category that is that is new and difficult. Right. Nobody, when I visit, when I go and I talk about this all the time, when I sit in a room with enterprise software architects and they're like, well, how do we do payments with streams? It's hard because you didn't have a class in that and you have to think through it anew. And so this is everybody's struggle. And and how do you convince you once you have the intuition that this is going to work, how do you convince your leadership? So this is awesome.

SPEAKER_00

It's also very much a uh engineering or or technology-driven paradigm shift. Like engineers, people on the front line, they can see the value of event streaming platform. Leadership, it's it's a harder sell. It's so it's such an esoteric thing still. So I yeah, there's this a lot of this bottoms-up kind of convincing of leadership.

SPEAKER_03

Right, right. And um yeah, well, terrific. Okay. I forgot, I forget where we even were. You you uh were fixing logging and you decided to convince um this non-trivially sized organization to adopt a new architectural paradigm. You succeeded. So take us from there.

SPEAKER_00

Right. So once the Vision Doc, you know, had this stamp of approval, we just got to work. So um we stood up Kafka, we stood up Kafka Connect, we decided which connectors we were going to install, and we built a few services on top. So uh back to this whole idea of self-service um and making a platform self-service for data scientists, we really wanted to have uh a very nice UI, like um, you know, GUI, like a web-based user interface to interact with the platform. And we built um uh uh you know HTTP utilities, so uh data scientists could produce data through HTTP. You could also use the HTTP REST proxy, but we wanted to have more control over the the published side of things.

SPEAKER_03

Sure.

SPEAKER_00

And we built um we built this admin service, which sat on top of Kafka Connect mostly, but it kind of was the glue that tied everything together in the platform. And it was the back end for the UI also. And that admin service was really, really useful. So that would be um one thing I would highly recommend if you're going to build a uh a platform for yourself. Like at least put one uh system in there that you you have complete control over if you want to have like a really nice cohesive sort of like custom feel to it.

SPEAKER_03

Sure. Would that um if I were to use the somewhat trendy term control plane to refer to that in terms of your service, would that be more or less precise?

SPEAKER_00

Yeah. Yeah, that would be kind of your control plane. Pretty accurate, yeah. We called it uh Caltrans because it was part of the data highway theme.

SPEAKER_03

Of course, Caltrans sounding like a train network in a certain geography.

SPEAKER_01

Right, right. Yeah.

SPEAKER_03

For those of you not familiar, it would be Caltran.

SPEAKER_01

Um also Caltrain Park.

SPEAKER_03

Oh, Caltrans is a button of the state government of California.

SPEAKER_00

Yeah.

SPEAKER_03

See, when you live in one of the square states, you just don't know these things. Um the GUI and the self-service aspects. I would like to dive into those a little bit.

SPEAKER_00

Right.

SPEAKER_03

And that again, I I don't want to not talk about interesting Connect things. I'm gonna dig into some Connect technology stuff in a little bit. But it strikes me that that that self-service thing is as much, again, a social and organizational hack as anything. And I think it's absolutely critical to initiatives as ambitious as yours being successful. So talk about your motivations there and what were the impacts and how did it work out?

SPEAKER_00

So let's see. Like what was our ambition with the with the um with the GUI?

SPEAKER_03

Yeah, why did you do that?

SPEAKER_00

So we wanted it to feel like a data science product. Like it was a product designed for data scientists with their needs in mind. I think this is one thing that's an opportunity for Confluent and for Kafka in general, is that um there the Confluent ecosystem is pretty far away from the data science ecosystem. So what we wanted to do was build a UI that had first class integration with the Hive Metastore, with Spark, um, with uh um like data governance, uh you know, um like it would integrate with our data governance system. Like it was really just uh the glue that tied the data science world to the Kafka world.

SPEAKER_03

Um when you got into um I I just let me follow up on that a little bit, making things self-serve for clients like uh data scientists, you're you're building plumbing and nice fixtures and things for data scientists to use to do their work. But I think giving them self serve tools gives them this low transaction cost ability to get data in and out and um just do whatever they think of without having to ask you for permission or allocate engineering resources or anything. You always want to like free up whoever the Clients are the system. You don't want them to need you after the infrastructure is built.

SPEAKER_00

Yeah, exactly. Yeah, that's huge. That was huge. Right. Yeah. Because like if they're coming to you all the time, like it's a drag on them, it's a drag on you.

SPEAKER_03

And they'll stop. Yeah. Uh is the thing. You won't get adoption and you won't get use if their transaction cost is high. They'll just do this one thing where once a month they dump things into their MySQL that they know and love, and it's nice to them and it's always been their friend and is a little bit of a dysfunctional relationship, but it's the one they know. And they won't get all the good stuff that you're trying to build for them. So that's Yeah.

SPEAKER_00

And with this, especially, like it was like night and day when we had the we had our, you know, our legacy logging system, which we had before I joined. And it was very um, it was a manual process. Uh a data scientist, if they really wanted logs, they would come to us. We would go, we would set some manual stuff up, we would create a stream for them, you know, and then a few days later it would be like, okay, it's all ready now. And we had, I don't know, I want to say like 10, 11 topics total. And then as soon as the data highway came out, it was self-service, it was like it shot up like 10x the number of topics and the number of users and consumers and producers like within a quarter. Like, and there was just so much need for it and so much desire and so much curiosity around it, but it it was just too much of a bottleneck, this manual process.

unknown

Right.

SPEAKER_03

So did that were there any infrastructure strains that you had to over when, you know, oops, I built this thing that everybody loves and they want to use it. Uh were there any challenges in scaling things uh when that big bump happened?

SPEAKER_00

Ooh. Yeah, I mean there were a ton there were definitely operational challenges that I would love to kind of spend some time really digging into if you if you want to if you want to go there. I would love to. Um tell us what happened. Okay. Yeah, yeah. So initially, I mean uh the the Kafka cluster was way over provisioned, which you know I I basically did that like from my experience at Logly, we we you know we ran a really tight shop and um uh the infrastructure we were we were running it pretty hot most of the time. And so with Stitrix, like uh I I just decided it's so much easier from an operational perspective to just throw more hardware at it and you have a lot less outages and there's the most expensive resource in any technology company is the time of the developers and the time of the people. So if you can just avoid having production issues and avoid having outages and things like that as much as possible by throwing a little bit more hardware at it, then like it's just so worth it. So that got us through for a long time, actually. So your your original question was like, did something did something bad happen when we first opened the floodgates? Um and the answer is no. Uh we just it it just wasn't it wasn't an issue for for a while, actually, like maybe a year or so. Um and then yeah, we did start to run into issues like with Kafka at least, like you can you can add more brokers, you can you can increase the size of your nodes. It's it's a pretty straightforward um scaling process. We would always just scale up when we started running out of disk space or we started running you know out of network. Um that was never too much of an issue. Kafka Connect, on the other hand, I mean, okay, first of all, I'd like to say I love Kafka Connect so much.

unknown

Yeah.

SPEAKER_03

Just want to be clear.

SPEAKER_00

Yeah, it works really well so much of the time, and it's just a really well-engineered system. It it I you know, I love the the API for writing your own connectors. It's so easy to write your own connectors. You can yeah. Um you you can you it's just it really like okay, so I I've been working with Kafka for a long time, writing Kafka applications, even before logly, I was at LinkedIn uh working with Kafka. And Kafka Connect just cuts down the amount of work you have to do to make a nice dreaming pipeline by so much. So yes, I love Kafka Connect. It did have some issues for us because we we started to have so many connectors. We had tons of like thousands of connectors in our Kafka cluster.

SPEAKER_03

Um That's a lot.

SPEAKER_00

And most of those issues were related to rebalancing.

SPEAKER_03

Okay.

SPEAKER_00

And that's actually Connect fixed node.

SPEAKER_03

Uh connect rebalancing. Yeah, that was a thing I think that changed in 2.3 or was it 2.4? It's recent. I know it's recent. I think we have summary videos of these things so you'd think I'd remember. But it's uh definitely not 2.2, and 2.5 doesn't exist at the time of this recording. So it's one of those two versions.

SPEAKER_00

Yeah, I think it was 2. Yeah, either 2.3 or 2.4.

SPEAKER_03

Tell us uh for those who don't know, uh walk us through what connector rebalancing is and why it was bad for you.

SPEAKER_00

Yeah, so in Kafka Connect, you have a bunch of connectors, and each connector is composed of a number of tasks. And those tasks are threads, basically. They're threads that run in your um in your cluster. And anytime you would create a new connector or delete a connector, or change the number of tasks in a connector, Kafka Connect would try to balance the tasks across the cluster so that you don't have hot spots or hot nodes or things like that.

SPEAKER_03

Seems perfectly reasonable. Let's make sure everything's fairly distributed.

SPEAKER_00

Yeah, yeah. Uh the issue was this is before um cooperative uh incremental rebalancing, which is recent. So um if you're just starting with Kafka Connect now and you're using a more recent version, this is not going to be a problem. So this is really applies to people who haven't upgraded their Kafka Connect cluster yet and are starting to run into this problem. Um so what would happen is if you add a connector or delete a connector, all of the threads, all the tasks would shut down and they would all come back up again. It was a stop the world rebalance. And um if you have a few thousand connectors and a few thousand threads and they're all shut down and they all come back up again at the same time, you can start to have like these thundering herd issues where uh and it it manifested in in several different ways. So we had several different sort of um issues and outages that were caused from that.

SPEAKER_03

How long does a rebalance take in a cluster like this?

SPEAKER_00

Ooh, gosh, like it like before uh incremental cooperative. Like again, before incremental, I mean it would be like 30 seconds to a minute sometimes.

SPEAKER_03

Okay. So that is uh significant. That's not there's there is no real time anything here. And how big how big were your connect clusters in terms of nodes?

SPEAKER_00

Ooh gosh.

SPEAKER_03

If you can talk about these things, I realize we're talking about kind of somebody else's stuff at this point.

SPEAKER_00

True, yeah. I guess I should say um pretty uh big, not extremely massive in terms of the size of the nodes or the number of nodes. Our connector, we had a ton of connectors, but most of those connectors were pretty low in volume. So we ran into issues that you get when you have tons and tons of connectors, but not necessarily a lot of volume. I mean, it would happen if you had a lot of volume too.

SPEAKER_03

Yeah, either way. But these are clusters, you know, these are connect clusters. Yeah.

SPEAKER_00

Yeah.

SPEAKER_03

Uh cool. Okay, so um I keep uh I keep uh I keep distracting you. This is the uh if we're talking about the rebalancing uh the the the rebalancing problem. That was one of the things that you saw go wrong right away. And I think it's important to note that this is a this is a pretty beefy connect use. You know, if you've got hundreds or thousands of connectors in there, that's a lot. So you are you are all in on this.

SPEAKER_00

I think like I think we're sort of on the on the vanguard. I don't think it's I don't think it'll be unusual to see connect clusters like that in the near future in like the next year or two.

SPEAKER_03

As things, as uh as adoption continues. Uh how about tooling, you know, connect specific tooling. What were some things you discovered there?

SPEAKER_00

So oof. So okay, being a data being a data org, uh a lot of our other tooling a lot of the tooling in the in the org is um uh like job-based. So we have like jobs that run every hour or so. So a lot of the most useful things that we found, we we sort of encoded into jobs. So we had a job that would restart connectors. This one was huge, actually. Because connectors, when you have a thousand connectors or more, the connectors will crash like a lot, like pretty often for um just like network blips or or you know, uh, we had the S3 connector, and the S3 connector would for some reason get an error from AWS, like transient error, and then the connector would crash and go down. Um and so uh we started to see this very early on with Kafka Connect. So we just set up this job that would automatically restart the connectors when they crashed, and uh that was kind of the last we ever had to think about it. So yeah, they would they would start back up again. It would, I think it would post to Slack when it was restarting a connector, so we could kind of keep an eye. Um and that's important too, because sometimes connectors will will crash for a real reason, you know, you don't want to just keep restarting them. Um like because they'll just keep crashing. Uh so that was a big one. Another one was uh an auto-scale job. So you ideally want to have a connector, you want to have the right number of tasks for the amount of volume in the topic that your connector is consuming from. This is for a sync connector mostly. Um since it was uh an entirely self-service system, people could go in and actually add partitions to their Kafka topics, and they wouldn't always remember to add tasks to their connectors, which like you know, they're data scientists. We don't really want them to have to be thinking about things like that. So we just set up these auto-scaling jobs to kind of make sure everything is all um, you know, working as optimally optimally as it can.

SPEAKER_03

Aaron Powell So the way that would work would be to periodically check uh the number of partitions on the topics of interest and make sure that the number of tasks in the connectors of interest match them.

SPEAKER_00

Aaron Powell Yeah. That was the first that was the first implementation. Ideally, though, you measure the volume in the topic and the the rate of messages. Uh it was a self-service system, so sometimes people are like, I'm gonna make it a hundred partitions.

SPEAKER_03

And then they send like Because they're not very good at this.

unknown

Yeah.

SPEAKER_00

So um so if you're you know measuring and you're like, okay, this topic has a hundred messages per second, it really doesn't need more than like one task, really. So um even if it has like a hundred partitions. Um yeah, that was those were some some some of the most useful ones that um that are coming to mind right now.

SPEAKER_03

Nice.

SPEAKER_00

How about yeah, we could talk about metrics and things too.

SPEAKER_03

I wanted to get some metrics. Tell me about uh what you collected and how you how you used it. I mean with metrics, there's always two interesting parts of the story. There's there's usually a how do you get at the thing that you want to know? Um, and then what are you actually doing with the metrics to give additional understanding to someone?

SPEAKER_00

Right, right. There are three metrics that were by far the most useful. And you'd want to have these metrics even if you weren't running Kafka Connect, but it's they're especially useful for Kafka Connect, which is the topic volume like metrics. So you um really need to track the number of messages per second and also the bytes per second. It's really important to have both of those. And then uh you really, really need to have the consumer lag. So you need to know how far behind are your this is like for sync connectors, how far behind are your sync connectors falling in their um in consuming from their topics. And then uh oh, you know what? No, there's four, there's four metrics. Okay. The fourth one is um heap usage in your Kafka Connect cluster. So that was always a leading indicator of when we needed to scale our Kafka Connect cluster. Okay, that makes all the sense in the world, but that's yeah. Yeah. So the Kafka Connect is pretty pretty heap intensive. It does a lot of buffering. It's you know, reading and pulling in data from Kafka, buffering it, writing it out to other systems. Unless you have like complicated transforms, you're not gonna be super um CPU bound. Um yeah, maybe network, but like always, yeah, it was always heap was the was the limiting factor.

SPEAKER_03

That is interesting. Uh that's a good uh that's a good thing to know. How about some other um I I know you know you you made it clear that you're a huge fan of Kinect, and that's very clear from your background and from uh your work here and kind of where you'd like your work to go. But I the the interesting things are not always here's what's awesome, but here's what wasn't. What what was another gotcha that would be a good warning for others?

SPEAKER_00

I would say the API itself. So Kafka Connect has an HTTP API that you can use to create and modify and delete connectors. And that's really useful. But the API is really weird. Like it's just weird. Like uh and I I understand why they did it the way they did it. Uh it makes sense knowing what I know about Kafka and how um properties and configs work in Kafka. And I you know, I think they were just working with the constraints that they had. Uh but if you were just going to start from scratch and build an HTTP API that makes sense for Kafka Connect, uh the API it has is is not what what I would have designed. Um so that was one of the things that we did in our like control plane application that I mentioned earlier. We actually wrapped the Kafka Connect API and exposed a much simpler API that just really abstracted a lot of that awkwardness from the users and just gave them a really nice clean HTTP, you know, sensible HTTP.

SPEAKER_03

The abstraction that that was uh that that gave your users what exactly what they needed and no more.

SPEAKER_00

Right, yeah. Um and I think one other thing too is that um Kafka connectors can be paused, uh which is nice if you want to stop if you want to stop processing data or you want to sort of like uh maybe you have a sync connector and you're writing to Elasticsearch and Elasticsearch is having an outage or it's really overheating and you want to pause all your Elasticsearch connectors. Um that's great. But if you want to do something like um uh upgrade your Elasticsearch cluster or do something where like ideally you would actually uh release all of your resources, release the TCP connections, things like that, uh you really want to stop the connector. You want to shut down the thread, have it release everything, uh and just kind of you know completely shut down. But there's no there's no API to do that. You have to actually delete the connector and then recreate it. That's the only way to stop it. So that always kind of got to me. I think I should like uh bug some people here now that I'm now that I'm at confluent.

SPEAKER_03

Maybe you uh have access to some people who work on that. Um you always just get on the mailing list, but uh you know you know there are there are more people that you can come in uh easy contact with now, I guess.

SPEAKER_00

Yeah, yeah. Well, and I'm I'm kind of scraping the bottom of the barrel here because there's not that many things I don't like about Copka Connect.

SPEAKER_03

So what do you think was the biggest win from that thing that you built, whether connect related or um you know, to end to end on a good note, what was uh what was your favorite part of the outcome? Your favorite thing that you did?

SPEAKER_00

Yeah. So at Stitchfix, we had this data org, which had all the data scientists and all the algorithms and and all the you know all of that. We also had a engineering org, which had all of the um engineers that worked on uh on on the product, like they're product engineers, they're web developers, um Ruby and Rails developers. So uh they they were since they were in a different org, they had their an entirely separate set of systems, entirely different infrastructure. They had they even had their own platform team building separate tools for them. And the only uh you know, they had their own C-level exec. So getting data from their engineering org over to the data org was so difficult before the data highway, because we didn't have this nice system designed for moving data from A to B. Uh we had tons of little like one-off things. Um it was just I remember actually one of my first months there, I I overheard a conversation between a data scientist and somebody on the engine in the engineering org, and they were talking about like, okay, engineering org has this data that the data scientists want. How can we get that data from their org to the data org? And I was kind of eavesdropping, and they were just going back and forth, and it and it just the conversation just ended in frustration. Like they didn't have a good way. It just didn't exist. And after the data highway, like they probably wouldn't even needed to have a conversation in the first place because they could just go to the to the UI and set it up. Like it could have been a probably like a Slack conversation.

SPEAKER_02

Like so, like a reminder of what's that what's that topic called? Uh oh, okay, thanks.

SPEAKER_00

Yeah, yeah, exactly. Yeah. So like you know, the engineering org could uh could produce all the data and they just send it into the data highway and the data highway a lot of the time would just um sort of transparently pick up the data. Like the engineering org doesn't even have to um they didn't even have to explicitly publish it into Kafka uh because the source connectors would would just kind of suck it up like a like a I kind of call it like a like a big brother kind of system that's just watching you and watching everything the engineering org does, um, grabbing that data and and siphoning it off into the data org.

SPEAKER_03

Except within the context of the actually shared interests of the organization. Yeah. And boots.

SPEAKER_00

We're not like stepping on human faces. Yeah.

SPEAKER_03

That's a different novel. We're supposed to be making Pride and Prejudice. Uh we're off into 1984. Significantly less happy novel.

SPEAKER_00

Yeah. Um I think yeah, that was that was such a huge one. It just became so much easier to for the data scientists to get the data they needed. It was so much easier for engineers to to give the data. It it just it it eased so many things. Like just a hundred different use cases were s were were made so much easier from it. So I'd say that was the biggest win for sure.

SPEAKER_03

And that that really is the vision. Uh it's great to hear you say that. And that's again, sometimes I have to point these out and say that this is not there's nothing rehearsed about that question or that answer. That, however, is the vision of the event streaming platform as the central nervous system of a company. That data that's over there that you would like here, you don't have to ask.

SPEAKER_00

Yeah.

SPEAKER_03

Um I mean, there can be governance rules. This particular kind of organization is not especially regulated. You know, there'd be regulations around personal um information and things like that, but otherwise clothing is not all that regulated. Yeah, different industries have different constraints there. But subject to governance, the data is there. And you don't have to ask permission or get resources allocated, you just get it.

SPEAKER_00

Yeah, it's uh you know the real data democratization vision. Yeah.

SPEAKER_03

Indeed. My guest today has been Liz Bennett. Liz, thanks for being a part of Streaming Audio.

SPEAKER_00

Yeah, thank you so much. This is a really great conversation, Tim.

SPEAKER_03

And there you have it. Hey, it's Kafka Summit time again, and you get another discount code for listening all the way to the end. Kafka Summit London is coming up on April 27th and 28th of 2020, and you can get 30% off your registration if you go to Kafka-summit.org and use the discount code KSL20Audio during checkout. Just enter KSL20 Audio while registering at Kafka-Summit.org, and that 30% off is all yours. I would love to see you there. And anyway, I hope this podcast was helpful to you. If you want to discuss the podcast or ask a question, you can reach out to me at at TLBergland on Twitter. That's at T-L-B-E-R-G-L-U-N-D, or you can leave a comment on a YouTube video or reach out to us in Community Slack. There's a Slack sign-up link in the show notes if you want to join that group. And while you're at it, please subscribe to our YouTube channel and to this podcast wherever fine podcasts are sold. If you subscribe through iTunes, be sure to leave us a review there that helps other people discover the podcast, which we think is a good thing. Thanks for your support, and we'll see you next time.