Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov

Transparent GDPR Encryption with David Jacot

Confluent, original creators of Apache Kafka® Season 1 Episode 46

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 16:45

The General Data Protection Regulation (GDPR) has challenged many enterprises to rethink how they deal with customer data. Viktor Gamov chats with David Jacot about a unique approach to inter-broker traffic encryption that he implemented for his customer’s sidecar pattern use case.

EPISODE LINKS

SEASON 2
Hosted by Tim Berglund, Adi Polak and Viktor Gamov
Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
Music by Coastal Kites 
Artwork by Phil Vo 

  •  🎧 Subscribe to Confluent Developer wherever you listen to podcasts. 
  • ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
  • 👍 If you enjoyed this, please leave us a rating. 
  • 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
SPEAKER_00

The General Data Protection Regulation, or GDPR to its friends, has been a challenge to a lot of digital enterprises and made them rethink how they deal with customer data. Now, Victor Gamov recently chatted with David Jacot about a unique approach David has been taking to interbroker traffic encryption that meshes with his customers' sidecar pattern use case. Victor's going to dig into that on this special episode of Streaming Audio, a podcast about Kafka, Confluent, and the cloud.

SPEAKER_02

Hello and welcome to this episode of Streaming Audio Podcast. And this is Victor Gamov, and I'm coming at you from Kafka Summit, London, where I'm interviewing uh speakers who present in this event. And today, with me, uh my special guest, David Jacob.

SPEAKER_01

Hi, David, and welcome to Streaming Audio. Thanks for having me.

SPEAKER_02

Yes, it's it's uh it's exciting uh that um I was uh host in the monitor of the of the session where David was speaking, and um we were chatting about this the what we're gonna be talking about in this podcast, and I think one of the topics that uh David covered, apart from the technical aspect, we can talk about this as well. Um, is uh the aspect of uh GDPR and how to deal with the uh encryption and how to deal with the um expiring some old data from Kafka and so far and so on. So, David, welcome to the show. Um, can you talk a little bit about yourself, what you do, uh what you um engineering, what do you do in life, uh what brought you here in Kafka Summit, uh etc.

SPEAKER_01

Sure. So you know I'm regular like uh software engineer with like a regular background in this field. Um I have been working quite a long time in various companies uh with Kafka for a few years now, you know, in small startups but also in large enterprises like in the telco business. And uh yeah, uh you know I've been using Kaka for such a long time. I like the technology, I think it's really great piece of software altogether, and it uh really helped companies to you know to think about all this event slash streaming processing in a different way. So I like it a lot.

SPEAKER_02

So what's uh what was the Kafka version when you started um working with Kafka?

SPEAKER_01

Yeah, to be honest, I think it was 06 or 07. I don't I don't remember exactly, but you know, basically more or less when LinkedIn open source uh the the thing on GitHub back then we started to use it you know in a small startup I was working with.

SPEAKER_02

Um what was the uh use case?

SPEAKER_01

Why you think that you know basically it was a company doing like um you know data processing in a social media space. So it was basically about you know timing the Twitter firehose, you know, ingest a lot of social data. And back then, you know, we we used Kafka as like already like uh you know backbone to do microservices uh data processing around it. And and you know, we we needed that for scalability reasons back then, right? And it was like a perfect fit for us.

SPEAKER_02

Yeah, and uh I guess zero six, there was no like there was different Kafka, but it was really, really different.

SPEAKER_01

I mean it's uh it's really it's really amazing to think about uh the changes in that, you know, that there were no replication at all. Also the the the thing, you know, there were no logical offsets, for instance, uh you know you had to use like really the offsets, you know, in the you know, in the files stored on the disk to to to fetch your data.

SPEAKER_02

It was really really different, but still but interesting, like how um I think the messaging systems that time still were you know quite popular and they were like widely available. Uh what makes you think that, or like what makes you to bet on Kafka like in such an early stage? What was the uh things that fascinated you or like you said, oh yeah, this is something that we're not getting from anywhere else?

SPEAKER_01

Yeah, I mean uh at that point, I mean that the main thing was really about scalability. You know, we were using I think like RabbitMQ back then, and uh we we were struggling. I mean it was just not possible to you know to keep up with the scale we were operating at. So we had to to rethink our backend, and uh you know Kafka came out more or less at the same time, luckily. So, you know, uh you know, we we always follow like uh what LinkedIn was doing, and uh I mean they are really serious and they were doing like pretty good open source software already before that. So, okay, you know, if they are doing that, it it might be a good bet for us.

SPEAKER_02

Yeah, um there was a I guess before that it was a project uh volunteer mod as a like a distributed database, yeah.

SPEAKER_01

But we haven't used that, you know. We at that time we were like more like a Cassandra shop.

SPEAKER_02

Okay, so um your talk, can you talk a little bit about this, why this topic is important to you and why you decide to uh submit to Kafka Summit? Accepted, and uh uh now we're you know here.

SPEAKER_01

Yeah, so to give you a bit of context, you know, I've I've spent the lot last you know six years like working in a big telco in Switzerland, and there you know uh I brought Kafka in as a main data backbone uh to power like the you know so-called big data platform and all the beta streaming internally. And you know, in this journey, you know, we we went through like different steps, you know, from you know just adopting Kafka, maturing, then we had like a huge security uh slash governance uh part where we have to think about you know role-based access and all these stuff. And when I led the company last year, you know, one of the unsolved challenges uh we we had was about you know tackling it uh GDPR or more like privacy in general, so having like strong encryptions, uh you know, doing like fine-grained access control on top of the data and such things. And uh I've had this idea you know in the back of my mind for quite a long time, and that I took like a few few months off. I mean it was a good time to put together like a prototype. Yeah, and as it worked, you know, I just submitted it to uh to the submit, got accepted, and here I am.

SPEAKER_02

Nice. Okay, so is it uh like an open source project right now, or it's still in a kind of like incubating phase where you kind of prepare in for opening?

SPEAKER_01

It's more like uh incubating phase uh for sure. Uh it works perfectly, but I mean, you know, it's uh I would say it's an alpha software, right? I mean, just to prove the point uh that there will, I mean, a bit of work is needed to make it production ready. But uh, you know, depending on the reaction of the community, I think it makes sense to piggyback on this and to try to do something maybe in the open. Um you know, I I don't know, to be honest. Yeah, let's see what happened.

SPEAKER_02

All right, so let's talk a little bit about detail. So one of the things that uh I've seen that you were talking about is uh how to use the concept of uh like a service mesh into this kind of like a in this kind of uh plane. Um is it something that uh would be like relevant only for like orchestrated environment or anyone who uses just uh like a plain deployment also can benefit from it? So so basically just uh a little bit talk about this because we will definitely recommend people to watch your presentation and once it will be published, it will be available on the Kafkasummit.org uh website. But in general, so for the sake of conversation, can you um provide some of the key details of this implementation?

SPEAKER_01

Sure. So you know you know, uh obviously like this uh sidecar approach works best you know when you use containers because you know you can deploy that you know automatically, usually, you know, if you use Kubernetes, for instance, you can inject containers um quite easily. Yeah, but the cool thing, like you know, it's just like a again, it's just like a daemon you can run and it works everywhere, right? So if you use like a VMs like or like a physical box, you can just deploy the the process next to your Kafka client, yeah, and it will work uh uh out of the box, but on Linux only as of today, right?

SPEAKER_02

So the sidecar um is is a is a process that will be export you know, quitting your the main application somehow, right? So opposed to the ambassador where your application is talking to some service, like for example, local host, and this ambassador provides like external access, right? Um I understand this correctly. So with the um what is the requirement on Kafka broker side of things, you know, to work with this? And what kind of APIs this sidecar will be calling, like, or what's the prerequisite for having this tool um running?

SPEAKER_01

So to be honest, uh more or less nothing. So the the way I did it is like I've really implemented like a layer soft layer 7 proxy for for Kafka. So it supports the Kafka native protocol. And I think this is the pretty cool thing about that that you know basically if you have an application with a client talking to Kafka, you can just drop the proxy on the box uh and it will start to intercept the traffic and send in to Kafka on your behalf, right? Uh without doing much. It just works.

SPEAKER_02

Is it uh written in C or is it a Java? In Go. Okay. Because um there is a uh one of the conversations that community has in the project uh Envoy is to support the Kafka native protocol. And uh so it can be routing all this like Kafka traffic. So and they're doing this in C Envoy written in uh C. Um so I think the idea is also the same.

SPEAKER_01

They they they're working on the Kafka kind of converter for for for for their I think they just landed the PR last week and got merged in uh Envoy. You know, to be honest, I I haven't looked at the details yet, but I mean, you know, might be a good way to move forward. The thing is, like, you know, I did mine like nine months ago, right? So that there were nothing back then. So you know I'd see Envoy might be a better show, better choice.

SPEAKER_02

So, what other aspects of the um dealing with the privacy and uh dealing what what what are other aspects of GDPR that are relevant for Kafka users that you see and you know, you know, or you've seen on previous projects and the things that you implemented right now?

SPEAKER_01

I think there are two aspects which are quite important. The first one is the you know the famous right to be forgotten, right? So whenever you get such requests by customers, you need a way to delete data, uh, which is really tricky in Kafka. You know, as we know, you know, it's like uh you you get like uh initable uh log of events. So unless you you you use like the so-called like compacted topic, I mean you are screwed basically, right? So you have to wait that the event expire. But what I what I've seen a lot you know in big companies is that you know, as Kafka matures and grow in adoption, you know, people tend to store more and more data there and for a longer period of time, right? So I think it's quite common nowadays to have data stored for six months plus in Kafka. So there you cannot wait anymore, right? And I think this is where like you know stuff like crypto shredding you know ships in, where you can basically not physically delete the data but use something uh to cryptographically delete the data from your cluster. This is one aspect. The second one, and it's something which was really tough where I work uh before in the in a big telco in Switzerland, was you know about all this content management. Uh, because one of the strong aspects of Kafka is this uh you know what I call like ingest one reuse many times in different workloads, but uh you know the same data you might or might be not able to use them uh for the the different workloads or purposes you want to do, right? And there usually what I've seen in practice is like you either end up by you know creating topic per purpose, meaning that you have to duplicate all your data by pipelines, not easy to maintain in big enterprise, or you assume that the application you know will apply the whitelist and blacklist.

SPEAKER_02

Uh but you know so in this case you need to dictate uh to application how to deal with this, and there's no control like from the perspective of data, um data group who manages all this data and how this data will be used in application.

SPEAKER_01

Sure, I mean you don't see that. And the second thing, I mean, I'm a developer, so I shouldn't say that, but you know, relying on developers for you know applying such things uh might be good because you know we all write we all write bugs, you know, it could be like a configuring issue, whatever. Yeah. So I mean I think we need a better way to solve this at an enterprise scale.

SPEAKER_02

Right, yeah. Makes a lot of sense. Okay, so um, so apart from the uh decoding some uh open source projects and speaking on the Kafka Summit, what do you like to do like for your uh personal um entertainment or maybe like a to decompress from all this nonsense?

SPEAKER_01

That's a very good question. You know, I I got a daughter like almost two years ago, so I'm pretty busy with her now. I mean it's uh hitting uh a lot of free time, but it's super cool. So this is what I'm doing now, you know, just being uh nice uh dad and you know spending time with uh with her.

SPEAKER_02

So um one of the things that I usually ask um the our our guests in the podcast is to recommend something to read. Usually it's a book. Uh if you read this uh one of the books you read uh recently and it was like uh influential or like it makes some some impact on you, or maybe some paper, or maybe some blog post that um you think that the listeners of this podcast usually can, you know, audience of the Kafka Summit and all these streaming people. Um so uh if you have anything like this in mind, or maybe some interesting open source project that you currently found and you think that maybe um you you can plug it. So anything that uh you think that our listeners can take away.

SPEAKER_01

Oh wow, it's an interesting question. Uh so from a book perspective, not much. But I mean, you know, the the the the technology trend I follow quite a lot nowadays. I mean it's everything happening around service meshes. Uh and I think there, I mean, taking time to look at you know Isteyo, Envoy, Linkerd is really valuable. And uh I spend a lot of time digging in those uh those service meshes uh because I like the fact that you know we need to abstract away stuff from the application nowadays and bring everything back into the infrastructure. So uh I spent a lot of time you know in those areas uh in the past few months.

SPEAKER_02

Alright, so we will um we'll put some links to this like open source projects, and hopefully when you will open source your project, we also will um publish it. So um as always, uh you can find our podcast in um in iTunes. And once you find this there, uh it's a streaming audio uh by Confluent. And uh once you find this, don't forget to rate us, uh, write a comment if you like it, uh, and show your support uh to the show. And David, thanks again for uh your time on first of all speaking at the Kafka Summit. It was awesome. And thank you for your time for joining Streaming Audio. Thank you for having me. And uh as always, uh it was Viktor Gamov, developer of the Kit at Confluent, and as always, have a nice day.

SPEAKER_00

Hey, you know what you get for listening to the end? A Kafka Summit discount code. Kafka Summit is coming up on September 30th and October 1st in downtown San Francisco, and you can get 30% off if you go to Kafka-summit.org and use the discount code AUDIONETINE during checkout. Just enter AUDIO 19 while registering at Kafka-summit.org, and that 30% off is all yours. I'd love to see you there. But hey, I hope this podcast was helpful to you. If you want to discuss it or ask a question, you can always reach out to me at at TL Burgland on Twitter, that's T-L-B-E-R-G-L-U-N-D. Or you can leave a comment on a YouTube video or reach out in our community Slack. There's a Slack signup link in the show notes if you want to register there. And while you're at it, please subscribe to our YouTube channel and to this podcast wherever fine podcasts are sold. And if you subscribe through iTunes, be sure to leave us a review there. That helps other people discover the podcast, which is a good thing. Thanks for your support, and we'll see you next time.