Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov

Connecting Snowflake and Apache Kafka ft. Isaac Kunen

Confluent, original creators of Apache Kafka® Season 1 Episode 101

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 31:46

Isaac Kunen (Senior Product Manager, Snowflake) and Tim Berglund (Senior Director of Developer Advocacy, Confluent) practice social distancing by meeting up in the virtual studio to discuss all things Apache Kafka® and Kafka Connect at Snowflake. 

Isaac shares what Snowflake is, what it accomplishes, and his experience with developing connectors. The pair discuss the Snowflake Kafka Connector and some of the unique challenges and adaptations it has had to undergo, as well as the interesting history behind the connector. 

In addition, Isaac talks about how they’re taking on event streaming at Snowflake by implementing the Kafka connector and what he hopes to see in the future with Kafka releases. 

EPISODE LINKS

SEASON 2
Hosted by Tim Berglund, Adi Polak and Viktor Gamov
Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
Music by Coastal Kites 
Artwork by Phil Vo 

  •  🎧 Subscribe to Confluent Developer wherever you listen to podcasts. 
  • ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
  • 👍 If you enjoyed this, please leave us a rating. 
  • 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
SPEAKER_00

You know, I bet you've heard of the cloud analytics platform Snowflake, but do you know how it integrates with Kafka? Isaac Kunan, a product manager at Snowflake, is going to tell us all about that on today's episode of Streaming Audio, a podcast about Kafka, Confluent, and the cloud. Hello and welcome back to another episode of Streaming Audio. I am your host, Tim Berglund, and I'm joined today in the virtual studio with all of the social distancing you could possibly ever need by my guest, Isaac Kunan. Isaac, welcome to the show.

SPEAKER_01

Hi, Tim. And I will confirm I cannot see or touch you.

SPEAKER_00

Yeah, yeah. I don't even know. Where are you located, Isaac, right now? I'm in Seattle, sort of in the north. North end of Seattle. Yep. Okay. Secure bunker in your panic room in Seattle. That's good. I am in my basement in suburban Denver, Colorado. And uh these things seem to be working out for the both of us. It doesn't matter much where we are anymore. Really doesn't. It really does. Through the magic of Zencaster and podcasting, we can make this work. Isaac, we are going to talk about Snowflake and Kafka and the integration. We're going to dig into some Kafka Connect things today. I would love, as always, to hear about you first. How did you get to do what you do? What is it that you do? And uh how how does one get there?

SPEAKER_01

Yeah, so I'm a product manager at Snowflake. I've been working in the data industry for a long time. So I I uh started off really doing data work with SQL Server back many years ago. Uh I was there for a few years and uh also worked on a another data product at Microsoft called Stream Insight. Uh made my way through Tableau where I built uh Tableau Prep and then off to Snowflake where I've been working on various things, mostly in the connector space. So both our Spark and Kafka connectors and then also some extensibility work that we've got sort of uh in the in in in the works.

SPEAKER_00

Cool. So um I'm going to assume that everybody listening to this, if you want to listen to a pod uh Kafka podcast, you've heard of Snowflake and you probably know, yeah, cloud analytics thing, right? Um maybe give us a little something a little bit better than that. Um what what in your opinion uh is the minimum a developer or architect should know uh about Snowflake to be able to think about it as a component, as a tool in the toolbox, a component in the cloud ecosystem. Right. So I might that's a very long way of asking what's Snowflake.

SPEAKER_01

What is Snowflake? Yeah, so I I might I might give you a three-part answer. So the at the at the simplest, you can think of Snowflake as really almost like any other SQL data warehouse that you can connect to. So if you're used to using another data warehouse, uh Snowflake will look superficially very similar to you, which is great because you know, obviously SQL has been around since you know prehistoric times, the 1970s. And uh so being able to integrate with that with that ecosystem is I think a core a core strength of ours.

SPEAKER_00

And I might say the name evokes the the classical data warehouse.

SPEAKER_01

That's that's right. You know, I I get you know when you say you work for Snowflake, people give you funny looks, but really from a from a database perspective, that's a pretty pretty standard term. Plus, our founder is very much like skiing. So I think that that also fed into this.

SPEAKER_00

There you go.

SPEAKER_01

Um so so that's sort of the most superficial layer. But of course, uh we like to say that Snowflake is really the data warehouse that was built for the cloud. And so if you dig down just a little bit, you you find that our architecture is really very substantially different than anything that's out there. We really enable our customers to scale their data uh entirely independently from their compute. And this this really enables some real magic. I've worked on systems before where you know I could run my query before about seven in the morning, and then I couldn't after that because other people were playing with it. But with Snowflake, I can just always spin up my own compute and be very isolated from everyone else that's doing their work. So so we really scale almost infinitely on the compute side, and likewise on the data side, we really scale uh very independently and very, very well. So I mean we have customers that have single tables that have over a petabyte of data, and that just works, works fine. And then maybe the third answer to that would be that we're we have really started to build beyond our um sort of traditional data warehouse routes, and we like to talk about now being sort of the data platform for the cloud. And so we've we've we've started to really extend ourselves in terms of the kind of uh integrations that we do with things like Kafka, right? So people building apps uh can integrate with that that nice streaming solution that we're here to talk about today. Uh we have data sharing and uh data marketplace uh or sort of ways to share data both internally out uh and externally. Uh lots of sort of features that are moving beyond that that core data warehouse route so become a much more complete data platform.

SPEAKER_00

How do people generally see the data that is in Snowflake? I realize that's well after any sort of streaming components that we're gonna talk about, but um what what's what's a typical visualization front end?

SPEAKER_01

Well, yeah, so I mean obviously the the the easiest way you can get at your data, the maybe the most basic way you get at your data is we have our own worksheets where you can log in and write SQL and see the results. But we have a huge array of partners, uh, everyone from Tableau, where I used to work, uh, that that is a very common front end to products like Looker or Power BI. So all of these tools that people use to do BI or visualizations very naturally plug into Snowflake. And so you really have your choice.

SPEAKER_00

Cool. So any of the uh one might say the usual suspects, including a native interface of your own.

SPEAKER_01

Yeah, and maybe even some less usual suspects. But cool. Uh I am not I'm not maybe the expert there. But yeah, we've got we've got quite quite a number of tools that plug in and work with Snowflake.

SPEAKER_00

So the the um there's a question that's so softball, I I won't even ask it. You know, why would you integrate Kafka with Snowflake? And I we I think we get that, right? There's data in Kafka, there's various kinds of real-time analysis that the Kafka platform is good at doing, things in Kafka streams, you can do computations in Kafka streams, you can do queries in KSQL that that do you know that that perform computations uh that you know, right? When you have worked out the query in advance and you want something uh to be calculating that result on a message-by-message basis, like the the Kafka ecosystem understands that um and and performs very well at that. But then when you don't know what the query is and we need you to explore, and I it by the way, you're welcome to expand on my my uh analytics meta use case. Yeah, no, I would I would just kind of how I see things.

SPEAKER_01

Yeah, so from the Snowflake perspective, is streaming data is um is becoming certainly much more important, and we see Kafka as probably the most common. I wouldn't even say probably, it is certainly the most common streaming solution that we see. And so integrating there becomes very important for us. And of course, all the transformations along the way that you're talking about, uh Kafka really excels at. From Snowflake's perspective, you know, we we really uh I would say excel at, yeah, those ad hoc queries, they can queries too, that's fine. Um, over your your huge store of that persisted data. Like I said, you know, we have customers with single tables in the petabite range, and we we scale very well, handle that. And so it's a fairly natural, I think, marriage to say, hey, you know, that streaming data, as it's as it's in flight, you want to transform it, you want to extract some value. But I want to land that so I can do historical queries on it, for example. And where do I land that? And where do I run those? And that's Snowflake. So it's a pretty it's a pretty good uh sort of mesh of technologies.

SPEAKER_00

Got it. Tell me, so I I assume the uh the integration between Kafka and Snowflake is a Kafka Connect connector, yes? Correct. That's right. Nailed it. All right, good. Now, uh, we will put a link in the show notes to some shows on Kinect. Uh we did a lot of them in the summer of 2019. Well, spent a lot of time just talking about Connect from various angles. So there's plenty of stuff in the back catalog. If you don't know what Kinect is, but I'm gonna assume that you basically do, um, and Isaac and I are gonna talk about it. Um yeah, so tell me about the the process there. Um and I'm gonna I'm gonna assume you were you were PM of this connector and this whole interface. Correct, yeah. So how did it come to be? How did you make the case? Tell us the history of the connector.

SPEAKER_01

Uh you know, I think that one there's not much of a story to tell. Uh the the case sort of makes itself. You have enough customers.

SPEAKER_00

Why is there not a connector right now?

SPEAKER_01

Yeah, why is there not a connector? How do I make this easier? Uh we had a lot of customers that were going sort of through their own um machinations to get their data from Kafka into Snowflake. The most common route would be to use an S3 connector uh and then and then trigger ingest in in one of our various ways. But um, you know, there's some complexity there, especially if you want sort of reliable ingestion, you know, exactly once delivery of your data is is not the easiest thing to engineer with you know Kafka and our ingestion, sort of marrying those two. And so by building the connector, we're able to provide a lot more uh simple way to get the data in. And that's that's really what our customers were looking for with this.

SPEAKER_00

What was interesting about the process? Did um like did the engineering team learn anything interesting about Connect or the interface? Were there any special challenges that came up that were that would be fun to talk about?

SPEAKER_01

Um good question. Uh you know, I think we learned a lot about how our own ingestion story works, and and so that's that's always interesting to sort of dog food your own interface there and figure out how to make it work. Um really puts yourself in the customer in the customer's shoes. And and I think we learned a lot about how, like I said, our own ingestion APIs always work. As for the the connect APIs, uh I think it's been interesting to just sort of sort out how they work, but I don't I don't know if there's a whole lot to say there. In some sense, they're fairly straightforward um in terms of of spinning up the connector and uh yeah, maybe some challenges and making it work work well with our situation. That is, I think the the connect framework assumes certain things about startup and shutdown that that uh don't always hold with with R and Jess story. So dovetailing those two was was somewhat challenging, but not not rocket science, I wouldn't say.

SPEAKER_00

Sure. I mean it is ultimately it's uh it seems like it ought to be pretty easy. And connect is not as APIs go, and I always tell people when I'm introducing them to Connect, like you don't want to write S3 integration code. Okay, you just don't want to do that. You don't want to write the code that reads data from a JDBC database data source. You know, this is not uh why you've been brought upon this earth. That code has been written, use it. But when you do come up with an interface that's not a connector that exists you know, when you have to write a connector, it's not a bad API. No, it's fairly straightforward. It's not it's not that much to it. There's not that much to it, exactly. Yeah. Uh which is always the thing when we talk about connect that that it it it is one of those things that on the surface seems like couldn't wouldn't you just write that code? Like why not just write your own consumer or producer or whatever and and just make that happen? It's deceptively simple, right?

SPEAKER_01

Yeah, it's a little deceptively simple. Like I said, you know, I think the big the big challenge is in giving those guarantees about reliable ingestion. And that's not so much about the Connect API as it is about dovetailing that with the APIs that we have. Right, right.

SPEAKER_00

And why that and not just JDBC? I'm assuming Snowflake is you know bristling with JDBC functionality on its its exterior. Uh so why not that?

SPEAKER_01

Right. I you know it's fairly typical if you if you look at most databases, they have their own very optimized ingestion route. Going running data into a database via JDBC is pretty inefficient. And you know, in particular with Snowflake, adding single records at a time is not uh is not an efficient way to get data in today. Uh maybe someday, but not right now. And so we have our own bulk ingestion facility, which really lets you write a batch of records to a file store like S3, and then ingest those in a much more efficient way. So you you really want to go through that, or your performance is going to be poor.

SPEAKER_00

To batch. And batching uh but you said some people had kind of hacked this with with S3, which I assume would be a batched solution also. Um That's right.

SPEAKER_01

And so you can you can do that, and in fact, that's really at some level what our connector does is nothing that a user can't do themselves. So we've actually added no core Snowflake code to build the connector. It's all built using the APIs that we expose already. Nice. And so we have functionality, uh, we call it Snowpipe, that lets you do what we call continuous, you should really think of as sort of uh ongoing batches of loading of data into into Snowflake. And we do that by by uh using S3 and then triggering these ingests through a REST API. Uh you can also do what's called auto-ingest, where you just drop files in a bucket and we will we will do the ingest from there.

SPEAKER_00

You'll find the bucket.

SPEAKER_01

But yeah, that's right. Um we'll we'll we'll listen for notifications on the bucket and pick files up as they as they're dropped, but you can you can trigger this either that way or through through a REST API, and that's what we use in the connectors, is that REST API. So, yeah, I mean before the connector, what customers would do is they would drop files in an S3 bucket and then either rely on auto auto-ingest, which um for S3 was not always reliable, but is now, right? Those notifications uh from S3, we there were some issues there, but those have now been resolved. So that's uh it is a reliable way at least to pick the data up. Um or they'd have to go and trigger trigger that ingest themselves, which is programmatically um sort of a pain. Uh and then of course what you have to do is is make sure that you monitor that ingestion and you know advance your offsets in Kafka correctly once things are ingested. And you have to worry about what happens if you know either Kafka Connect or Snowflake were to go down and come back up. And how do I make sure that my records got in? How do I make sure that they only get in once, et cetera? So there's a bunch of bookkeeping that one needs to do on both the Snowflake and the Kafka side to keep these things aligned, which is tricky. And so uh to be honest, I'm not sure I've ever seen a customer that really invested in getting that logic right. And so people doing this through S3 would ingest their data and it performs fine, but would generally sort of swallow the fact that that reliable the reliability of that ingest was not always perfect.

SPEAKER_00

Well, you just said some really important words there that you've never seen a customer invest in getting that logic right. I I I might be getting the cat quote wrong. Uh that's right.

SPEAKER_01

And it's because it's tricky. Uh we went into this thinking it would be easy. And uh it ended up taking us a bit longer than we thought it would take to get that done. And so now we've done it. Uh, the nice thing about us doing it is that now all of our customers can go and just pick this up and get that logic once rather than having to reinvent the wheel over and over.

SPEAKER_00

Exactly. And that is such a great description. And this is the thing that fascinates me about Connect. I I don't get the question often, but I I have gotten it in the past from some what you know appear to be fairly junior developers. You know, why would I use this? Why wouldn't I just write this? And the problem is that the the block diagram, like if you have to whiteboard what you're talking about to get someone to understand what you mean when the task is to get data from a Kafka topic into Snowflake, it doesn't take you a long time to draw the diagram. And it does not take you a long time to get everybody in the room to understand exactly what you mean. It's so simple. There's data in a topic, and I want it to be in a table in Snowflake. Oh, okay. You know, so like the concept is there, it's not like some fiendishly complex financial payment settlement mechanism or no, that's right, you know, real estate transaction modeling or some something that you really have to bend your mind to understand the business process. You just don't. You got some stuff here and you want to put it into that thing there. But the particulars of it are really hard.

SPEAKER_01

That's right. And when you start asking questions like, like I said, what happens if Kafka goes down or Kafka Connect goes down, or Snowflake errors out, or there was an error in the data, and what do I do with that if I don't want to lose it? Or how do I scale this? Right? All of these questions start to complicate the picture. And I think we've got a fairly good sort of answer to all of these with the connector, but it isn't trivial. Um, so yeah, I mean, some developers might might sit there and say, Yeah, can I can I just build this? And of course the answer is yes. But I can't remember the exact the exact quote here, but what is it the you know, uh good developers borrow and great developers steal.

SPEAKER_00

Yeah.

SPEAKER_01

And here we're saying, yeah, please steal steal the whole thing.

SPEAKER_00

And even better developers, you know, just figure out a way not to write the code. But yeah, uh that's right. And this this is a case for stealing where it's or using the dang component because actually getting all that stuff right, like you said, offsets and failure modes and all those cases, um, that's that's hard to do.

SPEAKER_01

Uh that's right. By the way, and on the on the on the point of stealing, we should just mention that the connector is actually up on GitHub. It's all open sourced. So yeah, please go steal.

SPEAKER_00

What's the the license? Do you know off the top of your head? It's Apache. Awesome. Okay, well, that's wonderful. So, you know, it's not stealing if you have given it away.

SPEAKER_01

I guess that's right. Yeah. Uh my my father always said it tastes better when you steal it as he was taking my food. So I like to, you know, go go ahead and steal it, it'll taste better. But um, yeah, no, we're we've really given it away.

SPEAKER_00

He is as, as it turns out, not the first person to have said that. Reminds me of the um Apple store buy the item on the app thing. Of course, this is all notional right now because uh there's that's not a retail function that is available at the time of this recording. Uh, but you know, you can you can on the the Apple store app walk in and like buy uh you know a cable or something like that just by scanning it, paying for with Apple Pay and walking out, not having spoken to any human being. Feels like stealing. So it's like all of the you know, the the naughtiness without the actual unethical behavior.

SPEAKER_01

Uh yeah, maybe maybe they sell more because it feels good to just take stuff. I don't know.

SPEAKER_00

I didn't talk to anybody. I just walked out with a pocket. I absolutely paid for it. I didn't steal property. That's a bad thing. But yeah, so if if uh somehow cloning the repo makes you feel that way, uh it shouldn't. But hey, good. That's good to know that that's a little code up there.

SPEAKER_01

Yeah, and send us send us pull requests. You know, if you've got if you've got items that you you would like in there and are are are willing to contribute, we're happy to take those too.

SPEAKER_00

Nice. Talk to us about some and and actually let me not let that comment go too easily. That's amazing. Uh that's great that you guys are open to that. So please, yes, if you're using this uh and you want to change something, drop a PR. We like that.

SPEAKER_01

For for what it's worth, this is this is our standard pattern with with uh our connectors and our drivers for the most part. Not all of them, but most of them are actually up in uh up in GitHub and available. And yeah, this is this is our common pattern.

SPEAKER_00

Tell us about some maybe some of the more interesting pipelines you've seen. Uh again, one of the struggles here is that that conceptually it's so simple. There's there's topics and there are things, and then I want to do analysis on them in an ad hoc performant way in the cloud. And so I use the Snowflake connector and I get it into Snowflake. What are some of the more interesting use cases that you're able to talk about that you've seen?

SPEAKER_01

Well, I'll just say the the most general pattern we see, which I think is a very uh powerful one, is that we'll have customers who run whatever events. And it's uh it's a I I I'm using events here to be quite broad, right? These could be click logs or you know, IoT device events, or you know, really almost anything. Um so much of our data is these kinds of events, the these events, and you know, Kafka is a is a great tool for flowing those around. Um so we have customers that that will drop those, and then they want to do some sort of post-processing on them in Snowflake. So we've got a very nice link with uh two features in Snowflake that work well together. So it's sort of three features that work well together is uh Kafka uh ingestion and then and then streams and tasks. And so streams are our way in Snowflake to capture, you should think of it as uh CDC, right? It's a way to capture uh diffs on a table and then act on them. And then tasks are a way to schedule that. So what we have is a lot of customers that will take some sort of uh event data, uh flow it into a table, and we ingest it as what we call a variant. It's just sort of we represent the structure that's in that Kafka data. But then they want to, for example, do something simple like flatten that out into a table. And so you use streams to capture the diffs and then put a task that may run every minute or every five minutes, whatever your latency requirements are, to pull that data and parse that out and spray that into a table. So that's a very common pattern uh that we see people using uh over and over. In terms of the the most Interesting things. I think the hardest, maybe I I would say the hardest ones for us to deal with are we have uh some customers that are really using Kafka as a medium for doing uh sort of database replication. And uh I'm not sure we have a great solution for this when the number of tables that people are replicating get very large. So they'll do things like like uh you have a database with 20,000 tables and figure out a way to shove it through a Kafka pipeline. And those are those are challenging. I'm I'm not sure we like I said, I'm not sure we've got a great recommendation of how to how to deal with those today. Uh if you've got a few a few of these that are flowing, it's pretty easy to deal with. But uh those are challenging.

SPEAKER_00

What makes that a challenge? The I mean the large number of tables, but where where in that does it?

SPEAKER_01

Uh right, yeah. So so how do you do this? You you either have to run each table, for example, in its own topic, which uh will work, but you end up with an awful lot of topics, and Kafka, I think, has a hard time dealing when the number of topics gets too large or the number of partition gets too gets too large. There's some fixed overhead in Kafka for dealing with that. And so what users tend to do in these kinds of situations is they multiplex them into a single um into a single topic. So I've now got to demux them on the other side, and that demuxing is a is a is a challenge.

SPEAKER_00

Gotcha. You'd need you'd need some sort of correlation identifier to indicate the the source topic.

SPEAKER_01

Yeah, and you and they all have different schemas, and how do I write the code to deal with this? It's not a it's not a it's not an easy task.

SPEAKER_00

Is there anything in the connector? Like this would be a uh single message transform sort of thing, but does the connector have any features or SMTs or anything like that that have grown up around it to help accommodate some of those interesting corner cases?

SPEAKER_01

Yeah, so SMTs, I'm not sure how SMTs would really help too much here because you you still on our side end up with single.

SPEAKER_00

You still have a single right. It would it would help on the input.

SPEAKER_01

Yeah, it might help on the input. We actually don't support SMTs in the connector right now, and we're working to try and get those in anyway. So excellent. So that's what yeah, yeah, absolutely. No, it's it's in flight, actually. And so we're working uh actually with the the Kafka folks to try and to try and get that online. But uh yeah, no, it's it's a it's a tricky problem, right? This this is the system didn't really grow up around this use case. So it's sort of stretching it a bit and it's uncomfortable.

SPEAKER_00

Right, right. And um I guess a uh word of precision on the uh partition thing. Current versions of Kafka, you should be able to, it's a it's a number of partitions per broker. You should be able to get about 20,000 partitions per broker. So if you had 20,000 topics and they all had to be 100 partition topics, for example, I mean that that that's unlikely. Uh those would be really big topics. That's just a super big cluster, right? That's like a hundred or two hundred node clusters. That's right.

SPEAKER_01

And you might be surprised as a lot of real data warehouses have many tens of thousands of tables. Uh you know, I've seen I've seen data warehouses that have in the hundreds of thousands of tables. Um and so it's it's sort of shock shock shockingly large. Most of those are fairly quiescent, right? They're lookup tables of various sorts. Sure. But um some of them are going to be very active.

SPEAKER_00

As long as those tables are named in with a six-character naming scheme uh with abbreviations, then that should be okay, right? Of course. Yes, we would hope. Yes. Um only only the most humanizing schemas that that that respect the dignity of the developer uh who will be using them. That's our goal here. Yeah, sure. No, so for that that end of the spectrum, that's a lot. That's still uh you still totally do that, uh depending on you know the number of partitions. Like you said, for a lot of those being probably small reference data e-type things, you could make that happen. But I still see that that integration problem as an interesting one if you're muxing them into one topic.

SPEAKER_01

Yeah, it's an interesting one. Like I said, that's that's an outlier. Most of our most of our uh customers are flowing, you know, it's it's web logs or it's yeah, like I said, sensor data that they're flowing in. And these these work really well. So I think we're we're dwelling on maybe maybe an outlier, although it's one that's come up a few times that's interesting.

SPEAKER_00

Yeah. Well, I mean the outliers are fun because they illustrate where the costs are in the system and you know where the kind of growth edges of the technology are gonna be. Because if they're hard to do, well then that's the thing that you need to build next.

SPEAKER_01

Aaron Powell Yeah, they're also the ones, of course, that we hear about the uh at some level the most, because it's when customers pipe up and scream.

SPEAKER_00

Well, that's the thing. We should do an episode on the psychology of being a product manager at some point. It's not gonna be this episode, but it's one of those jobs where no matter what, every decision you make is wrong to like a significant minority or plurality of your constituency. So, yes, you you hear about these things because you're a PM.

SPEAKER_01

That's right. And you know, uh this is one that we've heard about a few times, but it's also one that we've not terribly prioritized because, like I said, you know, the the typical use cases are the ones that we don't hear about, and those actually seem to be working pretty well. Nice.

SPEAKER_00

Um on that point, um, where does the connector go in the future? Or I a little more broadly, um, what are you excited about that you see in the future of Kafka Snowflake integration and deployments? And uh what what do you want to see happen there?

SPEAKER_01

Well, I'll say the the most immediate thing is we're working hard to get the connector up in the Confluent Cloud. And I think that we've we've that actually should be going live relatively soon. So that's that's something we've heard from a number of customers that that um you know they're using Confluent Cloud and they want the connector. So so yeah, that's that's maybe the the the baby step. Uh longer term, yeah, I think the biggest sort of um oh let's say the the largest request that we have is for writing to to streams as well, or writing to writing to Kafka topics rather. So um you know, right today uh this is entirely what what Kinect would call a uh sync. You run your data from Kafka into Snowflake. But as we talk about being a real data platform and people are building more of their apps on Snowflake, they want to do things like monitor what's going on. We talked about this notion of streams. Uh so a stream I can run on any any table and capture changes, and I might want to flow those back out into a Kafka topic so that I can pipe them to some other application and deal with them there. So that's a pretty common request is how do we how do we write out? And so from the Kinect perspective, how do we become a good source?

SPEAKER_00

I love that. And that reflects, I think, a certain maturity of the architectures that are participating in this Kafka snowflake integration because the I have events and I need to get them somewhere so I can do valuable things with them is is Kafka thinking of a few years ago. I have events and events are central to my thinking about my system, and those events need to be in Kafka topics, and they need to be in other places that are better at doing certain kinds of things with them, like say analytics. So they need to be in Snowflake, but then the results of those analytics, well, dang it, they need to be back in Kafka topics because there's some other microservice that needs to consume them. Uh, and so I like that desire for uh source connector.

SPEAKER_01

Yeah, that's spot on. There's if if you view Snowflake as sort of a dead end, uh, and not a dead end in that you're not going to use it, but you know, the I'm gonna land my data there, and then I have consumers that you know come in through Tableau or something like that, but it's really in some sense it's a dead end for your data, then a sync connector is all you need. But Snowflake is not usually a dead end. Uh people are going to run analytics, they are going to take the results of that and want to do something with it in a programmatic way. And so how do I then do that? And uh yeah, if you're if you're a customer that has invested in Kafka, probably a key way that you want to do that is by running things back out through Kafka.

SPEAKER_00

That's such a great point. I I um you know when when I was thinking when I said that and I'm thinking about the architecture, I'm thinking, well, uh gee, good, that's a more mature use of Kafka, but really is a more mature use of Snowflake, too, if it's in the loop, if there's more computation that happens after it adds its value and it's not just a dashboard. That's right. My guest today has been Isaac Kunan. Isaac, thanks for being a part of Streaming Audio.

SPEAKER_01

Absolutely. Thank you very much.

SPEAKER_00

And there you have it. I hope this podcast was helpful to you. If you want to discuss it or ask a question, you can always reach out to me at TL Burgland on Twitter. That's at T L B-E-R-G-L-U-N-D. Or you can leave a comment on a YouTube video or reach out in Community Slack. There's a Slack sign-up link in the show notes if you want to register there. And while you're at it, please subscribe to our YouTube channel and to this podcast wherever fine podcasts are sold. And if you subscribe through iTunes, be sure to leave us a review there. That helps other people discover the podcast, which we think is a good thing. So, thanks for your support, and we'll see you next time.