Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov

Turning Chaos into Push-Button Provisioning with Dhiraj Suri| Ep. 14

Confluent Season 2 Episode 14

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 21:13

Viktor Gamov talks to Dhiraj Suri (Confluent) about his career in systems engineering and stream governance. Dhiraj’s first job: software developer at NetApp. His challenge: working at Splunk to stitch together disparate systems into an event-driven provisioning platform.

SEASON 2
Hosted by Tim Berglund, Adi Polak and Viktor Gamov
Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
Music by Coastal Kites 
Artwork by Phil Vo 

  •  🎧 Subscribe to Confluent Developer wherever you listen to podcasts. 
  • ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
  • 👍 If you enjoyed this, please leave us a rating. 
  • 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
SPEAKER_03

From building system software for hardware appliances in C to building data governance software in cloud. This is Confluent Developer.

SPEAKER_02

And we had a very uh serious problem. This was one of the projects that made me realize how different systems need to be able to communicate with each other. We did look at that and realize that that still was not enough.

SPEAKER_03

Hello and welcome to this episode of Confluent Developer. My name is Virgo Gamov and I will be your host today. And with me in our uh real studio today, uh my colleague from Confluent, Dirac Suri. Welcome to Confluent Developer. Thank you. So, Diraj, um, can you uh tell us our listeners and our viewers uh what do you do with Confluent? And uh, you know, uh it would be great to learn a little bit more about your current uh current work.

SPEAKER_02

I'm a I'm a staff software engineer on the Stream Governance Platform team. And our uh the primary uh focus of our team is a product called the Stream Governance Platform, which um essentially allows um users of customers of Confluent to manage their schemas and and you know make sure that the data that they have that flows in from disparate sources is compatible uh across their ecosystem.

SPEAKER_03

Yeah, so talking about schemas and schema registry is probably my second favorite subject to talk. Uh apart from talking about Flink and talking about Kafka. I um I start um kind of introducing uh schemas for polyglot developers, usually kind of like a Java developer, just like, okay, well, whatever. We're gonna do whatever we have. Uh maybe even Java serialization, maybe Jason, but when you start dealing with um the systems that written in different languages, you need to come up with the cross-platform protocol to talk. Yeah. Um so in uh in the conflict developer podcast, what we like to do, we we're exploring origins of technical excellence. I I like like how I like to put this, meaning that we are chatting with the uh practitioners and the people who've been around and building some of the um complex systems, and we would like to learn like what they learn over the time. And essentially we like to schedule structure everything around uh two concepts. One of the concepts is kind of like what was your first job, how you start understanding that your skills can actually you can you can do something with those. And if uh if you have an example, how you end up in technology, like for example, what was your first language that you developed, like your software, or maybe something that you developed and you saw this, that would be an interesting thing to learn.

SPEAKER_02

Yeah, definitely. Uh my first job right out of college was as a software developer at a company called NetApp. Um it's still obviously pretty big. It's still there, yeah, of course, it's still there. Uh, but um, you know, back then it was the cloud revolution was was kicking off, you know, into high gear. I was, I think in a way back then I thought I was unlucky, but looking back, I was very lucky to start my career uh writing code in C and C. Because I think for a computer, any computer science major, that's the perfect language to take off and you know have your career start off in. Uh uh, because uh it's about as close as you can get, well, at least in the past couple of decades, that you can get to how a computer actually works rather than uh a lot of the languages that are around today. So yeah, um uh I was writing code that runs um you know storage software on the actual NetApp uh you know appliances. Appliance, the actual appliance. Yeah. Um and that's when I I I started understanding um, you know, like the scale of uh you know, the power of the cloud, really. Yeah. Um and NetApp was right on the cusp of you know being being a company in the pre-cloud era, transitioning into the cloud era. Yeah and that gave me the perfect opportunity to understand um and actually really see what the next couple of decades uh were going to look like. And and you know, and here we are.

SPEAKER_03

Yeah. So um you think the is um is C still good language for running cloud applications? Like if you develop a cloud platform.

SPEAKER_02

Yeah, yeah, certainly. That that's a that's a great question. Actually, after my my second job right after NetApp was actually for a video ad serving platform company called Brightroll, it got acquired by Yahoo. And we were running the entire ad serving stack in C, in web scale C like billions of transactions a day or requests a day, really. And um we couldn't have done it without C. I don't think the company in that in the form that it existed.

SPEAKER_03

So you think you think that the Java would not be able to uh uh to do uh that's you know because Java, it yeah, mean Java, it's it's it's not a thing.

SPEAKER_02

For those margins to work out, yeah. We in fact uh one of the main things I did in my two and a half years there was to rewrite just one component of the ad serving application, ad serving stack from PHP into C. I approved that.

SPEAKER_03

This is this is yeah this is very honorable uh work to do. Um and it turns out that the PHP actually it's not uh uh that bad. You can actually compile PHP to C with the hip hop and things like that on the Facebook, right?

SPEAKER_02

We Yes, we we did we did look at that um and realize that that still was not enough uh to get us to the margins and the scalability that we needed to.

SPEAKER_03

Okay. So um uh so now uh what what um um when I talk to people who start their like a career from learning like a serious like a system level language, uh rest of the kind of like world for them is kind of walk in the park, you know. Learn high-level language like Java or like learn something like or.NET would be much easier and kind of transition to the different world would be. So, what language do you use today?

SPEAKER_02

Uh well, schema registry, um, a lot of schema registry is written in Java, as uh folks listening to this would know. Um, yeah, so I I do use a lot of Java and then uh Golang obviously are a lot of uh for control plane, yeah. Control plane.

SPEAKER_03

Like things, I think that you would agree that uh with uh emergence of the Golang and Kubernetes, the the shift for like cloud software, system software kind of like shift shifted, like things shifted into the world where the go would be kind of a de facto kind of standard for right now.

SPEAKER_02

Basically it's become yeah, it's basically become the de facto standard. So yeah, uh to answer your question, yeah, I do find it. I I I I think the inverse would have been much harder than uh than going from C to Golang or C to Java. Yeah, so I do find that I don't find the complexity of the language to be a barrier for me data to day.

SPEAKER_03

Every time when we're talking about this uh new or cloud uh languages, we cannot not talk about Rust. Yes. What's your thoughts about this? And uh do you think uh it is a great tool for writing? Like since you've already been in the cloud space and touch different like a system components.

SPEAKER_02

Yeah, I I do I do see Rust being used for a lot of applications. Um, you know, uh but but I I as far as I've seen, it's still somewhat specialized. It's it hasn't broken through to becoming uh you know general purpose in the same way. Yes, exactly. That's my point.

SPEAKER_03

The same with Golang or or well, obviously Java, but Golang has become so the one of the one of the things that uh we also would like to discuss with our guests, and I think the you have uh um great uh years of experience in in this space because uh because you came from the cloud companies. What was the the most challenging thing that you ever done? Um it can be, you know, the before Confluent, in Confluent, like what's the kind of like a very interesting technical challenge uh that you solved or you didn't, you know, and you're proud or you're not uh to talk about.

SPEAKER_00

Now a quick word from our sponsor. Confluent developer the podcast is brought to you by Confluent Developer the website, which has everything you need as a developer of data streaming systems. And it's completely free. We've got curriculum, hands-on exercises, executable tutorials, the online data streaming engineer certification, also free, a way to find a meetup near you, those are free. Everything is there. I really want you to be successful in your journey as a data streaming engineer, and this is the site that has what you need. Check it out at developer.confluent.io. That's developer.confluent.io. Now back to the show.

SPEAKER_02

Yeah, yeah, I think one of the one of the most interesting challenges that I've had to deal with, um, I think in terms of scale and impact on business, I think when I look back is is uh at my previous role at Splunk, we had a very challenging problem of uh needing to be Splunk is an enterprise software company, much like uh Confluent, and we had a very uh serious problem uh with needing to provision customers quickly, uh large enterprise customers very quickly, and giving them their Splunk deployment um, you know, in a reasonable amount of time. Uh and the barrier to doing that was the fact that um we had disparate systems across engineering and IT engineering, you know, software engineering, which is you know our organization. And we had to stitch all of these different systems together in order to produce a coherent outcome for the customer. And that's when, you know, that's where my team, you know, my expertise, uh, having been at the organization for several years, uh, that that came into play. And I was able to, we were able to architect a solution that put together uh, you know, systems like Salesforce using API integrations like MuleSoft and then other message queue systems uh that essentially uh took something that was seemingly non-deterministic, that required a lot of manual touch points into what eventually became a practically a push-button provisioning system. And that's what really that's what really um you know uh uh showed me the importance of you know having asynchronicity in these large, you know, very complicated enterprise software systems. Yes.

SPEAKER_03

Can you touch a little bit about like architecture of the solution? Uh, because our listeners would also appreciate kind of like learning some of the technical details. I guess maybe not super deep if it's uh you know uh some proprietary secrets, it's it's totally fine. But in general, like let's imagine um you need to architect the you know, I would say that it's kind of self-service portal for for customers where they can provision different environments and stuff like that, right?

SPEAKER_02

So yes, that's pretty much it. Yes. Uh I think um in terms of really from a high level, what the solution involves, you know, what a solution like that would involve, not without getting into the specifics of that, is you know, you you want you want asynchronicity and really event-driven processing. Um and really in a in a way, uh, this is really about streaming data from customers, right? That the inputs that customers give us across different you know databases, like you know, we were effectively, you know, Salesforce effectively functions as a database. And from that, you want a queuing system so that you know their back pressure doesn't actually lead to a backup that users feel. Um so yeah, we had MuleSoft, we had Salesforce, uh, we had a lot of AWS. Uh we weren't using Kafka yet, but we had SQS, SQS, SQS and SNS a lot. And then of course, um Bing Splunk, we use Splunk itself for a lot for observability.

SPEAKER_03

So how how you think the the you know the now you're working with the schemas and how your you know knowledge from the you know the technologies from the past uh translated to technology that you're working on this today? Like what uh kind of experience you bring in into building data governance tools?

SPEAKER_02

Um Yeah, yeah, actually this was one of the projects that made me realize how different systems need to be able to communicate with each other, right?

SPEAKER_03

Um, you know, it's just because another it's Conway's law, you know, just because uh different uh departments will be uh kind of like uh correspond or like will be um mirror software that will be implementing.

SPEAKER_02

Correct. Conway's law is sometimes quoted as a good thing, but in my opinion, it's not a good thing.

SPEAKER_03

Depends on the organization, like how everything is established.

SPEAKER_02

You might need to work against Conway's law, right? You you will you don't want disparate systems to be built and then to reside within disparate organizations. You want to break those barriers as for the company to grow and serve customers better. And and I think something like schema reg schema management, it becomes very important when you have APIs that are designed by different engineering paradigms, in fact. Forget languages, forget uh services, right? Different engineering paradigms. And that's where you need API contracts. Um and that's one of the things that we I was very uh keen on enforcing and really evangelizing within the company on what REST, you know, yeah, exactly.

SPEAKER_03

Open open API schema, exactly swagger and all the things.

SPEAKER_02

Exactly, exactly. And that's uh that really got me thinking about how you know message cues in in the abstract and specifically something like SQS or Kafka uh would do very well with having um you know schema management and message.

SPEAKER_03

I um I spent a couple of years working uh in API management uh layer a little bit. I worked at a company called Kong. And um the one of the products was like API management, the API gateway. And lots of conversations that I had with many customers is how to bridge um like API request response approach with um with messaging, streaming, event driven. Like what would be your kind of like a take on this? So you already you spent some time in API integration and now that you too like in event driven, what would be you know uh would be your advice for uh for customers who you know like to understand better how they move from the request response to event driven?

SPEAKER_02

Yeah, yeah, definitely. Um I think I think having um uh just having uh you want to loosely couple systems, right? I think I think in in the in the era of microservices and you know having um obviously like a lot of event-driven and data streaming use cases, especially for these days they're calling these like AI agents. Exactly, exactly. You you if you need to have something agentic and and really uh um under the hood, that's just about loosely coupling and having event-driven systems, right? Uh in a new form, a new iteration of uh you know event-driven systems. I think one of the uh key aspects of having something like that succeed is to build fault tolerance into it, um, is to have eventual convergence and asynchronicity built into those uh architecturally. Uh just and and that has to be end-to-end, right? And and not just it can't just be, oh, I'm just gonna have a for loop that just retries five times. That's that that won't cut it for once a system gets complex enough. You need something like um, you know, uh if we're talking about under the hood, right? You need something like a Kubernetes operator. You know, if you're trying to, let's say, provision a Kafka cluster, obviously we already have CFK, yeah, uh, but you know, something else like that, you need you need a you need a you need a uh framework like Kubernetes, or it could be anything else that's similar, but you need to build reliability and fault tolerance.

SPEAKER_03

Um one of the things that uh we built and we used in in our times using the service mesh that provides kind of like a virtualization of networking that allows to define different routes, how the traffic in um can can flow, and also you can have uh built-in uh kind of resilience um uh components that allows you to do kind of like a uh the back pressure, can do kind of like um uh retries, like you mentioned, kind of like exponential retries. If you don't want to kind of like cripple the system constantly, like trying to you get like a 404 or getting like a 503, and you continue to hitting the system with this.

SPEAKER_02

Um you don't want thundering hurts, you know, those are definitely bad, yeah. And and and and really um there there are um depending on the use case, there are off-the-shelf tools these days. So you have Argo, you know, workflow engine, you have a lot of tools like that. And those are the kind of things that uh you you definitely don't want to reinvent the wheel if there's a solution uh that exists out there.

SPEAKER_03

It's it's kind of funny because like we we had the episode where we were actually talking the times where um the tools were non-existent, um and the lots of things need to be kind of like a reinvented. And today, as as like 2025, 2026, you probably will listen to this uh podcast already in 2026. Um there's a plenty of tools available, and knowing these tools, um that's that's also important. What's your uh like favorite ways to learn those tools? Like what's your go-to, uh, you know, the mindset of learning about those uh new technologies and cloud native stuff and all these type of things?

SPEAKER_02

Um yeah, yeah, I I'm tuned into a lot of blogs. Um and I was tuning into Confluent before long before I joined Confluent. Uh uh and you know, on LinkedIn and obviously on the website.

SPEAKER_03

Yeah, the Confluent blog was known as a like uh resource even like before uh before the the cop become like a the big big company. The lots of technical excellence uh came from from the blog.

SPEAKER_02

Absolutely. And and the open source community is is obviously a big part of um you know how something like this well comes to be and then evolves and then and then you know expands, right? It which is what Confluent uh has become. Um so yeah, definitely I I I've always uh been keeping track of uh the Apache family of of uh of pro projects, and especially I don't know if it's if it's LinkedIn and there's been a lot of interesting um uh companies and projects that have been founded by LinkedIn alumni.

SPEAKER_03

Obviously, this one, but exactly, yeah, yeah. Um it's also interesting for me that like even, I don't know, maybe 10-15 years ago, people would say, Oh, like the the projects will go to Apache Foundation to retire or to die. It's definitely not the case anymore. So there's some um new projects and uh uh that came up in the in the in the field of data. There's a lot of like interesting things that happen around like streaming data, stream data processing in inside the Apache Software Foundation.

SPEAKER_02

And and in fact, I now see uh confluent confluent alumni now founding companies that that add to this ecosystem, especially in in you know, in the obviously data streaming and then AI agent use cases. So uh I I follow companies that are doing well because that tells you that those tools are useful. Yeah. Uh to me, that's the easiest way to keep track of um uh you know what is the latest out there that I can use. And that's I knew Kafka before Confluent, but obviously, you know, seeing Confluent itself grow uh, you know, made me realize the power of Kafka.

SPEAKER_03

Yeah. So um before we part, I would like to ask you one important question that maybe many of our uh listeners would like to know. Avro or protobuf?

SPEAKER_02

Protobuf.

SPEAKER_03

Come on.

SPEAKER_02

Uh it's seriously. It's the power. I mean, we used we use protobuf for ad serving. Yeah. Nothing else would compress the ad serving request that much.

SPEAKER_03

That's that surprised me because um during the uh during my some research on for some um some presentation, I learned that uh the the the people in Hadoop world came up with Avro because Protabuff was not able to um serve the needs uh in what way? I think it's all about schemas and how the Avro provides kind of like a read uh write schema and read schema so you can have a um get access to the data and the backward compatibility treatment is slightly more flexible than in Protobuff?

SPEAKER_02

Absolutely, yes. For that, yes, but given my background of like you know ultra low latency, you know, applications, I think nothing comes close to you know, protobuf, gRPC, that whole uh ecosystem.

SPEAKER_03

Yeah, that's great. Um, Diraj, thank you so much for being part of Confluent Developer. As always, my name is Victor Gamov, and as always, have a nice day.