Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov
Hi, we’re Tim Berglund, Adi Polak, and Viktor Gamov and we’re excited to bring you the Confluent Developer podcast (formerly “Streaming Audio.”) Our hand-crafted weekly episodes feature in-depth interviews with our community of software developers (actual human beings - not AI) talking about some of the most interesting challenges they’ve faced in their careers. We aim to explore the conditions that gave rise to each person’s technical hurdles, as well as how their experiences transformed their understanding and approach to building systems.
Whether you’re a seasoned open source data streaming engineer, or just someone who’s interested in learning more about Apache Kafka®, Apache Flink® and real-time data, we hope you’ll appreciate the stories, the discussion, and our effort to bring you a high-quality show worth your time.
Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov
Adventures in Data Infrastructure with Gwen Shapira | Ep. 11
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
Adi Polak talks to Gwen Shapira (Nile) about her career in databases and data infrastructure. Gwen’s first job: a side hustle fixing computers. Her challenge: figuring out why a production report at HP slowed down dramatically after daylight saving time.
SEASON 2
Hosted by Tim Berglund, Adi Polak and Viktor Gamov
Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
Music by Coastal Kites
Artwork by Phil Vo
- 🎧 Subscribe to Confluent Developer wherever you listen to podcasts.
- ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
- 👍 If you enjoyed this, please leave us a rating.
- 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
The clock changed and an entire production system fell apart. No code changes, no hardware failures, just time itself breaking the database. This is confluent developer.
SPEAKER_01Our CTO is expecting it every day at 8 a.m. He's getting pretty upset that it's not ready until much later. I think it takes some bravery to do it, but I can see how if your data changes significantly over time, it's a really good idea. So we turned on something that you basically never turn on in the database, and a lot of databases don't even have that.
SPEAKER_02I'm Adi Polak and welcome to Confluent Developer, where we explore the journeys of engineers who turn impossible programs into elegant solutions. My guest today is Gwen Shapira, database pioneer, author of Kafka, the Definitive Guide, and founder of Neil, a modern reimagined of Postgres for the AI era. From our teenage side hustle, fixing computers to debugging daylight saving time, Gwen's story proves that curiosity can turn even the strangest bug into a breakthrough. Let's dive in. Hi Gwen. Hey, so good to see you. Yes, so good to see you. It's been a while. It's been ages. Ages, I know. I'm curious, what are you up to recently?
SPEAKER_01Oh no, nothing special, just building a multi-tenant database and trying to market it, onboarding first customers to my product. Nothing super exciting. No kidding, it's super exciting, and especially just watching a product that you built grow is exciting. And obviously, I've loved databases for a very long time, and there is something special about having your own.
SPEAKER_02Yeah, it's amazing to see your vision comes to life. I remember we talked about it some years back, and you know, now all the pieces are coming together. So it's uh definitely amazing. Um for the people who didn't get a chance to know you before, maybe you want to say a couple of words about you uh so they can you know know you better.
SPEAKER_01Yeah, how far back should I go? Do I talk about? I left Confluent I think almost four years ago now, so it has been a while. I used to be the Kafka person and wrote a Kafka book and did a lot of Kafka talking. Before that, I was a Hadoop person. And now I'm kind of back to what I did before I was into the whole Hadoop thing, which is relational databases. And if back then I was mostly helping other people build products with relational databases, uh Nile, which I've been building for the last four years, is basically taking Postgres and updating it for modern workloads, things with AI, multi-tenant applications, all these kinds of things. That's fascinating.
SPEAKER_02I just want to call out one thing your book, uh The Definitive Guide for Kafka, is still a bestseller, and every time someone wants to learn about Kafka, this is the most recommended book. So just FYI.
SPEAKER_01We didn't know that because it's actually pretty old at this point. Someone should probably go and write a third edition. Uh, a lot of change, right? And especially with Confluent having Quora and like cloud native Kafka and all of that. It's uh there's a lot of new things to answer.
SPEAKER_02Yes, and cues for Kafka, it changed the whole architecture. But hey, this is for completely different conversations. Let's jump into it. So, in our podcast, we're gonna talk about roughly about some challenges that we solved uh as an engineers uh throughout our career. But the first thing that we have to know about uh Gwen is what was your very first job?
SPEAKER_01So it kind of depends how you count. My very first earning money was actually working for myself as a teenager. Everyone always asked me to for help with their computer, you know, fix my printer, my Windows has viruses, all this kind of stuff. And at some point, my dad told me, Look, you're helping all of my friends, you're spending hours helping all of my friends. Why don't you start charging the money? I'm like, I can charge people money for for helping them. I thought I'm doing it just by being a nice person. And I was like, Yeah, you know, you should probably save some money, you can buy nice clothes, this kind of thing. I was like, okay, let's do that. And it was so embarrassing to ask for money for what I do, and like it was stressful. What if I suddenly fail to fix their computer and now I'm charging them money and I'm failing? While previously I was mostly successful, but my fellas, it's like my fellows don't count if I don't charge them money. So I think this may count as a first job, even though I was kind of not working for anyone else. My other way of making money that did involve having an actual boss and working hours and all of that. I was working at the Hebrew University in Jerusalem as they were doing medical research, and it involved you know, small animals, mice and rabbits and all of those. Someone has to actually maintain these animals, you know, have to feed them, clean up the cages, pet them, and you know, make sure they feel uh taken care of kind of thing. And um yeah, I really it was an obviously didn't use any of the skills that I had except just liking small furry creatures, uh, but it was a nice way to uh spend time and make some money when I was uh yeah, I was in university anyway because I was getting my degree, so it was just a nice way to do it.
SPEAKER_02Yeah, skills that last till today. Um watching all your Twitter posts with cats.
SPEAKER_01Yeah, cats are way better than um those guinea pigs and mice, to be honest.
SPEAKER_02Right. I just want to highlight one thing that you said that um you know really hit home. It's the fact that uh your first job, it's when you're helping, you're actually delivering value to to someone, and it's okay to charge for it. Um you know, I I believe many people go through that phase of understanding that um delivering value equals uh you can monetize that. Very, very interesting learnings.
SPEAKER_01Yeah, it's funny because you grow up, you're asked to help people so much, and at some point it's like, okay, now it's okay to get to charge for it.
SPEAKER_02Right, right. It's um you're exchanging something for something, and uh that's always very interesting. Cool, very cool uh first two jobs. Uh one is very techie, one is uh, you know, uh love for furry uh pets. And uh that was very interesting. So I'm curious uh now that we gonna jump forward into your career and some of the things you build in tech, and you build a lot of things uh in tech. Um so what are some time where you hit like a tricky software challenge or technology challenge, and how did you figure out your way through it?
SPEAKER_00Now a quick word from our sponsor. Confluent developer the podcast is brought to you by Confluent Developer the website, which has everything you need as a developer of data streaming systems. And it's completely free. We've got curriculum, hands-on exercises, executable tutorials, the online data streaming engineer certification, also free, a way to find a meetup near you, those are free. Everything is there. I really want you to be successful in your journey as a data streaming engineer, and this is the site that has what you need. Check it out at developer.confluent.io. That's developer.confluent.io. Now back to the show.
SPEAKER_01Yeah, so I picked the specific story not necessarily because it was the hardest, but because it was, you know, how there is some things that are literally never so much never the problems that we're almost used as a joke about it cannot possibly be the problem. Like in Dr. House, it's never a lopus, and uh you know it's never a compiler bug. And if you have a bug, you can blame the face of the moon, but what are the chances that your bug is actually related to the face of the moon? So this is um actually a problem from many years ago, probably well over a decade now, and I was responsible for a large, very large set of databases at HP, and someone came over and said uh my query is running really slowly. We have a report that is it was over 10 times slower than expected, and it started happening when the time zone changed, like daylight saving time started, so we came over. Think it like was like end of merch, is like time we have daylight saving time since March 15, and now for the last two weeks, our report, which is supposed to be ready every day at 8 a.m., is actually not ready until much later, and our CTO is expecting it every day at 8 a.m. He's getting pretty upset that it's not ready until much later. And yeah, it's like how can possibly a report become slower because of daylight saving time? I would understand, okay, it's having the wrong numbers, maybe someone did the calculations that didn't take daylight saving time into account. This is super common. Every person who ever did reporting with data knows that time zones and daylight saving time make every report about 10 times harder. Yeah, but why would it be slower? Uh so the first thing you check is is it still cal getting similar results? And yes, does it still hit more or less the same amount of data? Did the amount of data in the database change? No, it didn't. What could possibly so and then you're like, okay, it cannot possibly be Del at 7 time, but you know, something may have happened two weeks ago. What did we do to the system two weeks ago? It's the same storage, we didn't really change anything. There were some operating system patches, but they don't really seem performance related, and the only thing it's really impacting is this report, and those are busy databases. What could it possibly be that only affects a single report and in and changed two weeks ago? So obviously, one of the if you have a large report, one of the biggest things that could change in a database is the query plans. And so I we had kind of sample logging for our uh query plans, and but this was a long and large enough report that we had really good historical samples for it. So that was good. That's one important tip. If you're responsible for performance, it's really good to have snapshots taken at regular intervals. Because if you have if someone comes to you and says, Hey, it's slow now, and it was fast before, and you don't have any information of what happened before, it's really hard to figure out why is it slower now because all you have is the bad state, you don't have any good state to compare it to. So we went back and looked at the what happened before, and the plan completely changed. The moment we saw it, it was like, Yes, it's clear and why it's slow now. It used to have a good plan, now it has a bad plan.
SPEAKER_02So the query plan, the the the actual query plan changed, like the physical query plan or okay.
SPEAKER_01Well, depends on the layers, there are different layers, but it's like the physical query plan of how the data essentially how we're so when you do in relational databases, when you run a query, you write SQL text, and then there is an planner and optimizer that says, okay, this join is going to be nested loop join because one side is very small and one side is large, or both sides are large, you're going to do a hash join. So all and usually in uh in a larger port, there is thousands of those decisions that the optimizer has to make in order to figure out how best to produce this data. And got it. Without the data changing, the plan became completely different. Now there could be reasons for a plan to change. Uh one of them is that the statistics have changed for whatever reason. Uh the plan is created based on data statistics. The data didn't change, and I checked the statistics, they also looked completely normal, which was very suspicious. The statistics didn't change, and yet the plan was totally different, which is absolutely insane. And so we turned on something that you basically never turn on in the database, and a lot of databases don't even have that. We're a bit lucky that back then Oracle had this. You can turn on a flag that it will not just log the plan, it will log every decision that the planner did and why it made a decision it made. So I'm choosing this because of those statistics. And the moment I looked at that, I could see that the statistics the plan we were using were not the statistics that I was looking at. It was it basically believed that the entire report was running on empty tables. Like, why would it think that my entire report is running on empty tables and plan for that? Well, turned out that we had a nightly job to refresh statistics because you know you load data, you delete data, you make updates, things change. So every night at I think midnight, uh a job ran that deleted all the statistics and collected new ones. The report was supposed to start at 1 am. I don't know if you see the problem. The late 70s started. The statistics for some reason stayed at midnight because there were this was one chrome job that basically midnight was stable there. The report generation moved to an hour earlier at to start at midnight. So it started at exactly the one point in the day which lasted no more than 10 minutes, in which we had no statistics. So it did all the planning based on the 10-minute interval in which it could believe that all the tables were actually empty.
SPEAKER_02Wow.
SPEAKER_01So yeah, the night saving time can cause the report to be a lot slower if the report runs at the wrong time. Which by the way is something that excited me, I think fairly recently, but like two years ago, Databricks published a paper and how they allow if they monitor the query plan during the execution, and if it discovers that the amount of data it sees while executing don't match the amount of data it believed it had while making the execution plan, it actually goes back and replans and modifies the execution in real time. It's called adaptive query something. And what when I read the paper, I was like, oh my god, I could have used that years ago.
SPEAKER_02Yeah. It's fascinating. So essentially, let me see if I if I got it correctly. Essentially, there was a you know some bug in the uh in the query, the query was super slow, um, something happened there, and you start investigating and looking into the data and what is going on, and then you turn on uh the query planner output so you can see exactly you know what's the the actual the physical plan that got out and what are the statistics that these plans was based off. Uh and then you realized, and I'm guessing you have to sift through like many, many days, right? Like to compare uh the different times where everything worked well versus when things start to go uh downhill. Um and then comparing like the metadata there and and seeing like what what's the plan, what's the metadata, what's the statistics, if that makes sense, only to realize that the query, the the job was planned to run at 1 a.m. And when the clock changed, you know, something that happens twice a year. Um, but it was a specific case, was moved back to uh uh to midnight when just before the metadata query that updates the statistics started to run, right?
SPEAKER_01Yeah. I mean it's uh it managed to move to exactly the point after the statistics job managed to drop all the old statistics and before it managed to collect the new statistics. It was incredibly bad luck.
SPEAKER_02Wow, that's uh yes, yeah. How did you fix it? What was what was the fix for that?
SPEAKER_01I mean the fix was easy, right? You just uh change the timing of the reports. I mean it's amazing when you the report starts 15 minutes later, it actually finishes four hours earlier. It's almost like you know, the time you leave home or work if you have to time it with rush hour traffic, and sometimes like 15 minutes change has a huge impact on when you actually arrive to work.
SPEAKER_02Yeah, especially in the Bay Area. It's like this 10 minutes in the morning really really counts. Every minute counts, absolutely. Every minute counts. And yeah, that's super cool. And the um the adapt adaptive uh query planner is also a very interesting solution. I also looked at it some years ago, actually, not recently. Um, and I was very surprised to uh to see that. But I do I am curious because with every kind of like magical thing that happens automatically uh under the hood, um what's the the rate of you know errors and mistakes? But I guess we'll never know because uh I agree.
SPEAKER_01I don't know if adaptive is good or bad. There is something nice about knowing that you have the same query plan for the same report day after day. Um but on the other hand, if you know statistics change, data changes, uh then you actually don't no longer want the good old one day after day. So it's uh definitely an interesting judgment call. Uh Postgres now has an extension for uh they call it a switch join, which is kind of the same idea, you know, joins have a lot of impact, and there is several normal join methods and you choose between them based on statistics. And the switch join extension basically means that if it chooses one of them, but if during executing the join it looks like it's getting way more data than expected or way less data than expected, it will switch over to the alternative plan midway. I think it takes some bravery to do it, but I can see how if your data changes significantly over time, it's a really good idea.
SPEAKER_02Yeah, it's fascinating. All the all the optimizations around joins like broadcast joints, hash joins, and so on are you know a whole universe that I think um not many software engineers get to dive into. So it's definitely an exciting part that you do. Um, I hope you know you'll write more about it. Uh I I would be the first one to read. I'm making a note. Here we go. Challenge accepted. All right. Gwen, thank you so much uh for joining us today and sharing your your knowledge and expertise. It was you know super exciting to to hear first about what you did early on and then later on some challenges you're solving. And of course, I'm very excited for Neil and everything that you're building. Um, I, you know, optimistically, you know, carefully optimistic, but I think this is the future. Um, so very excited for your vision to come to life.
SPEAKER_01Thank you so much. It's been a pleasure as always.