Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov

Inside $3M GPU Racks: Powering Modern AI with Bryan Oliver | Ep. 16

Confluent Season 2 Episode 16

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 30:18

Adi Polak talks to Bryan Oliver (Thoughtworks) about his career in platform engineering and large-scale AI infrastructure. Bryan’s first job: building pools and teaching swimming lessons. His challenge: running large-scale GPU data centers while keeping AI workloads predictable and reliable.

SEASON 2
Hosted by Tim Berglund, Adi Polak and Viktor Gamov
Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
Music by Coastal Kites 
Artwork by Phil Vo 

  •  🎧 Subscribe to Confluent Developer wherever you listen to podcasts. 
  • ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
  • 👍 If you enjoyed this, please leave us a rating. 
  • 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
SPEAKER_02

Today we're asking one simple question. Can your platform actually handle AI or will your GPUs melt first? This is Confluent Developer.

SPEAKER_00

We're constantly trying to solve and innovate for these problems right now. Okay, how do we track that variability across the entire, say, GPU data center? The problems that we had 10 or 20 years ago are the same problems we have now, just in a different context.

SPEAKER_02

Hello everyone. I'm Adi Polak, and you're listening to Confluent Developer Podcast, where we uncovered the human stories behind complex software systems. In this episode, I'm discussing with Brian Oliver, a platform architect and author, and one of the voices shaping how large-scale AI systems are actually being delivered in production. From teaching swimming lessons and building gaming in his teens to writing Lua add-ons for World of Warcraft, to now scaling GPU's platforms and influencing the future of cloud native AI. Brian's journey is all about performance, resilience in the pursuit of predictability in a world that is anything but Brian.

SPEAKER_00

Hi, how are you?

SPEAKER_02

I'm good. I'm I'm so excited to that we found the time to actually talk. I know you've been doing a lot of things recently and you've been super busy, you know, working on architectures and GPUs and platform engineering and so many good things. So what what keeps you busy these days?

SPEAKER_00

Uh yeah, so really deeply ingrained in large-scale GPU work right now. Um giving a talk at QCon in New York, QCon AI in New York about some of that, like doing chaos engineering and large-scale GPUs. Um also, you know, book writing and giving talks and all this other stuff. It's uh yeah, it keeps me pretty busy. For sure.

SPEAKER_02

Right. Book writing is is a labor of love.

SPEAKER_00

It really is. Yeah, we just finished our first one, um Effective Platform Engineering with Manning. Uh it arrived on my door a few days ago, which was pretty cool. And then uh my second book with O'Reilly is uh we the title may be changing actually, but currently it's um delivery systems, but we are reframing it to have some AI stuff in it. Um so it might be something like delivery systems in the age of AI or something like that. So but that'll be in about a year from now.

unknown

Yeah.

SPEAKER_02

That's really interesting. So it's our do you do you see like a profound change to our delivery systems with AI today?

SPEAKER_00

Yes. Um this is maybe also a precursor to the ThoughtWorks radar coming out pretty soon. Um I'm on the committee that writes that. And one of the central themes of the most recent radar session in Bucharest was um we talked a lot about AI scheduling, which is a deep problem right now. Um, one of the things going on in this space is you have these really massive data centers with you know thousands of GPUs and they all cost millions of dollars. Like a I think a GB200 rack that Nvidia came out with recently costs like three million dollars for one rack. Um and they're all network-linked GPU to GPU, and they're they're next to each other. Um and so one of the things that Kubernetes and other scheduling systems wasn't super aware of is like how to deploy jobs in a contiguous way, where it's like that because the jobs are split up across tons of GPUs, so you want the job to hit like a block of GPUs that's close to each other to reduce latency. Um so we call this topology aware scheduling, something that's existed in computing for ages, but now it's a GPU problem. And so all of our delivery systems are starting to adapt to this new requirement. Um it's really interesting. You see a lot of projects in the CNCF and in SLERM that are tackling this problem.

SPEAKER_02

That's amazing. I'm guessing I can read more about it in your upcoming book, right?

SPEAKER_00

Yeah, the O'Reilly book, we're gonna talk about it a lot. Um, in the more short term, we're gonna have some things on the ThoughtWorks radar for that as well, in the themes as well as some of the blips. Um yeah, definitely check that out.

SPEAKER_02

Yeah, I'm definitely subscribed to that, and every time it comes out, it's uh it always becomes a huge conversation uh with my team and in the company about all the new innovations in uh what is uh you know uh what has proven itself to uh to work well in production, what is still experimental, and uh I'm super excited for that.

SPEAKER_00

Yeah, it's there's pretty much this whole radar edition is all about AI stuff. It just dominated all of the conversations, so it's gonna be a good one. Yeah.

SPEAKER_02

Yeah. And hopefully it becomes more practical. Um so solutions are use cases are getting into production and uh being materialized there as well.

SPEAKER_00

Yeah, yeah, that was a lot of a like it seems like in the past sessions we it was lots of like, oh, experiment with this technology, try that. Now it's like a big, yeah, bigger focus on like all the production stuff and like scaling. Um, and those are all becoming more interesting topics now. Yeah.

SPEAKER_02

Exciting times, exciting times. For sure. Um, so I don't know, you know, uh we have a lot of people listening in, and a lot of people are curious about you and and your background and and more things. So maybe you can take us back in time to your first ever job.

SPEAKER_00

Oh boy. Um my first job was as a teenager. I I worked at a swimming academy. Um, and I started off with helping them build pools because they were expanding um their scaling. Um, and then I eventually got my I didn't like that part, so I got the the certifications to be a lifeguard and a swim instructor there. Um and so I started off with teaching kids how to swim and eventually moved into like teaching like um triathletes and stuff, um, like working on their performance tuning, and and I really enjoyed that. And then as I got older, I'm in like my late teens, I worked at a land gaming arena. Um, and that was a a really low-paying but really fun job. We we got into some pro gaming through that because we had teams that were hosted there. So uh I never made it onto uh like a starting squad, but I was an alternate on some like battlefield pro teams and call of duty. So I got to participate in some some like game battles tournaments, and it was really, really fun. It was uh it was a cool job. I enjoyed it.

SPEAKER_02

Sound exciting. So the love from like scaling pools and uh swimming lessons to scelling GPUs and performance tuning. Uh it's kind of like a theme there.

SPEAKER_00

I know, right? Yeah, it it's funny. I actually wasn't even into computer science or mathematics. Uh I but I grew up playing World of Warcraft from like age, I don't know, 13 to in well into my adult years. I don't play anymore, but I I would write um add-ons for the game in Lua. So I was experiencing I got some experience programming early on, but I never like considered it as a career or anything. I was gonna be an English major in college, and I ended up uh uh switching to computer science because I took this symbolic logic course, which was like a philosophy course. Um, but it's basically just discrete math but more abstract, which is the sort of core math behind computer science, and so I ended up switching because I enjoyed it so much. Yeah.

SPEAKER_02

It's fascinating. It's always interesting how you know discrete math and some uh you know aspects of mathematics can be it's more philosophy driven, and uh, and then we bring it into the computer science world, into the actual uh zero ones and uh soon, who knows, quantum and the rest of the things.

SPEAKER_00

So yeah, it was really fascinating. Like it in the philosophy aspect, it's lots of like set theory, and but it it ends up being very similar to manipulations you do in discrete math and techniques you use, um, which is really just it really drove me towards computer science and math at the end of the day. I went from wanting to grade English papers and teach British lit to to yeah, programming GPUs is quite the switch.

SPEAKER_02

I can imagine. And and and right on time. I think when I look into the future, when I hear you know, people like um Nvidia CEO and so forth and so on, I think GPUs might be dominate like the the future cloud.

SPEAKER_00

Yeah, yeah, I think so. Um we're gonna see a lot of I mean we already do, we see a lot of like large-scale cheaper GPU um scaling happening for the inference requirements. But what's interesting is the the model sizes are changing the definition of what small and large mean. Um like the what small language model means now is kind of a moving target to the point where it's like you have these um GPU racks that I was talking about earlier, the $3 million one, and that thing has 13 terabytes of memory, unified memory, and it acts as one GPU. That's now the new large in a lot of ways. Um, what does that mean mean for small language models? Um and the size is kind of just this like moving target. So we're starting to see like those scaled requirements I was talking about earlier with scheduling now are even being applied to small language models because they still need to be spread across multiple GPUs. Um so we're starting to see these distributed complex needs um occur even in everyday computing and companies that don't have access to large-scale GPU infrastructure. Yeah.

SPEAKER_02

Interesting. So we we didn't yet I think when we started, it was like a small language model should fit on one machine. So essentially beginning what I've you know experienced was it wasn't uh a distributed topology behind that. So now you know we're seeing that we're growing past that. But uh a little bit, yeah.

SPEAKER_00

Like you you have like uh what is I think Quen 72B or it's a I can't remember exactly, but there's there's some models that are like 72 million parameters, and you would now almost think of those as small when you compare to who's using 13 terabyte um GPU racks and even spreading across multiple of them. Um maybe that's just considered massive scale and we we need a new definition. Um but yeah, it's really interesting. It's definitely a moving target.

SPEAKER_02

Yeah, and it definitely pushes you know platform engineers to to think beyond uh some of the practices we had so far. Um this is why I'm I'm I'm super excited for your books. So uh I'm gonna order both of them and uh and read through them and have some notes. Um I'm gonna I'm curious, you've been solving a lot of tech challenges in in your career, and uh of course you're at the forefront of what's happening now in infrastructure and platforms. So maybe you want to share some challenges that you're working on and perhaps some learnings or or things that you know uh came out of them.

SPEAKER_01

Now a quick word from our sponsor. Confluent developer the podcast is brought to you by Confluent Developer the website, which has everything you need as a developer of data streaming systems. And it's completely free. We've got curriculum, hands-on exercises, executable tutorials, the online data streaming engineer certification, also free. A way to find a meetup near you, those are free. Everything is there. I really want you to be successful in your journey as a data streaming engineer, and this is the site that has what you need. Check it out at developer.confluent.io. That's developer.confluent.io. Now back to the show.

SPEAKER_00

Sure. Um, did you want were you interested in challenges I had many years ago or more recent?

SPEAKER_02

I think the recent one are super fascinating. Um, but you know, I'll let you decide. Cool.

SPEAKER_00

Yeah. Um yeah, I think the one of the interesting challenges now is um not just the topology awareness um like we were talking about earlier with trying to reduce latency for deploying AI workloads. Um there's this new kind of newer concept. There was a paper presented at Supercomputing Conference 2024 here in Atlanta on um variability aware scheduling. And I'd never thought about this before, but when you have, say, a hundred GPUs and they're all identical, um, the difference in their performance is um much larger than 100 identical CPUs. It's much less predictable, in fact. Um GPUs are just not as stable of hardware, not as predictable. So, like if you have a data center full of them and they're all identical, some are going to be performing better than others by a pretty significant percentage in some cases. And this could be related to some are being cooled better than others, some are maybe not getting the amount of power they're supposed to get. Like, there's all these different factors into why they're not quite as predictable. And so one thing we're starting to think about is okay, how do we track that variability across the entire, say, GPU data center and figure out, okay, this block over here has a few that aren't performing as well. Maybe the cooling's not doing great over there, but it's it's still working fine. We could still run things over there. So, how do we maybe like say, okay, if we're doing like a multi-tenant architecture data center, send like, okay, we have this really premium customer or really critical job we need to run. We're not gonna put them in a block of GPUs that has maybe some that aren't performing as well. Um, and then there's also different types of training jobs, even um where it's like some don't really care about those performance differences, whereas others do. So it's like you are also starting to maybe categorize the types of training jobs you're running and then scheduling it to certain areas of your data center and others depending on the needs. Um, so it's the complexity just keeps going up and up. And so we're we're constantly trying to solve and innovate for for these problems right now, and then new techniques are coming out, like new papers are coming out every day where it's like our team chat is like, read this paper, because this is really like it, this is almost a daily occurrence now, um, where we're like, oh, cool. Oh, yeah.

SPEAKER_02

Yeah, it's fascinating. So essentially, you know, I can buy the same hardware, I'm paying you know good money for it. And uh because my cooling system or electricity system is not uh top-notch across, I'm guessing. Um, this is when I'll start saying seeing like different performance coming out of uh these GPU, and I'm guessing, you know, even if you put all the GPUs in one rack, just the proximity of one GPU to the other and the the heat that comes out of that could be uh impacting the performance. Um and you mentioned that there are some algorithms that are uh okay with this uh you know performance differences and and variability, and some of them are are not. I'm curious, like what's uh you know, what's in the algorithm makes the difference? Um and how could people navigate that perhaps?

SPEAKER_00

It's not so much the the algorithm, even it's um it's like you need to have the the data. Um so like if you you're always running AI workloads inside of say, I'm using like one data center as kind of an abstract way of talking about this problem, but this is across hundreds. Um it's you need to collect telemetry on all of that hardware and then sort of centralize that into some sort of monitoring system. So then your scheduling system can be aware of like each one um individually so that it can then make intelligent decisions. Um then you have technologies like Slurm in the sort of old school. I would say old school because it's actually becoming quite modernized by cloud native requirements. Um but Slurm's been around for a long time, since I think 2004, that is aware of all of these different types of nodes that are available to it. Um and Slurm actually runs like a daemon on every single GPU server in your data center. So it's collecting information on that node, it knows all of its different um aspects, types of GPU, all that. And then it can make sort of good decisions on those basic hardware things, but it's not aware of the those variability differences. Like you you have to implement that yourself. Um, so we're starting to look at also CNCF-based projects like uh Q or DRA and Kubernetes um API advancements. Um there's also like the Kai scheduler, um Skypilot, which is uh Kates or non-Kates. Um all of these can be fine-tuned to be aware of the topology. Um, but the it takes investment and time from your team in order to take advantage of that. Yeah.

SPEAKER_02

Yeah. And is there a way to improve like a specific Jupyter performance? Let's say now you know we you don't send any workload to it, put it on idle for X amount of time. Is that something you've been looking into?

SPEAKER_00

Not as much, but we we are trying, we are looking into the talk I'm giving at QCon is about um maybe not driving more performance out of current ones. A lot of that has to do with like cycling them out and scheduling them for work or maintenance from, say, data center team, but um getting more predictability out of the performance by maybe doing chaos experiments um on your GPU infrastructure so you start to know ahead of time where the issues are instead of finding out while you're currently running workloads. Um we're starting to think about like, okay, how do we apply the concept of chaos engineering from the API world and start to move it into this GPU world? Um and it's very different. It's hard to put something like that together because you need access to really expensive infrastructure in order to really make it work. Um But we're we're starting to work with um Chronosphere on an open source project to try and get something that that's functional for doing that. Yeah.

SPEAKER_02

Yeah, it's it's fascinating. I wonder if you can walk me through like what would an experiment look like.

SPEAKER_00

Um yeah. For sure. Um so you you might have your your typical ones um where you're like say blackholing certain nodes. Um what black holing is is it's the concept of like things that can go in, but they can't come out of, say, a node. Um you might do like take nodes down, like a chaos monkey kind of thing, like all the normal stuff. Um but where things get interesting is you also want to do things specifically to the GPU. And one node might have, um this is quite normal, like one server might have eight GPUs on it. And the way you shard up a really large training job is you like say you're writing some PyTorch code, you might inform your training job or your scheduler, okay, you're going to be split up across, say, 16 GPUs, and this is going to be on two servers, so sets of eight. Um well, the way that that um model that's being trained gets divided up, if one GPU from the set of eight goes down, you then need to stop the job and switch over to another server that has a full set of eight. Like you can't just continue on with the odd number. You either have to tweak your job and your parameters or move over to another node that has a full set of eight going. So in the chaos world, what we're thinking about is like, okay, how do we go into the server and maybe bring down one of the eight and see how our scheduling processes respond? Um we could also maybe create noisy neighbor type things where it's like there's something going on with one of those GPUs. Maybe it's being consumed by another job or another daemon or maybe an agent collector. Um and then you can get really low-level and like eBPF or um NVIDIA has um DCGM, which is this um sort of low-level GPU management tool where you can interact and actually send commands to directly to GPUs, and you can send it faults. So you can pretend to create like heat or power faults through DCGM. Um so we're starting to think about things like that as well.

unknown

Yeah.

SPEAKER_02

That's cool. So Nvidia actually thought about it and it's like, hey, we know it's gonna be stress test, and uh people are gonna build different experiments, and uh, we're enabling people to uh to send faults to our GPUs as well. Um kind of, yeah.

SPEAKER_00

They they did it for themselves really, because they they need to test their GPUs um when they're building them. So they built the tool, I think, more for that. Um the DCGM has two projects, um the exporter, which was created for the Kubernetes world, which just collects metrics on the GPUs running, and then the DCGM like CLI tool, which actually like you can run commands against. And they really made it for themselves to test and work on GPUs and verify they're good. But it's just a CLI process you can build it into like a job or API, that kind of thing. Yeah.

SPEAKER_02

Yeah, it's amazing how in software, you know, sometimes you build one thing, it's just for us to test and validate that everything works, and then we expose it to the world, and the world's like, hmm. I have some other use cases just for that.

SPEAKER_00

Exactly. Yeah, exactly that.

SPEAKER_02

Very cool. So I I'm curious, like you know, it's um I didn't realize that. The beginning to be honest, like GPUs is going to have a different performance because I've been um you know of a working a lot with more CPUs, and GPUs are kind of like an additional to uh to the existing CPUs that I I worked with. And that that's really a surprise to see that there is uh variability uh in in the same machines. I'm curious if there were you know more things that kind of like are were less intuitive to you that you learned through working on these projects.

SPEAKER_00

Um mainly it was around the the differences between machine learning, like what a machine learning engineer wants and needs and an operations engineer wants and needs, because they're the the current like AI workload and supercomputing communities um have really been using this Slurm project for a long time. And um the Kubernetes community is really trying to move into supporting this space because the the AI operations engineers want to use Kubernetes and the MLE engineers just want their jobs to work. And that's been Slurm for a long time for them. It's just like they can use Slurm APIs and quickly write their jobs and they don't have to care, um, which is different, but also similar to like the problems we've heard from like when DevOps came around, where it's like devs just want their code to work, they just want to write code, and operations engineers want it to be predictable. And Kubernetes kind of solved that that problem of helping them meet in the middle. Now we're trying to do the same thing with AI workloads and get to the point where our MLE engineers can, um MLE is the acronym we use for machine learning engineer, um, can create jobs that are deployed to Kubernetes, and it's that same, like they don't have to care where it's running. Um I was at the uh Kubernetes um contributor summit in in Paris last year. And Tim Hawken got on stage and he was like, we have to start like all focusing on AI support, or our project is going to get left behind for something else. And from and that conference was so focused, that KubeCon was so focused on AI. Um and since that moment, it really has exploded in the CNCF space.

SPEAKER_02

Yeah, it's um it's fascinating how requirements get changed, but the kind of the nature of the two different types of engineers that work together in a company, it's like, you know, yeah, ML engineer wants to get things done, platform engineer want to get real real ability and uh predictability for for the platform. Um so beyond like we have the GPUs, it's kind of a game changer and they bring new um challenges, we'll put it that way. If we do technical problems or challenges that uh we got to solve. What about like working with Emily's are requirements? You mentioned PyTorch. Um are there other requirements that they're now focused on these days?

SPEAKER_00

Yeah, there is PyTorch is just you know one project of many that uh machine learning engineers are using. Um so we the the approach we're trying to focus on is um helping them sort of containerize their jobs and workloads so that they can be cloud native. Um what's interesting is you even have some different container runtimes in this space. Like um there's nroot, which is um one of NVIDIA's sort of container runtime. Um you also have like the latest and greatest GPU architectures um have some differences between maybe traditional um server and node architectures that require changes to these upstream projects in order to support them. Um so it's re you start to see like certain projects need to sort of adapt um to these ever-changing like the um the GB200 um has some differences in its hardware architecture that require um you to think a little bit differently about how you do AI and how you do training in order to make use of that really large-scale um hardware. Um, because it's technically like one GB200 is, I think, like 72 GPUs, but the way they've designed the architecture is it acts as one. Um so the way you write your jobs is a little bit different in that context. Um I've never written one myself for GB200, so I just know kind of how that works. Um but you know, getting access to that kind of thing is is very expensive. So it's all kind of theoretical for some of us. Yeah.

SPEAKER_02

I can imagine. And it's probably like the Emily needs to change some of the way, or the researcher need to change some of the way they're developing. Or is it like a platform supposed to be kind of like um catch them all, uh figure out uh you know which hardware we're running on right now and try to do the adaptations?

SPEAKER_00

Yeah, they definitely need to be aware. Um even you know, the MLEs that aren't working in that bleeding edge, really large scale. Like let's say they're using like, you know, the one of the standard GPUs you see in clouds is like the um, I think it's the Tesla T4 or L4, um, or maybe like some H100 um GPUs, which are quite large. I think they support like 80 gigabytes of VRAM per. Um the MLE needs to be aware of like what node server, like GPU server they're targeting and how many GPUs it has, how much virtual memory it has. Those are all things they're probably always going to have to care about. Um but the things we're trying to abstract are how that workload gets scheduled to that infrastructure. They shouldn't care if it's um Slonar or Kubernetes or whatever. Like they should be able to just write their training workload or their inference workload and deploy it. And so that's the abstraction we're trying to get to.

SPEAKER_02

Very cool. Very cool. Um one or two learnings that you want to share today from all these great projects that you're working on.

SPEAKER_00

Um or two learning. I think the it's funny to me how the same. I guess one of the things I learned is the problems that we had 10 or 20 years ago are the same problems we have now, just in a different context. So I'm learning like the problems repeat themselves just in different ways. Like the, you know, we were trying to bridge the gap between developers and operations 15 years ago when Kubernetes solved that, and now we're we're kind of dealing with the same problem. So it's like I'm I think what I'm learning is I'm starting to recognize these problem patterns um in computing and realize like, oh, this is similar to a problem we solved 10 years ago, it's just a different kind of hardware or context or whatever. So I'm learning to maybe recognize those things, um, and that helps me like think about how to solve them.

unknown

Yeah.

SPEAKER_02

Yeah, uh software is software is software. It's uh great advice. Always stick to fundamentals, you know, know the core, know the basics.

SPEAKER_00

Exactly. Yeah.

SPEAKER_02

Awesome. Brian, thank you so much for sharing your expertise and experience here today. I am super excited to read your books. And I know you also you also give a workshop, right?

SPEAKER_00

Yeah, so we um me and one of the other co-authors on the platform engineering book just announced a workshop on platformengineering.org. It'll be the platform architect certification for that website. Um, and then I'm teaching a workshop on platforms with AI at KuCon San Francisco coming up in November.

SPEAKER_02

Very cool. So if anyone is listening at home and want to catch up and want to become certified and knowledgeable in all this space and build a career there, they should definitely uh go to this website and uh check out your certificate and uh get your books. Thank you so much. Thank you. It was a pleasure.