Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov

Gunnar Morling Built a New Parquet Engine with AI | Ep. 31

Confluent

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 42:36

Tim Berglund talks to Gunnar Morling (Confluent) about his career in open source Java and data infrastructure. Gunnar’s first job: a student PHP developer in AMD’s e-learning group. His challenge: building Hardwood, a fast, multi-threaded Parquet engine for Java with minimal dependencies.

► The One Billion Row Challenge blog post: https://www.morling.dev/blog/one-billion-row-challenge/

SEASON 2
Hosted by Tim Berglund, Adi Polak and Viktor Gamov
Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
Music by Coastal Kites 
Artwork by Phil Vo 

  •  🎧 Subscribe to Confluent Developer wherever you listen to podcasts. 
  • ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
  • 👍 If you enjoyed this, please leave us a rating. 
  • 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
SPEAKER_00

Send PRs to Gunnar's Projects, you'll get a job. That's what I'm hearing.

SPEAKER_01

Okay, that sounds good, yeah. It's the first AI-based project that I'm doing. And, you know, it's AI first, I would say.

SPEAKER_00

Hey, Tim Bergland here. Welcome to another episode of the Confluent Developer Podcast. Today, I spoke to returning guest, Gunnar Morling, about Hardwood, his new Java-based parquet parser. Gunnar has applied some of what he learned in his famous 2023 viral hit, the One Billion Row Challenge. I'll put a link to that in the show notes if you don't remember it. It was pretty cool. He's combined that his extensive experience as an open source developer and project leader to build a next generation reader and writer of parquet files. That's what Hardwood is. And in this episode, of course, we talk about how AI augmented engineering has played into the project and kind of how he's using AI. And I think we just generally give the thing a pretty solid deep dive. I really enjoyed this conversation. Always love talking to Gunnar, and I hope you love listening in. I'm doing all right. It's good to have you back. Um, we are talking about uh your new woodworking hobby, as I understand today.

SPEAKER_01

Right, yes.

SPEAKER_00

Yeah, you want to talk about hardwood. I was confused, but now uh you've been working hard on this project called hardwood. Uh I've seen a lot of conversation about it. Maybe that's mostly from you, maybe it's for other people. I don't know. But I thought this is really important and I wanted to get it on the show. So uh what is hardwood?

SPEAKER_01

Right, yeah. Awesome. I mean, first of all, thank you so much for having me again and you know, uh taking the time to talk about this project. Yes, I went into the flooring business, you know, and I'm dealing with hardwood and parqueting.

SPEAKER_00

Are building a house, so this is all very interesting.

SPEAKER_01

Maybe we can you know find an agreement there. Um, the idea is, I mean, uh, you know, it's a new parser reader and uh hopefully also writer for what's called Apache Parquet. And Apache Parquet is a column of file format, and we can talk about what that means. And yes, hardwood helps you to parse and also uh soon write those parquet files. You know, and I will I thought I would be smart, and you know, the project is named, you know, it's like a lay on words, so it is file format Apache Parquet. I thought hardwood, it's another kind of flooring, and that's that's you know where the name is coming from. But it's terrible in terms of search engine results because if you search for parquet and hardwood, it's it's really bad.

SPEAKER_00

Exactly. I'm an expert at that kind of thing. Like, oh, this name is so clever, I'm so happy with myself. Uh yeah, you can't search for it.

SPEAKER_01

So yes, exactly right.

SPEAKER_00

Um but uh so obvious question. Um parquet, and uh maybe I'll throw a link in the show notes. Like if you if you don't you said it's calm or file format, if you don't know what parquet is, we'll we'll give you a link. Uh you you should read up on it. Uh but it's not like there weren't any APIs. I mean it's the kind of thing where there's pick your language, there's a language binding. There certainly was a Java one. So why did you do this?

SPEAKER_01

Right. Uh I would say two and a half reasons uh for doing that. Um so yes, first of all, there is a very well-established library for um parquet in Java, which is called Parquet Java. So um, you know, people have been using this for a long time, it has been around for a long time, and this is part of the reason of the reasons why I thought there is uh you know the need for something new, because it comes with quite a bit of dependency baggage. So essentially, if you pull in parquet Java, you pull in the entire Hadoop stack. Um, this is just, you know, because of where it's coming from. And now you want to parse a parquet file, and you don't want to have like a gazillion of you know transitive dependencies jars on your class path just to do so, right? So um it's a nightmare in terms of um just supply chain security, right? That's a lot on our minds these days. Um you just don't want to have all this uh this overhead of all those uh dependencies. So that's the one thing, first reason. Then this existing library is uh single-threaded. And you know, today, of course, all our machines they have like uh many, many cores. My laptop, 16 cores. I guess uh everybody has you know those beefy machines these days. And we want to make use of those uh CPU resources, right? We don't want to just parse a file with one core and let all those other 15 uh cores sit idle. So that's the second reason. And as I said, it's like two and a half. There's also an angle. Um, you know, I wanted to see how far I can get using those LLM tools like Cloud Code and so on for building such a relatively low level uh code base. So this also was you know just for me a little bit, okay. How good are those tools? Are they are they fit for purpose for this particular application? And this is you know also like kind of like very selfishly a theme for me.

SPEAKER_00

Yeah, okay. Um and I I since you started this, I have kind of felt echoes of the billion row challenge. I mean, it it's not exactly the same, but it did that experience inform you here?

SPEAKER_01

Absolutely. I mean, it also there was this idea, yes, let's put some of those learnings from the one billion row challenge into practical use and you know let's see what we you know can make with those findings and insights, and uh can we build an actual um you know hopefully production worth your project based on those insights? Absolutely.

SPEAKER_00

I'm I'm super interested in the multi-threadedness. Maybe we can come back to that later. Um that seems difficult to pull off. But uh first the zero dependency thing. Um you know, for first of all, John Donne would like a word with you. He's the poet who wrote uh No Man is an island and tire of itself. And you're like, No, actually I am. Uh and you know you know pulling in a bunch of dependencies. If you build stuff in Java, if you build stuff in JavaScript or TypeScript, like you that's just that's just life, right? You have a dependency and it has 600,000 other things it needs. And you made a point about about the supply chain aspect there. So what um tell me more about that and what do you feel like the trade-offs were? What did you lose in doing?

SPEAKER_01

Absolutely. Yes, I mean it's a it's a really multi-faceted uh uh question there. Um but so yes, why is it a problem? I mean, you know, just having all those dependencies, um, yes, so there is this uh security aspect, right? So there could be CVEs in any one of them. You need to uh make sure like you're on the latest versions, then maybe you just cannot simply upgrade like one dependency, uh like one transitive dependency, it could you know cause version conflicts. So there's that. Then there is just the question of like what's the footprint on your uh class paths, and like uh do you even need to load so many classes to do a certain thing? So so there's all that. But yes, of course, also libraries give us something, right? It's not like we just randomly uh pull in stuff for the sake of it, right? So we want to benefit from that uh functionality. And um, I think this is where the genetic coding tools they definitely change that consideration uh quite a bit. Um, because well, with those tools, I think you just can self-maintain much more code. So let me give you one specific example. So, you know, many oftentimes those parquet files they are stored on object storage. So you have a file, it's stored on S3, and now instead of downloading the entire file to get something out of it, you uh you know you can apply all kinds of neat tricks and statistics and so on to only touch on specific chunks from the file. And so it just reduces the I.O. uh quite a bit, right? Um so you only need to download specific sections of the file you're interested in. Now, typically you would do that using the uh S3 SDK from Amazon, um, which by itself is quite dependency heavy, actually. And now here what I'm doing is well, Java since version 11, it has a built-in HTTP client. This is totally good enough for that. I can do those range requests, that's fine. But now it has one subtlety. Uh, if you want to do those HTTP requests against the AWS um API, you need to sign those requests uh with a specific AWS um you know uh defined uh signing algorithm. Now, typically, this is something which I would not have rebuilt myself um traditionally because you know it's easy to get wrong and um yeah, you know, the S3 SDK would do this for you out of the box. But as it turns out, this signing algorithm it has a very well-defined spec and also like a test kit. So you really with an LLM you kind of can one-shot like this signing algorithm, and then you have it by yourself and you don't need to bear any dependencies. So that definitely you know it changes those, let's say, make or buy uh decisions.

SPEAKER_00

The the agentic uh uh augmented coding, whatever you want to call it. Yeah, yeah. That is a really good point. There are there are things that it just it seems seems cheaper to just rebuild.

SPEAKER_01

Um, and of course, but then you want to also draw a line, right? So I mean, like you have been using this stuff a lot, right? So you will know there's a gazillion of different ways than how you can actually authenticate against uh AWS. It could be like with just like your uh, you know, uh essentially like uh um user ID and and and um secret uh ID, but it also could be like an IM role, like a role attached to your EC2 instance and so on. Now I felt this is not something I want to rebuild because it's just like so open-ended and AWS, they could change this at any time. So there are actually I am falling back to something which they provide, but you can be much more conscious and more meaningful in terms of making um that decision.

SPEAKER_00

Is that the hardwood AWS auth module that's the optional module?

SPEAKER_01

Yes, exactly right. So um, you know, if you wanna let's say you want to use um IM identities, then you could pull in this optional module for this particular um because doing that in a zero-dependency way in your judgment was was just over the line.

SPEAKER_00

You you had to make a judgment call there.

SPEAKER_01

Exactly right, exactly right. So it's not it's not black or white, right? So um it's about trade-offs.

SPEAKER_00

Yeah, I this is uh a super interesting uh super interesting thing to do. I I'm guessing that you're learning a lot about about when to make that call. It it's uh it's sort of a uh the word we're using the word taste a lot, but it's almost a a taste kind of thing. Like I just feel like this would just be stupid to rebuild this, and so we'll take on a dependency. Um but defaulting to not taking on dependencies.

SPEAKER_01

Exactly. That's kind of the exactly right. That's the mental model. Um, you know, let's default to not. But then another example is um parquet files, they support all kinds of different uh compression algorithms. Um and again, there felt there's just no uh point in like reinventing the wheel and like re-implementing like uh Gs and all those uh different ways for compressing and decompressing files. So again, this is like an optional dependency, so you would pull in just those compression dependencies of the files you actually um have.

SPEAKER_00

So if the file is using compression, then you have a dependency to decompress. Exactly right, exactly right. Okay, okay, so nice. And it's not it's not a thing, um it's kind of a closure thing, is how I think of it. You know, this I'm not gonna have any dependencies, it's just gonna be a library, it's it's just not a thing I see in the Java community a lot. So I I appreciate your experiment here.

SPEAKER_01

I think that's I mean, it's it's it's um I think it should be more common to be honest. Uh uh one of my uh shower thoughts is this entire thing where we automatically pull in transitive dependencies, maybe we would be better off if that were not the case. So if you know you know you pull in library and instead this one also then pulling in like three other libraries, and those three, they pull in like you know, uh nine other libraries, whatever it is, right? So if you were not to do that, we had to be much more conscious, I think um that would be would be better, maybe.

SPEAKER_00

You were uh do you feel like you were a little late to the game or a little conservative in agentic development? Is this your first big project doing that? Talk about your journey there a little bit.

SPEAKER_01

Right. I mean, I guess we all I mean, I feel like we all are learning. I certainly feel personally I'm uh pretty much at the beginning of all that. Um but yes, it's the uh first like really uh heavy um AI-based project I'm I'm doing. And um, you know, it's um AI first, I would say. So uh 95% of the code is uh written by AI. But then I make a very strong point about like reviewing everything, and we can talk about the nuances there. So, you know, the tagline I have mentally is like it's uh built with AI, not built uh by AI. So I want to use those tools and I want to be better in terms of using them for code review and all those kind of things. But I don't want to, you know, give away the key to the kingdom just yet and you know just send it off. Hey, build a parquet puzzle, because I just don't see how would this uh how would this would work at this point.

SPEAKER_00

Right, right. Yeah, the models are not, I think the reality of we're recording this in the middle of May 2026. The models are incredible, the models are not fire and forget super intelligences that just go and build computer programs for us based on big prompts.

SPEAKER_01

I mean, my you know, my cloud.md file where you encode all those rules and specific guidelines, it keeps like growing and growing. So I add more and more instructions to it. Um, you know, it doesn't really care by default a lot, for instance, about maintainability. So you have, I don't know, you need to wrap words, maybe you know, if like UI dialogue and you need to wrap words, so it will happily have this code for word wrapping in like three different places, uh, unless you tell it, don't do that. You know, try to extract something and try to unify it. It doesn't really care about what is a public API and what is like implementation uh specific. So how can we uh uh you know keep those things apart? So there's all those kinds of nuances where you you can't you know you say taste, so you need to have that, and you to you need to tell it uh about all those conventions which you want to enforce.

SPEAKER_00

Yeah, yeah. I think in some cases it in any given processing of a prompt or a spec or whatever whatever tools you're using, it may not be aware that there are two other implementations. You know, it's it's creating one because that's just not in context right now. Um you know, for us, we have the benefit of uh a big associative memory where you see something and you go even if you weren't thinking about it, you go, Oh yeah, that's right, I know that. Um I can be really.

SPEAKER_01

It cannot be like I'm doing I'm the first one in this project who has this has to do this word wrapping, right? Something must accelerate. You kind of have like this intuitions, exactly. Right, right, right.

SPEAKER_00

So, but hey, I mean they're amazing, amazing tools, in my experience.

SPEAKER_01

Absolutely. And you know, I right now I'm working on my flow for code reviews because people send large PRs and I want to be on on top of them, of course. And so, you know, I have uh what they call like a custom skill uh for uh you know uh how I want to do code reviews. And again, the I the core idea is I always want to be like a human in the loop, right? So the idea is okay, if I ask the agent to do a code review, then it will write it down into like a markdown file, which has like the actionable items with like uh checkboxes. And then for instance, I have another skill where I then can say, okay, please you know, uh work through those findings and check off those boxes. Or maybe if there's like a decision to be made, uh, you know, there's like an open question, should we do A or B, then surface this question. And I felt this uh works actually pretty well to kind of uh formalize how I like to think about uh code reviews.

SPEAKER_00

Okay. So you are getting external contributions. That's wonderful.

SPEAKER_01

Yes, absolutely. So there's you know, as it's always is, there's a few people who do a lot. Uh actually, one of them we'd hired Justice to Kaufland, very excited about it. Hey Ryan. Um there's other people, uh, you know, maybe they do like a thing or two and then they go away. But so far, I think like 16 people or so have contributed. So yes, it's nicely uh ramping up.

SPEAKER_00

All right. See uh uh send PRs to Gunner's Projects, you'll get a job. That's what I'm hearing.

SPEAKER_01

That sounds good, yeah.

SPEAKER_00

Now a quick word from our sponsor. Confluent Developer the Podcast is brought to you by Confluent Developer the website, which has everything you need as a developer of data streaming systems. And it's completely free. We've got curriculum, hands-on exercises, executable tutorials, the online data streaming engineer certification, also free, a way to find a meetup near you, those are free. Everything is there. I really want you to be successful in your journey as a data streaming engineer, and this is the site that has what you need. Check it out at developer.confluent.io. That's developer.confluent.io. Now, back to the show. I wanna I wanna get back to the um your your parallelizing of ingest. Now, I I can I know enough about parquet to kind of imagine how that works, but um talk to me about that. Like I my my my guess is that the obvious first approach may or may not work. So what what tell me tell me what the obvious approach to parsing parquet in a multi-threaded way is and where you ended up if it's not that okay.

SPEAKER_01

So maybe uh we need to set the scene a little bit and talk just a little bit about the structure of those uh parquet files. So as I mentioned, it's a column of file, right? So unlike a row-based file, like let's say like CSV, which has like, you know, for each record a line or a row in the file, here we store all the values of one column consecutively, right? So let's say you have customer data and you would have like all the names of your customers, like one after the other, then all the birth dates and so on. Uh why are we doing this? Well, first of all, it uh tends to compress very well. So, you know, you could, for instance, do things like uh delta encoding. If you have like timestamps and they are like ordered, you could just store, instead of storing each timestamp, you could uh just store like the difference between two of them, right? And it would be much smaller. So it can be stored very efficiently. And also, if we think about like analytical use cases, oftentimes you want to aggregate all the values uh or filter all the values from one column, right? So now if all those names or all those ages or whatever are like written consecutively, we can just take that one column, um process it, project it. It's it's very efficient for analytical um uh workloads, right? So that's the that's the background. Now, how would you paralyze it? I mean, you could think about it. Well, um well, I could paralyze at the column level, right? So I could have, for instance, um two threads and one processes the name column, one processes the um age column. And this works.

SPEAKER_00

That's the approach that came to mind, right? When you said multi-threaded, I'm like, oh, that's cool. I guess you would just have a thread per column.

SPEAKER_01

Right.

SPEAKER_00

And I I sense a trap.

SPEAKER_01

So Right, yes, that's a trap. I mean it's it's it's how I started. It's a very quick win. It is a you know a tangible uh improvement over a single threaded, but also it hits a limit pretty quickly. Because the thing is, um, maybe you just don't have that many columns to begin with. Um so you have those 16 cores, but maybe your file just has like eight columns, or maybe uh you're just interested in two columns, right? So it just is you know not fine-grained enough in terms of uh parallelism. So that's that's one problem. But the other problem is that also different columns, depending on the type, the data type, and also depending on the compression and depending on the encoding, they are more or less CPU intensive to decompress. So it just you know, if you want to read a thousand values from one column depending on all the things, it might take twice as long to do it for that one column in comparison to another column. Which means if you Pallads just at a column level, essentially, you know, one thread would be done and it would just sit idle and wait for the slow column to finish. So that's that's also a problem, right? And what I'm doing here is uh um I'm actually I'm going one level uh further down, and uh the thing is those columns again they uh uh store the file uh the data in what's called uh pages. So you know, a page essentially is like a self-contained uh block of data. And now this is the level of parallelism I I apply. Now let's say you know I have two columns and uh one of them maybe has like 20 pages and the other one has maybe like two pages, then I can assign essentially more threads to that column with the many pages, or maybe with those pages which are slow uh to process and give less resources to that uh faster column. And that's essentially the core uh idea which I'm applying.

SPEAKER_00

And pages are an inherent part of the parquet format.

SPEAKER_01

Exactly right. Yes.

SPEAKER_00

Okay. So you just pull in a bunch of pages and divvy them up to threads and let the threads keep the threads busy, as many threads as well.

SPEAKER_01

Exactly. That's that's the uh you know, on a very high level the idea. Then there is this notion of faults called um predicate pushdown. So um, you know, you want to query those files, and now you don't want to load all the data uh and just then to discard the the chunk of data you're not interested in, right? So let's say um you're just interested in customers uh you know who live in this, we have a specific uh zip code for the sake of the example. Um and now what parquet files also can have, they can have uh statistics. So this parquet file will tell you very efficiently uh which of my pages actually have data uh in that zip range I'm I'm looking for. And now if you think about like remote files on a three, if I don't have to download those files, uh which I'm not even interested in because I know they only contain zip codes I I don't care about, then of course uh you know things are much more efficient. And again, like all this predicate f pushdown and uh then also um you know like filtering the remaining records, all this is happening in a multi-threaded way.

SPEAKER_00

So Park, uh sorry, um Hardwood is aware of the predicates in the query or potentially aware of the predicates of the query.

SPEAKER_01

Exactly right. So there is uh a simple um a query API. So it you know it supports projections, it supports all kinds of um like filtering less than, greater than, equals, and so on, um logical like end or or uh grouping. Things. And now if you were to build like a SQL parser on top of that, or a SQL query engine, you could employ that predicate push down capability to do exactly that.

SPEAKER_00

And of course, there is there's already any number of such query engines that'll run queries on parquet. And hardware presents itself as now a more efficient way to access the underlying parquet files.

SPEAKER_01

Exactly right. So that's definitely one of the goals I would have. Once dust has settled a little bit, we have uh maybe done like a you know f a stable release. We also have write support, so currently we only have free support, but let's say once we have write support, then it would be very interesting to see okay, what would it take to integrate this into Flink? And you know, I did some explorations already uh with the Flink uh parquet modules. Um because then again they would benefit like from like zero dependencies and so on, right? So I think it's a very interesting proposition.

SPEAKER_00

Are there any more uh parallelism wins that you can see that are on the horizon now that you've you've got pages figured out that that seems optimized, but uh what what's what's remaining there?

SPEAKER_01

Right. I mean there is of course this entire uh notion of uh ZIMD instructions, uh single instruction, multiple data. So you know, CPU instructions which essentially work on multiple values all at once. So let's say um again, you know, you want to filter data, you're only interested in specific, I don't know, customers who are older than 18. Now instead of looking at each age value one by one with such a ZIMD instruction, uh the CPU will actually be able to look at eight values or 16 values or maybe uh 32 values all at once and essentially give you this data at a very fast uh pace. And this is something which you know Java um not really supports at the moment. There's uh what they call it. That was my next question.

SPEAKER_00

How do you make Java do this?

SPEAKER_01

Right. So they have like what they call like an incubating API, uh um which gives you that SIMD capability, but um yeah, you know, I really hope it's gonna be like stable at some point and that we really can rely on that and and utilize it.

SPEAKER_00

Um it'd be nice if the JIT just figured that out and just did it. Right.

SPEAKER_01

And it all it actually does that. So yes, that's a very good point. So there is what's called auto-vectorization. So the Java uh hotspot uh uh engine and the compiler is actually very smart. So if it detects certain looping patterns, it will sometimes auto-vectorize the code and actually execute it into you know a ZIMD code. But um it's kind of brittle. So you know it's very easy, for instance, to mess this up. So let's say you have this code uh and then by means of looking at the uh actual machine code, you have figured out okay, this gets auto-vectorized, and you feel very happy. Then maybe next day your LLM comes along and it changes just a little bit of that code, and then for whatever reason the compiler cannot auto-vectorize it any longer. So uh yeah, it's you know it's tricky to rely on that. Got it.

SPEAKER_00

So it'll be an annotation or something like that.

SPEAKER_01

Uh uh, yeah, big comment. Don't mess with this code, leave it as is. It's it's written in a very specific way for a very specific reason.

SPEAKER_00

Well, I'm sure we should direct a tweet at Brian Getz and say, you know, this would be so easy if you would just make this one change, it would fix everything. I know.

SPEAKER_01

Oh, yeah, he would he loves those. He loves those comments. He loves those.

SPEAKER_00

Yeah. It's so easy. Just do this.

SPEAKER_01

Um Have you ever thought of doing that, Brian? Yes. Oh, yeah. Right.

SPEAKER_00

Yeah, well, you know what? Uh yes. That is that is his answer.

unknown

Right.

SPEAKER_00

Um nice. Okay, that's exciting.

SPEAKER_01

Uh yeah, absolutely.

SPEAKER_00

Okay, so this is fundamentally an API. Uh not your first rodeo there, but um what uh what are you what are you learning there? I mean, like there's a choice between sort of row and column. I I know it's column or format, but but you can still expose a row API, um, because that's how a lot of people think. So uh what's your thinking there?

SPEAKER_01

Right. I mean, we actually do have both. So we have what we call a row reader API and a column reader API. So, yes, because many times data, you know, you want to reason about it in terms of uh rows, and actually Parquet also supports nesting, and you can have like a record which has a list of lists and or list of structs and so on. So it can be like, you know, really like uh let's say JSON mentally, like deeply nested data. And this you know, um can be expressed nicely in a uh row-based uh fashion. So we have that. Um, but there's also a you know um um a price uh to pay to assemble that representation. And for uh use cases where you just want to crank through the numbers and do this analytic workload, then we have uh this columnar API, which essentially gives you yes, uh batches of the actual values, so like you get like an array of doubles or you get an array of ints, and then you can you know process them. Not box actually.

SPEAKER_00

Those are primitive.

SPEAKER_01

Absolutely, yes. So that's actually also another thing. Uh, you know, there's this project in Java what they call Valhalla, uh, which is about uh what they call value types, and uh at some point it would allow you to have like generic code which operates on uh primitive data types. We don't have that yet, so in uh at this point in time we actually have to duplicate quite a bit of code because we want to have like you know execution flow which is optimized for int arrays and for um double arrays and so on. And we don't want to box this into array lists of integer because it would be like very slow.

SPEAKER_00

Yeah, and that's not if anybody knows you. It's not how you do things.

SPEAKER_01

Right, and you know, that's I mean that's one of the goals to be able like real fast. We also want to like you know um uh set up like um uh you know regression detection in terms of uh performance regression. So uh you know if I don't know I do a silly change or my LM does a silly change uh and it uh like causes a slowdown, we would want to be able to figure this out right now. It's a bit more manual and we have a few tests which you have to think about to run um after like changes which may be perf impacting or not. Um but yeah, you know, if this was automated, it would be much nicer, of course.

SPEAKER_00

Yeah, yeah. And on the uh the the topic of the API, I think you recently reworked the the read, the reader. Well, there's not a writer API yet, so yes, it's the only API, uh, into being a builder. Uh right. And uh so aesthetically, I'm a fan, I'm never gonna complain. Um but what what do you feel like you got wrong in the first past that?

SPEAKER_01

Yeah, I mean, so the thing is, you know, uh if you build this uh reader, uh more and more app options uh uh we added to that. So first of all, it was just about okay, give me a specific set of uh columns, so like a projection in terms of uh query languages, right? Then later on we added this notion of uh predicate pushdowns. So we had like an overload of that get me a query builder, uh get me a row reader method, which takes like this filter expression. Then we added the capability, I want to have just a certain set, a certain number of um uh rows. I only want to have like 10 rows or 100 rows, uh, because we also actually have a you know like this interactive CLI tool which uh lets you take a look at your parquet files. And for that, yeah, you want to have just like uh you know like a window of data, also maybe with the starting offset. So, long story short, we added more and more options, and then we had like, I don't know, like five or seven or even more like overloaded methods uh to get a query, uh like a row reader. And at that point, uh we decide, okay, you know, let's go back to the drawing board and let's have this builder and make it a more pleasant experience to you know provide those options in a safe way and also like in an intuitive way, right? Instead of having like all those different methods with all kinds of different parameters, and it's just hard to figure out which one you want to call.

SPEAKER_00

Yeah, yeah. That makes sense.

SPEAKER_01

And we are still like pre-1.0, so you know, um, we take kind of the liberty not to nearly really change stuff, but if there is like a meaningful improvement or we figure out we have like uh cornered ourselves, uh yeah, then we just take the liberty to to break uh yeah, you know, you're in the process of exploring the space right now.

SPEAKER_00

So you that's that's you uh you have that right. Not again, irresponsibly, but you you you don't know what it's supposed to be until you build it.

SPEAKER_01

Exactly right. And we want to move fast, right? So we also don't want to like iterate for four weeks uh upfront and then ship something. We want to you know be like really lean and mean and get something out, and then yeah, uh just I mean uh honestly that that that just kind of as a somewhat outsider to the project, it really looks that way.

SPEAKER_00

It looks like he's gonna have this thing built in nine months and it's gonna be pretty much complete and incremental improvements after that. I mean, it it you just I could be wrong, but um you've uh you know you've done this kind of thing before. And it looks like it's going real well.

SPEAKER_01

Yeah, I mean, thank you so much. I mean, that's that's the idea. I also like you know, again, like amazing uh people who are working with me and you know they add to the pace. But actually, just today I had a very interesting conversation with somebody who also uh contributes to the project, but not a lot. More like, you know, uh things uh here and there. And so he actually said, uh, I have a hard time to follow along because like you guys are working with such a high pace and stuff changes and you rework the architecture, and it's it's a hard time to follow along. I felt actually that's a very serious point because with those you know, LM tools, you are just so fast and you can ship this 5,000 lines of code change, uh, which is a meaningful improvement. Um but yeah, if you have those like every second day, it's really hard for people who are not like you know every day involved uh to follow along and stay in the loop. Absolutely.

SPEAKER_00

What's the You mentioned AI again, and I can't help myself. What's the worst thing AI has done in the project so far?

SPEAKER_01

Oh man, there I don't know. Sometimes I'm wondering whether it's me or whatever, but I've had like really funny things. So, you know, I mean there's this thing, I mean it's trained on human uh artifacts, right? And sometimes it just uh shines. So for instance, what it initially did is uh sometimes it would just um you know not fix uh certain failing tests. It would say, okay, I'm doing this change and I see this test is failing over there. Um it's not my department. It literally told me that, not my department at some point. And you know, I just live with that. And then I literally added the rule to my Claude.mt file. You know, that's not how we want to do it. If we see something which is broken, then we gotta fix it. Also, if it's not our current work which may have been been causing this. At one point, it uh you know, I was iterating with Claude on some design or whatever, and then I said, after a while, okay, yeah, I feel we have uh reached a good point. Let's let's call it a day. No one's like, do it. No, I just now I also just want to build it, you know. So don't hang up on me.

SPEAKER_00

What um uh I I I identify with Claude there on the You know, uh those tests, they're okay. I I make sense, I'm fine. I don't really want to not into it.

SPEAKER_01

Yeah, I mean, you know, I guess I guess it's just uh human, but then also there's a thing, so you you know, sometimes you do like this uh involved design, and then it says, Okay, this is gonna take like two weeks to build. Then you say, Okay, Claude, go and build it, and like then 10 minutes later it comes back, okay, I'm done. Like you always wonder, okay, where are like those super weird estimates coming from? Right, right.

SPEAKER_00

And do you reserve tests for humans to write, or do you have Claude write tests based on specs you give it?

SPEAKER_01

Yes, uh so um it's mostly written by Claude, and actually that's also one of the reasons why I think this kind of project that really lends itself very well towards uh agentic coding because actually the Parquet project and the upstream community they provide a very, very comprehensive uh suite of uh test files, so you know, which expose all the different encodings and types and uh compression algorithms and uh variant and so on. And so you know, you just can implement like really meaningful coverage, test coverage by uh essentially uh using those files. And then I you know I just compare the output which uh hardwood gives to the output uh of the upstream Pocket uh project. And of if there's a difference, well then obviously uh it's on us uh to fix it. Um and this worked uh really well. So initially, you know, until we had like the full coverage of that, you could just tell, okay, now you know look at this test file and make it work and it would go uh and then uh implement like the changes in the engine.

SPEAKER_00

Um a few more questions. We're we're getting close to time, but I I can't help myself. There's a few things I want to know still. Um I mean this has been a fair amount of work, it seems like, even with augmented agent-augmented code. So the what's the lesson here for the people who say, well, like SaaS is dead, there's no moat, we can just vibe code whatever we want. Um you're you're heavily relying on agentich development. Right. And yet this is a months-long project with an interesting roadmap ahead of it. So just talk about moat and AI and with respect to this or any any project. What do you think?

SPEAKER_01

Absolutely. Yeah, I don't know. I mean, I I mean, personally, I first of all, I I think I also underestimated uh the uh effort it would take. Uh so initially I thought, I don't know, maybe by this point in time, I started you know, around January, December, January, and maybe by this point in time, May, uh, we are already further ahead. But yes, I mean, as it turns out, this file format it has just a very long tail of you know different capabilities. There's uh things like Bloom filters, which we don't even support yet, there's encryption and so on. Um, so there's just a long tail of stuff you just need to work through. So there's that. It just is a large surface. But then also, um, yeah, you know, I think without supervision, it just wouldn't get uh those things right, like what's a clean and well-defined API, what is a consistent API, and we don't do like you know similar things uh very differently and different ways in the in the code base and in the API. How do we keep this maintainable? Um so I think it's it's very important to have somebody in the loop who you know is on top of that and who then also like codifies all those things in skills, uh, Claude MD and so on. Um but yeah, also um, you know, it's just a lot of stuff to get right and to make it efficient. So I mean, yes, you could, you know, you could say, okay, I want to have a parquet reader for like those particular files which I have, and you could point Claude and it would give you something like pretty quickly. But you have like a comprehensive engine which really supports that long tail of uh features, which is efficient, which is maintainable. It's a lot of work, and I don't see you know any at any point soon that you would say uh Claude built me that, and it would you give you that what we have done now in multiple months with multiple people in a half a day or what?

SPEAKER_00

Yeah, yeah, yeah. Do you recommend it for prime time yet? Again, this is uh the the the date is important. It's May 11th, 2002. I think it's the 11th, 2026. Um so should who should use it? Who should not use it?

SPEAKER_01

You mean uh hardwood, right? Uh yeah or uh yes. Um hardwood.

SPEAKER_00

Did I say parquet? No.

SPEAKER_01

Uh no, I must show uh um for the record.

SPEAKER_00

Uh it's it's probably really safe to adopt parquet. Uh it's fine. Go ahead. Um, and I don't know.

SPEAKER_01

You don't want to rush into these things, but I thought you might uh refer to the uh LM based uh coding tools.

SPEAKER_00

Oh, oh right. No, also agentic development. Uh let's do it. Please get it. Right. Get get on board. Um no uh hardwood is it ready for prime time.

SPEAKER_01

Right. I would say um you know people should probably wait for 1.0 to put it into Anger for their uh reading use cases. But I mean by by now I think it's pretty stable, uh definitely efficient. So I would love for people to give it a try and you know parse their files and uh use it. Uh and you know, if they run into any kinds of problems, uh re-report back. That would be great, right? Um, what I also realize for people to really adopt it, they want to have write support, right? I mean, let's say uh we want we wanted to integrate this into a flink. Um as long as we only support reads, it's just not a good purpose, right? Because then they would actually have like two dependencies, even more uh you know, hardwood for reading and the existing one for writing. So that's not very attractive, right? So we need to have write support that's on the roadmap for 1.1. And then yeah, I would hope uh people will adopt it quickly. We also have a compatibility layer. Uh so essentially, you know, we kind of like re-implement the existing parquet Java APIs on top of hardwood so that it's kind of like a drop-in replacement. Um, yeah, I think I feel pretty good about it.

SPEAKER_00

Good, good. Wow, okay, that is that is awesome. Uh benchmarking. Um, uh are you afraid of embarrassing parquet Java? So you don't want to run benchmarks against it? Are you afraid of embarrassing yourself? You know.

SPEAKER_01

Yeah. Um, no, I mean neither. So the thing is, there's this entire notion of uh benchmarketing, right?

SPEAKER_00

And uh or what they call like uh liars, damned liars, statistics, benchmarks. That's Mark Twain gave us those four.

SPEAKER_01

Yes, exactly right. So that's all that. So I want to be very, you know, uh conscious and cautious about whatever numbers uh we uh publish. And you know, the first thing I think which could happen is you publish a benchmark, and you know, let's say we uh we uh show pocket Java does X and we do Y and we say hey we are much better and then somebody says oh yeah this is because you have hold it you've held it wrong and you should have done you know whatever option or knob in PoKJava and this would have been much better. So you know, um actually this comes back to what I call the benchmark benchmark or paradox because the thing is, you know, if you wanna if you wanna compare different tools, typically you don't have the deep insight into each or every one of them to like uh tune it well and use it perfectly well. Um so there's a challenge. On the other hand, if you do have that uh level of insight for one of the tools, then maybe you don't want to compare it in a benchmark where you know where another tool would win. So there is like a there's a bit of a problem there. But ideally, yes, we get some numbers and we would run it uh by the parquet community and we would invite uh you know uh people to to contribute and and validate. So that's that's the idea.

SPEAKER_00

Yeah. Okay. So you are um this is open source, right? It's and and yes, you're you work for confluent. Is this a confluent thing? Is it a you thing? Is it what what how do people think about it in that way?

SPEAKER_01

It's a it's a good question. Uh it's not an official conflict uh project at this point in time. I mean, you know, I've been uh working on it uh on my uh confluent time, but it's um not uh yeah, not like confluent supported or whatever. Um I don't know, maybe at some point it would be, uh would would be great.

SPEAKER_00

Um but yeah, at least they'll they'll use it, you know, it's uh tableflow uh.

SPEAKER_01

Right, exactly right. So I mean tableflow. Actually, I spoke to those now that you say this. Uh um I think it would be very interesting for them to integrate this into Tableflow and make use of that engine.

SPEAKER_00

Cool. Uh last question. Most interesting thing you've learned about parquet. Uh, not not just the JVM, but about Parquet, the most surprising, interesting thing in the process so far.

SPEAKER_01

About parquet. Let me see. Um I think, yeah, really that you know, just uh that fact that it is this massive thing which has grown over a long time and it gets you know uh more and more capabilities. Um so I completely did not expect this. So for instance, just recently we added what's called uh support for variant. So that's uh like a self-descriptive type, which kind of you know takes um it's a bit like JSON mentally, it can it can be like arbitrarily structured, um, and you can have those variant columns within a parquet file. And then at the same time, uh you also they could actually be uh manifested in uh what's called a shredded column. So then again, it's actually stored as a proper uh parquet column. So I realize that's confusing. But so yeah, you know, there's this long time, sort of this long tail of like all kinds of capabilities, and I don't see them stopping adding this new stuff. And this uh, you know, it's not like okay, let me uh implement this thing and then I'm done. Uh no, it you know, it is an ongoing thing. And I think this was definitely something which I didn't expect to that extent initially.

unknown

Yeah.

SPEAKER_00

Awesome. Um, well, my guest today has been Gunnar Morling. Gunnar, thanks for being a part of the Confluent Developer Podcast.

SPEAKER_01

Jim, thanks so much for having me. This was great fun as always. Thank you.