Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov

Deleting Architecture for Better Systems ft. Daniel Doubrovkine | Ep. 19

Confluent Season 2 Episode 19

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 26:05

Adi Polak talks to Daniel Doubrovkine (Shopify) about his career building data‑intensive systems. Daniel’s first job: delivering pharmacy medications by bike. His challenge: building Artsy’s Art Genome and auctions as simple as possible.

SEASON 2
Hosted by Tim Berglund, Adi Polak and Viktor Gamov
Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
Music by Coastal Kites 
Artwork by Phil Vo 

  •  🎧 Subscribe to Confluent Developer wherever you listen to podcasts. 
  • ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
  • 👍 If you enjoyed this, please leave us a rating. 
  • 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
SPEAKER_01

Today we're exploring what happens when art meets code and why simplicity might be the highest form of engineering. This is Confluent Developer.

SPEAKER_03

What really mattered was data and not the actual implementation. Reducing complexity and abstracting away is, I think, the key when starting something earlier. I made so many mistakes. I remember this being like I will need to automate my way out of this one day.

SPEAKER_01

Hello everyone, I'm Adi Polak and welcome to Confluent Developer, where we uncover the human side of software. My guest today is Daniel Dobrovkin, or DB, a lifelong builder whose story spans continents and generations, from refugee roots in Moscow to the hacker underground of 1919's Geneva to leading the creation of Artsy Art Genome Project that exists till today. We talk about building search engines before Vector Database existed, keeping developers happy, and why sometimes the best architecture is the one you delete. Let's get to it. Hey Daniel.

SPEAKER_03

Hi D, how are you?

SPEAKER_01

I'm good. Or whether I say DB.

SPEAKER_03

Oh, that's right. You could call me DB.

SPEAKER_01

That's awesome. You know what's the story behind DB?

SPEAKER_03

Oh, it's uh it the B is after Alexander Bloch, the Russian writer. I uh joined this BBS in the early 90s called Boris Co. in Geneva, Switzerland. I needed a username, so I I couldn't come up with one. I asked, I thought I thought about like what what kind of pseudonym can I use in this uh software pirating BBS, and uh uh I used Block, which is a uh a pseudonym that my father used for quite some time, uh, but I spelled it B-L-O-C-K. And then it needed to be six letters, so it became D block, which stuck with me through high school, college, and so on and so forth. So now I just go by DB for the last 30 or some years, and even my kids call me DB sometimes.

SPEAKER_01

Wow, that's uh that's quite a story. Um super interesting. Thanks for sharing. I'm curious how how the kids react to that. So it's uh definitely mostly roll their eyes. Yeah. Cool, cool. So welcome to uh Conflict Developer Podcast. I'm super excited to have you on board. I know you have many, many years of experience. Uh, but the fact that I know you doesn't mean that our audience, you know, are know you well. So maybe you can share a little bit about yourself uh so they they know you better.

SPEAKER_03

Yeah, most definitely. Um I was born in Russia, in Moscow, and my family immigrated to Switzerland when I was 13 years old in 1990. We're refugees in Europe. Um I had no friends, so I picked up computers as my friends very early. Uh, joined uh some hacker types uh, you know, who also had no friends, and we we made quite a group. Then uh eventually I started writing software, copying my very first one from SVM magazine. I remember copying like two pages of x86 assembly, but I had no idea what that meant. And then it all it all worked out. So I ended up having a couple of companies, a long career in tech, um, mostly alternating between IC and managerial types. Uh so I kind of go back and forth, and uh I think that's a that's a common pattern in my history. I live in New York City in Manhattan, have two teenagers that mostly roll their eyes on my on my nickname.

unknown

Yeah.

SPEAKER_01

Or or in anything. I mean, teenagers is uh is a pleasure, right?

SPEAKER_03

We could talk about teenagers during this entire podcast if you want, but I am probably the wrong person to give any advice about teenagers. I'm doing my best. It's still a mystery to me. None of them codes, and this is you know, I'm not sure if it's because I tried to convince them that it's a good idea or because I didn't try hard enough. I don't know. We can talk about it.

SPEAKER_01

Yeah, it's uh it's interesting, but it's I I guess that's that's a different mystery. Maybe one day we'll figure it out, but I'm not I'm not sure if we if you know parents will ever, even though we ourselves been teenagers uh a while back.

SPEAKER_03

So anyway. In fact, they they do remind me of my parents who used to yell at me, get uh get away from the computer, you know, this is useless, nobody cares. What why are you wasting your time spending your entire day in front of the computer? My my kids say that a lot. It's like, what do you even do?

SPEAKER_01

Is that a real job?

SPEAKER_03

Exactly. Like you're doing nothing all day, just sitting in front of your computer.

SPEAKER_01

Right. By the way, jobs. I was curious, what was your very first job? Do you remember it?

SPEAKER_03

Yeah, I uh I did some non-tech jobs very early. I I delivered pharmacy um medications on my bicycle in uh very, very early in high school. But I because I quickly got into computers, I thought I'd get uh a job in computers and had an internship at this Swiss bank. Uh I think I was 15, and my job consisted for three weeks to copy uh numbers from a paper ledger where somebody wrote them with a pen into a very early version of Excel. And I made so many mistakes. I just remember this being like I will I will need to automate my way out of this one day. There's no way this could be could be a job. Um that was my first proper corporate job with you know a paycheck, uh, an equivalent of you know UW2 and something like that. And then um I quickly pivoted into working in this computer store, mostly uh uh assembling PCs, uh being paid cash under the table because I was you know kind of underage, but I knew about computers, so there weren't that many people in the early 90s that uh that did. So I I really I really enjoyed that job is working with people and uh selling and assembling PCs in one of uh Geneva's um computer stores.

SPEAKER_01

Yeah, I mean you gave them value, and so it's uh back in the days it's hard, like you mentioned, it's hard to find people that that expertise. And even today, I'm not sure if you know the new generation knows how to assemble things or to just go to a shop and and buy stuff.

SPEAKER_03

I probably don't know how to assemble a modern PC anymore. It's been it's been a while. It's really fun. I I made so many friends in in that little nerdy crew of uh, oh, you know, you've got the new 486, it's got so much RAM.

unknown

Yes.

SPEAKER_03

Compared to today, 40 megabytes by second 250. I think I spent like $2,000 buying that.

SPEAKER_01

Wow, wow. Crazy. Today's like, you know, can you run AI on it? No, okay, what is it worth?

SPEAKER_02

You can run this very slow AI on it. Very early LOM, maybe.

SPEAKER_01

Yeah, extremely slow. Uh not a lot of uh tokens to process and so on.

SPEAKER_03

Yeah, things were simple then. Things were very simple.

SPEAKER_01

Yeah, that's true. Things were very simple.

SPEAKER_00

Now a quick word from our sponsor. Confluent Developer the Podcast is brought to you by Confluent Developer, the website, which has everything you need as a developer of data streaming systems. And it's completely free. We've got curriculum, hands-on exercises, execution tutorials, the online data streaming engineer certification, also free. A way to find a meetup near you, those are free. Everything is there. I really want you to be successful in your journey as a data streaming engineer, and this is the site that has what you need. Check it out at developer.confluent.io. That's developer.confluent.io. Now back to the show.

SPEAKER_01

Let's jump into the future for a little bit. Um, what was some of the software challenges you faced? And if you can walk me through like how you solved it and what's some best practices and learnings that you can share.

SPEAKER_03

Yeah, I think um a good job to talk about that had a lot of technical challenges that I dealt with very personally. Was my time at Artsy, it's ARGSY. I was uh I joined the company when it was in 2011, and it was maybe a seven people company, and we will go and build the largest fine arts marketplace and most red online arts publication over eight years, having raised you know hundred million dollars and having many hundreds of people working at the company. But um, some of the the initial challenge of that of that project was was super interesting. Uh, it was called the Art Genome Project, where art historians would classify works of art according to a uh quote-unquote genome. So you'd have these features, um, let's say abstract expressionist work, and you'd have a value zero to a hundred of how abstract expressionist it is, or let's say pop art, and artists and artworks would be classified that way. So, you know, Andy Warhol would get 100 in pop art and so on and so forth. So there were about 1200 of these genes, and what we were trying to do is a similarity engine. And I didn't know much about uh KNN vector search at the time. And some people that preceded me built a very heavy prototype on um JBoss. They wrote this like crazy thing in Java, very over-engineered, very enterprise software-like, uh, that you know took 20 minutes to boot, load all the data, try to build this nearest neighbor graph, and then you'll be able to search through the graph. But almost nothing worked. So uh when I took over the project, uh uh we we we restarted from scratch. Uh we I went around New York startups asking how do people build startups today? Everybody told me, you know, you should be using Ruby on Rails, MongoDB, and that's how we rebuilt the initial version uh of search. It's a was a very brute force K and N vector search uh in its version one. And then uh eventually it became more sophisticated. We used something called locality-sensitive hashing, and then finally an algorithm called NN Descent. Um spent probably two or three years building various iterations of the search engine. Um I wrote quite a bit of it uh at the time myself, but I guess the conclusion of all of this was that in the end, what really mattered was data and not the actual implementation. The data made the engine really successful, like high-quality data entered by humans. And then finally, all of this was replaced by off-the-shelf nearest neighbor uh search implementations. You know, we used Elasticsearch more like this package, and then now you have gajillion uh nearest neighbor vector search engines that can do much, much higher dimensions and return results almost instantly. So all this all this exciting new technology that we had to develop from scratch is completely commoditized today and is available almost instantly. Um so that was a fun technical challenge because we had to uh to produce results that were not only fast and but also good. And good was very subjective. We'd put art historians in front of uh of the results and say, Hey, what do you think of uh of of this search result? Here's an artwork, here are similar works by other artists. Like, how do you feel about it? And they would they would feel so passionate and yell, like, oh no, how do you dare put artworks between this artist and that artist on the same page? Like that is absolutely unacceptable. So they'd say, like, wow, how did you connect how did you find this amazing connection between these two artworks? Like this is magical. Uh it's like a museum-curated show. So that was fun.

SPEAKER_01

Yeah, it's a it sounds like a very hard thing to do. I was curious, you mentioned there was different things that you had to re-implement. But first, you know, the fact that someone or other companies already implemented and made it a commodity canon and other algorithms, it's it's fascinating by itself, I think, because you know, if you look at the world of open source, there's so many things already implemented. And for us, we can lead with data and maybe tweak some things that already exist there. Um but I'm curious because you implemented a lot of core stuff. Is there anything that you remember kind of like profound, like maybe some of the hashing capabilities or anything like that that you want to dive a little bit deeper, or perhaps um, you know, meeting some of the founders in in the space and learning what are they doing differently?

SPEAKER_03

I mean, we we we I eventually hired people who were better than me. Like I'm my my computer science background is limited to college. Like I did graduate in computer science and then promptly forgot all the math uh that that came with it. And so eventually we hired people who were you know PhDs in math and things like that to do all the all this work. And um most of the time it involved reading papers and trying to implement the papers as opposed to um necessarily like invent our own. So it's not like we invented some new way of doing nearest neighbor search. Uh what we did have to struggle with is the is performance of these things. Because it's one thing to take a paper and re-implement an algorithm, it's a completely different uh different performance pattern when you actually execute it. So when you're working on real-time systems, you have to return search results within some milliseconds later. And that that was definitely challenging in the early days of of these vector search implementations. Um and I didn't know the future of LLMs which use this kind of technology. Like none of that was we we we could have never guessed it would be all that. So it became commodity later. But while we were working on it, there was there was nothing off the shelf or very little off the shelf that worked for us. And we felt like we needed to build our own to have to for it to be like our core technology. In the end, it it it was only our core technology for a very brief moment, and then the world kind of flew by with off-the-shelf software that that did it. So can't claim any uh any scientific breakthroughs in the in this work. And I think it's appropriate for uh for a startup uh where ultimately the goal is that the is a delightful user experience as opposed to something you know deep and complicated. And so we focused on that. Uh I built the first version of the bidding engine uh for auctions, and uh I I tried to apply the same approach of like the the nearest neighbor search was way too complicated when we started. We thought it would be like such a hard problem, and so we looked for a very hard solution. And when it came to building uh a live bidding engine, um, I heard the same things from all kinds of people. I went to Wall Street and everyone was like, oh, an auction, a live auction, that's you know, that's what we have on the stock market. That's very, very hard. You need low latency, you need a lot of memory, you need to solve for a lot of concurrency. And then I went to art companies and like, yeah, we use paper, and uh you write your bids on the bottom of the page. Like, what's it what's in between? So I opted for almost the paper version, so something extraordinarily simple, just built on top of our uh of our existing implementation for e-commerce. And um and practically this was this was just a single-threaded background job that would run the auction while people could place bids. So I separated the user experience from the actual part that needed to be really fast. And uh this this was this was a good idea. Uh in the end, people would not notice that something was happening very slowly. We faked it in the user experience, like in the UI, basically. And the auction would catch up as as especially the last like five seconds of a of an auction, people are trying to snipe. It's when they try to place a bid as l as late as possible to get their to get their winning bid, to be the max bid. Um so the in the background it did we we just like run the auctions, you know, there's an animation, and uh 15 seconds late it tells you, oh, you won. So you don't need to know it took 15 minutes, 15 seconds to calculate. Like it just it looks instant. He's waiting for for this moment. So a lot of the user experience of something that needed to be super fast could be faked to look fast, just didn't need to actually be fast. Eventually, we built a system that was fast. Uh but that was maybe a version five of the system.

SPEAKER_01

Yeah, so if I understand correctly, you had to build to it, right? Because it was you knew you needed real time, but real time is hard to it takes time to build. And so building up to it with other versions until you get to uh to a real-time experience that was was that was the path, right?

SPEAKER_03

Yeah, and it's it's really about keeping things simple. And we say, oh, the KISS principle, keep it simple, stupid, but very often we don't actually do it. So our final version for the bidding engine is Acker-based, actor-based system, like real-time events, and it's it's incredibly complicated to understand what's going on in the system like that. And uh, of course, it scales incredibly well. But the first seven versions of the of the system were written in Ruby in like hundred lines of code, and that that carried us through many, many hundreds of auctions before we hit the actual scale problems uh of uh of our implementation. So it's premature optimization and keeping things simple. These these are it was it's pretty amazing to experience that in real life. Um I I know I know about these challenges. I faced them a little bit before, but starting a system from scratch and then seeing it actually happen was uh was very educational. And that that's why you go work in startups is to um is to to start something greenfield sometimes.

SPEAKER_01

Yeah, that's exciting. Greenfield and also with a very small team. You mentioned that you started with seven people and then the company grew. Uh the whole company was seven people.

SPEAKER_03

I think I was the seventh person there.

SPEAKER_01

Oh wow. Wow. And then you raised $100 million, right? Yeah, something like that. Back in the times where that was like a fortune, like multiple. It's like raising multi-billion dollars today. Uh equivalent.

SPEAKER_03

Yeah, we didn't raise it in one round, though. We did serious ABCD. So it took its time. The company is very much around today, artsy.net, it's it's an amazing marketplace for fine art. I still buy art on art. Uh you know, I I collect some some prints like Joseph Albers and uh uh artists from the 60s from uh from the past century. And I I I buy it all on art. So you have access to like all the galleries, all the auction houses. It's it's it's a pretty great.

SPEAKER_01

Very cool. So I'm curious if you had to kind of like rebuild the solution today, giving you know all the different solutions that are available on the market, like out-of-the-box canon and the data streaming solution and and so on, like what would you do differently?

SPEAKER_03

Well, I think a lot of the software today is is available because the um the time has come. You know, these systems are scaled and uh work very, very well. So now probably half a person in like half a day could assemble uh Arts version one very, very, very, very easily. And there is endless options for the software uh for vector search and things like that. I think I'd still start with my my all-time favorite, like Ruby Ruby on Rails, uh, and then pick the actual search engine, like I don't know, Elasticsearch, maybe not, uh, OpenSearch. I worked on that fork for for a number of years. Um, maybe like a simpler, newer vector search engine. Uh I still love MongoDB, still use it for my pet projects. I think these are good simple startup systems, and they um they're very low maintenance. Um I don't know why people hate MongoDB so much, but you know, it's just like garbage in, garbage out, set it and forget it, it works great. And scales to like 95% of projects that will die anyway because you know your company may not have that that future. Um so that's still my preferred stack for uh for getting something bootstrapped very, very quickly. Rails, whatever database, Postgres or MongoDB, and then maybe something specialized for the specific problem of VectorSearch.

SPEAKER_01

All right. Would you consider like a cloud service, like something that was already managed and built out, or would you rather like have the solution managed in-house?

SPEAKER_03

Well, it has to run somewhere, and of course you'll be running things cloud-based. So for my pet projects, I use DigitalOcean a lot because the app platform is super simple and you don't have to maintain the hardware. So that's a that's a viable option to get started. Absolutely. Um I I would I'm not in the camp of premature optimization, like throwing things onto your own hardware. Then you have the you have to have people who maintain all this. stuff. So reducing complexity and abstracting away is I think the key uh when starting something early. You want to give yourself options, but I I like the idea that things can run locally very, very easily. And I like the idea that my my infrastructure, cloud-based infrastructure is is uh painless and that's worth money. I just don't want to deal with it. So I want I want something extra extraordinarily simple and there is so many of these things out there that all work quite well.

SPEAKER_01

Simple and painless.

SPEAKER_03

Yes and and cheap. I think it has to be it has to be cheap and it's important to to know when these things start hitting limits of cheap. Like Hiroko used to be really really cheap to get started. Arts started on Hiroco. And then uh when you start scaling out like it becomes significantly more expensive. You pay overhead there. So like today I I would hesitate starting on that platform because I know that the limit is not that far off in terms of like cost benefit. So it's good to know where where these limits are and then avoid future migrations. But a simple architecture always wins and then something portable between services where I can say oh you know like DigitalOcean doesn't work for me I'll redeploy on some something else versal or whatever other platform out there. I don't think it's that important. But one thing that I always optimize for is developer happiness. And uh you should take the team you have like if db is the team Ruby on Rails and Ruby it makes me happy. Okay. If you are the team maybe it's Node.js TypeScript and you know something else. Others will want Python and Flask who knows I think that you should you should work with the people you have and be like what would you choose? What is your um what is your preferred framework? Like some some want to optimize for something super new and super exciting that everyone's talking about like today like the cutting edge. And that's a great choice. And others like oh I'm very conservative I want the thing that I know will work for sure. So that also works. In fact all of it works.

SPEAKER_01

Yeah it's very interesting uh engineering happiness it's kind of it could it could be a metric for engineering or it could not because it's not always you know it's hard to um explain it I think sometimes um like measuring developer happiness so like are you happy now?

SPEAKER_03

Like how happy are you?

SPEAKER_01

Yeah how you used Ruby on the Rails are you happy now yes well if you're being paged every night uh for a couple of times a night probably not happy um but yeah so many people love Go so much and I think its error handling is insane.

SPEAKER_03

And so I I make fun of sometimes of Go programmers. I'm like so you know it's more lines of code to do the exact same thing. Are you happy though?

SPEAKER_01

That's why right that's true. All right Divi thank you so much for sharing your experience it was a pleasure if there's one thing I'm taking away out of that is you know make things simple and uh easy to use and focus on the experience for the customer because that what's really brings the value uh at the end of the day. So thank you thank you thank you.

SPEAKER_03

Yeah and don't forget to make developers happy as well with what they're working on. And thank you for having me.

SPEAKER_01

Yes absolutely we always want to make developers happy uh of course till next time