Backblaze, Inc. Class A Common Stock 0 Earnings Call

NASDAQ:BLZE · Jul 15, 06:57 PM

Hi, everybody, welcome back to the Q1 2026 Drive Stats report. Today, I am joined by Laquie Campbell. I'm Stephanie Doyle. Let's do our meet the team, jumping right in. I'm Stephanie Doyle, Senior Manager of Market Intelligence, affectionately called internally as the keeper of stats. I'm here for Drive Stats, here for performance stats, and here for network stats. Today, I've got Laquie Campbell with me. Lakey, you want to talk about yourself a little bit? Let everybody know who you are?

Sure. Hi, I am Senior Product Marketing Manager focused on Media & Entertainment. Been in the creative technology role for a while. I'm also part of SMPTE, I'm the Director of Education for programming with them. Right now we're focusing a lot on AI in media workflows and infrastructure.

I'm super excited to have Lakey with me today because I feel that Media & Entertainment has always been a really important industry through which to view hard drives. We were talking about it before we went live today, it is a space where both the individual drives very much matter, as well as this top-level orchestration of data workflows. You really see multiple lenses on a hard drive when you're in media. Thanks for joining me today, Lakey.

Thanks for having me. Drive Stats are super popular when I go to shows.

Love that. A few housekeeping notes we do every time. Webinar is recorded. Please feel free to ask us questions. We'll leave some time at the end. Don't miss the resources and attachments. We'll drop them in the description section of BrightTALK once this webinar is over. You should see links and all that good stuff as well. What is Drive Stats? If it's the first time you're joining us today, Backblaze has been collecting our raw device metrics for over 13 years at this point, and we publish the failure rates of hard drives. If you go to our website, you'll see a maintained data set that has daily logs from each of the drives in our pool, and then we also filter those out and look at them on a tabular level so we can talk about annualized failure rates.

What happened this quarter, very exciting. We've got about over 340,000 drives now. We had about 1,000 drive failures and a bunch drive days, which is how many drives existed per day in a single quarter. You can see the drive population by manufacturer. It's about a third, a third, a third, which is interesting. We'll talk about some of the quarterly data here. This is the big eye-crunching table that if you go to our blog, you'll see that we actually have a much more friendly version of that this quarter with responsive tables. We're very excited about this. Top level here, you've got 1.24% annualized failure rate for the quarterly stats this time around.

If you want to just look at what's been happening quarter-on-quarter for about the last year, you can see that it's up from last quarter, from Q4 2025, but down year-over-year. We've got new drives. The 26 terabyte WDC is now in production. You can see some pretty interesting stuff. This last quarter, we deployed a little over 10,000 drives, and of those, 9,400 or so were more than 20 terabytes. We've been keeping an eye on that population because of course, as you deploy larger drives, it's got implications all the way around. This quarter, the AFR for that population specifically was 0.85%. Don't over-index on that. These are pretty young drives. We know for sure that younger drives fail fewer times. There's that. You get a couple drives here with zero failures, which is pretty exciting when you look at the age of those four terabyte drives.

Those guys are, I don't know, they've been around for a real long time. If you look at the average age in months, the four terabytes drives are 106.1 age in months. We do measure our drives in toddlers in months. That's a little under 10 years, which is pretty cool. There's only 186 of those remaining, which says to us that they're on their way out, and you'll see that they've dropped off the lifetime table this time around as well. Lifetime data, these have slightly different exclusions here. They need to be hanging out for more drive days, and have more drives in the population. You'll see it's slightly different.

What's interesting this quarter is that our annualized failure rate is 1.39%. We've been holding pretty steady, actually, for the last several quarters at right around 1.29% or so. This is a bit of a jump. The jump is actually because some of those much older drives came out of the pool. The four terabytes have dropped off a lifetime table, and like I said, those had 10 years of data behind them. As those exited, you'll see a little bit of a fluctuation because things have changed. We've got three new drives entering this as well. You'll see the 22 terabyte and the 26 have made the exclusion criteria to be tracked in the lifetime drives.

What's interesting this quarter is actually, the data's always interesting, but it's a little bit something that we were able to flag when a drive was not failing that crazy. This is the disturbance in the Drive Stats, what I'm affectionately calling this. I think what's really important is that, and something we've talked about before, is that all systems that we talk about in the reportings that we're doing here, they represent a managed system, right? We're monitoring our drives at all times. We can actually affect whether things have higher and lower failure rates. If we see something, we can take proactively or reactively change it. This is a really interesting instance of this, because what happened is that we had a pretty acceptable drive model, and we started seeing failures climb.

We actually saw it was a little bit of a confusing investigation, because once the investigation was complete, at the end there ended up being two separate issues that were happening that were very distinct from each other. The drives themselves could experience one or the other, or both of these failures, right? We identified the power cycle issue early on. We noticed that if we'd shut them completely off, they would have trouble coming back online or maybe not come back online at all. What we did is we took a mitigation step that we just reduced the power cycle frequency to the vaults of those affected drives. Once we did, the failure rate came down to a reasonable level. This is obviously not real numbers here that you're seeing.

If you can imagine, what we did is reduce the risk of one of these types of failures. That left us totally fine. Once the investigation completed, we saw that there were multiple issues happening. From our perspective, when I went to Drive Stats, I was like, "Hey, would you look at that? That failure rate is not very high, but I know something happened." I had to do a bunch of investigation on my end to make sure that we were properly capturing failed drives because I didn't want there to be a situation where we were under-reporting where we should. What that led us to is a really interesting lens on how we define a failure. We've talked about how we define a failure in the past.

The quick and dirty is essentially that there's a C++ job, the custom program that collects the SMART stats for drives each day. There's some things about the exclusion tables or inclusion tables that are maintained by humans that affect these things. At its most basic, it's conditional logic. Was the drive there yesterday? Is it gone today? If it is gone and it was here yesterday, we log it as a failure. If we look at if it's a capital F failure versus a lowercase F failure, let's say it that way, that's where those exclusion or inclusion tables come in. You can imagine if you swap a drive out for normal routine maintenance, we're not going to call that a failure because the drive did not fail. We just swapped it out.

The other thing is that if the serial number comes back online before the end of the quarter, it's no longer a failure. We took the drive offline for some reason or another and brought it back. There's a little bit of squishiness there. The quarter end cutoff means that if the drive comes back after quarter end, potentially you've got one or two false failures in there, you got to cut it off somewhere, right? What this does mean and what came into play for this specific incident that we were tracking is that if you have day one failures, you actually don't log those, right? Because it has to exist before it can not exist the next day. From our perspective, I think it's an interesting caveat to make on the dataset.

Lots of folks look at the dataset as a source of truth. If we're undercounting day one failures, we want to make sure that's well understood. On the other hand, how often do you get a day one failure? In this case, with this specific drive, there actually weren't that many, it was something that I wanted to double-check on before I came out with failure rates. On the other hand, it's pretty rare for drives to have day one failures. There's qualification drive manufacturers do. We personally do qualification on the drives. It's not something that happens all that often. Again, because this is used so widely, there you go. We've got a new caveat to the dataset. If you're using it, make sure you're checking logline data and see what's happening. All right. I think that brings us to resources here.

Got the report. Of course, follow the series. These are many ways you can interact with us. We've got the Drive Stats home base on the website. You can join the Insider's Newsletter or reach out to us directly. DriveStats@backblaze.com is a monitored inbox. I'm there. I'm always happy to respond to real humans. We've got socials in the comment section. Yeah. Let's take this time and talk questions. Lukie, anything on your end that sparked while we were going through the data?

Yeah, absolutely. From my POV, I told you many times, I always get stopped at shows and when I bring up Backblaze, they're like, "Oh, yeah, Drive Stats. I look at that every time." It's because in media, and when I say media, I don't mean just film and television. Every company these days is a media company, whether we're talking about corporate marketing, bio. It doesn't matter. Everyone's trying to tell a story visually, and those files typically are larger. In a traditional media library, a video file might carry maybe a few hundred metadata attributes. Now that everyone's working with AI, trying to figure out how they can efficiently use it in their workflows, that same asset can generate 3,000, and that's just the metadata.

Yeah. Every stage of the workflow is producing more data, more iterations, more outputs.

Another tangible example would be of AI upscaling. AI has made the ability to make an older show, remaster it, and make it look great quicker-ish. If you have a show that was shot natively in 2K and a decision is made, let's make it 4K for OTT or streaming, or even 8K. Let's say there's a premium tier or you're trying to future-proof it. That 2K original isn't going to disappear. They're not going to trash it. It's the source of truth. It's the legal master. What we're seeing obviously is now you have that one asset with three retained versions, the 2K, 4K, 8K. The 2K, let's say it's a one-hour show, and it's ProRes. Let's go with that. That might run you like 100 gigs.

That 4K upscale typically is going to need its own render passes, going to need its own mezzanines, it's going to need its own QC deliverables. That one new version is realistically going to bring 3 to 4x to your footprint per title.

Yeah. That's before 8K. All of that's happening, and that's great.

The best. I mean, it's great that you can be able to do that, and faster.

However, if your infrastructure wasn't great before, you're going to start to see the cracks. I think Drive Stats is becoming even more important to this audience because they're having to bolt together really smart systems and infrastructures.

Yeah. The story is not whether we do cloud or on-prem anymore.

That's dead. It's more, how do we put all of this together to work efficiently, to be performant for capacity's sake? Because obviously budget is a factor, how much can we really put on this drive and it not make all of our originals disappear? Those are always going to be very front-of-mind thoughts. This is why Drive Stats is really important in my domain.

I love it. Yeah, it's always good to hear, and it's interesting you mention AI because we've got a question from Raghav here in the chat. How is AI impacting Backblaze? How do we think it'll affect small and medium businesses when it comes to storage based on what we're seeing? Just based on what you are saying here, and sort of the conversation we had before the call, I totally agree, right? People in M&E space have to consider individual drives, but also all these other moving parts of their architecture. As far as AI goes, I think it's a very interesting conversation. I can talk about it from data center perspective, but I'm interested to hear what you have to say in the M&E industry or even in just content generation at all. Thinking of content broadly as multimedia files, I should say.

Yeah, it's just that continuous derivatives that are being created that people are thinking, "Oh, I can be more creative. Our team can deliver more.

Again, if you don't have that IT team, if you don't have that technical person that is thinking about, that's awesome. I love everyone being creative and wanting to grow and develop, but how do we store all of that efficiently? How do we get to it when we need it? How do we monetize it, and be able to find it quickly? All those things start to come into play when you are doing all this iteration. Another example would kind of be different grading. I don't know if people think about when I say grading, I mean color grading. Sometimes, and we've all seen it, the warm grade for a certain type of movie versus a blue color grade. There's a lot of testing. Directors, DPs, the director of photography, they do a lot of testing. Now we're maybe doing AI color grading upscales.

Now we have the upscale and we have the color grade.

Maybe there's four different ones. Let's add on HDR to SDR grades- Yeah A/B test audience variants.

It's amazing. AI is great because you have the ability to create these things possibly faster. Again, if you don't have the infrastructure built out and the processes to manage it and figure out where everything is and make sure everyone's getting all of the files that they need to create, that's when it can become an issue.

It's just like an explosion of data on the most basic level.

Managing data, particularly in active archives, has always been a conversation in this space. I think when you're talking about how it's affecting, if you're saying how is AI impacting Backblaze, I think that's a pretty complicated question because it kind of depends on what lens you view it through. We certainly see this huge demand for data. On the flip side, we also see access patterns for these things moving in really different ways. If you look at the network stat series that we work on too, you'll see just a high volume of data moving all at once, and then being processed a lot of times through, sometimes even through traditional CDN providers who are converting to or are already neo clouds.

That, in conversation with traditional media tooling and where a lot of that data lives, you're seeing a lot of these integrations become super important to your point, Lakey, where it's like you have a lot of moving parts, and they all have to work together.

Absolutely. That's why I love the ecosystem that we've built, Backblaze. We have amazing partners that are doing just mind-blowing things with AI and technology. I think that's another way that Backblaze is being impacted and also being part of the story, is that the openness of Backblaze. Yeah, it's an amazing time, but it's definitely one of those fire hose moments of how do we manage all this? I think the infrastructure was built so well here that I think that we're just excited about what we're seeing and able to ride the ride with people and our partners.

Yeah. Agreed. Next question, two kind of related ones. About how to use the stats to buy a drive or what sort of brand we'd recommend. I'll give our standard language here. We really do not play favorites when it comes to recommending drives because, and I think this is an important part of the conversation, in some ways, what we do is sort of the ultimate litmus test for a hard drive. They are running at max capacity for their entire life until they die. That is what we do in a data center. On the other hand, your personal use case might look quite different. We also have redundancy both on the software layer and with secondary drives to be able to manage potential data loss.

The way that I would recommend using the stats is to identify what size of drive you need first, and then see with what you're comfortable with, look at reviews, and then build out for backup. As always, I always recommend having at least, at bare minimum, a 3-2-1 backup strategy. With parts of it in the cloud, parts of it offsite. If you want to have your onsite for a traditional air gap, that makes sense, too. All of that gets complicated. Data management on a personal level is something that.

Yeah I have a lot of respect for.

Yeah. really, you can get as durable as you want to on a personal level and your drives are certainly part of that conversation.

I think these drives really help, trying to help people figure out what reliability at scale looks like for them.

When they're feeding data into, whether it's automated pipeline or, we're not just talking about the human editor aspect of it. If people are doing a lot of modeling and training against their footage and their archives. undetected drive failures, mid-data set. It's not just. Yeah you lost a file.

It's like silently corrupting and the whole training model, right? yeah. This is, I think, a great report for people to look at and just try to assess accordingly what's going to work for them in their current workflows.

Yep. Totally agree. All right. Looks like we've got another question from Hank Guns. There we go, guys. Excuse me for mispronouncing your name there, Hank. Member of a group of 200 photographers. Awesome. Love it. Cool. You are the techie, and you showed them how to back up on-site, but you want to have a presentation on why they should use Backblaze to back up off-site.

Hank, with no personal motivation, I direct you to our blog. I say that because I've written on the blog historically for quite a long time these days. We've got quite a few articles on the benefits of off-site backup. I think when you're talking about off-site, it's really important when we think about disaster recovery, and that sort of thing. The other piece that I think is important specifically for this audience is a little bit what Ryan is talking about, where you have lots of things everywhere and you might want to be able to access them remotely or with different tools or in different ways.

You might want to, I think just simultaneous and distributed access is an important part of this conversation, too. Beyond backup, those active archives that folks need to work with. Right. Anything else you're thinking about, Ryan? Anything top of mind for you?

No, I don't think so.

Yeah. I think there's a lot to look at here, and many different lenses to look at the impact of hard drives and AI. I think what's particularly interesting to me is people have started to look at their tech stack and understand how much they need it to move with them. Right? Those active archives are always part of the conversation in my neck of the woods.

Yeah. No one wants those single points of failure. We're always trying to figure out how to make a workflow smarter.

Yep. In your archives, there's literal gold.

Each frame has such high value, it's just really important to make smart decisions at this point as you're growing as you iterate.

Yep. I totally agree. All right. Well, if we don't have anything else from the comments section, I feel like we've had a great time here today. Appreciate everybody for showing up. As always, feel free to reach out if you've got additional questions. We're happy to answer them.

Full transcript, live translation, and audio in the StockNow app.

Get Started