Titre inconnu
YouTube transcript, YouTube translate
A quick preview of the first subtitles so you know what the video covers.
Nvidia just put 748 GB of RAM into a desktop, and that one number changes what local AI even means. A 70 billion parameter model, full precision, no cloud, runs on this thing the way a browser tab runs on your laptop. The hardware ceiling just moved way further than most builders realize. So, this is the DGX Station, right? Announced by Jensen Huang at Nvidia's Computex 2026 keynote. That was in Taipei, end of May. And it's, you know, it's a tower you stick under a desk. And inside it, there's a chip called the GB200 Grace Blackwell Ultra. One package, two processors kind of fused together. A 72-core ARM CPU and a Blackwell Ultra GPU. And they're joined by Nvidia's own NVLink interconnect. Okay, quick reality check on the headline though. You'll see 768 gigs floating around, and that's even on one of MSI's own listings. But the real spec is 748. Nvidia's official number, every primary source, they all say 748. So, that's the figure I'm going to use, and honestly, you should be a little suspicious of anyone quoting the round one. So, here's the spec that actually matters, and I'll tell you why in a second. That 748 GB, it's unified coherent memory. So, roughly a quarter terabyte of fast GPU memory. The rest is high-bandwidth system RAM, and the whole thing just sits in one pool. Either processor can read all of it, no copying back and forth. That's the part that does the work, and that's where I'm headed next. So, the spec everyone fixates on is compute, right? But the spec that actually unlocks local AI, it's the memory. Let me make that concrete. If you've ever tried running a big open weight model on your own machine, the wall you hit, it's not speed, it's RAM. The model's got to fit in memory to run at all. And a 70 billion parameter model, full precision, needs around 140 GB just for the weights. And that's before you add the KV cache, which is, you know, the working memory the model uses to track a conversation as it's generating. On a normal gaming GPU with 24 gigs, you can't load it. So, you quantize it down, you shrink it, you accept worse output, or you just give up and rent a cloud GPU.