AMD has shown at IFA 2026 a tower you put under your desk with data-centre accelerators inside. It's called Threadripper Halo, it ships next year, and it's the first time AMD's Instinct parts have been sold in a workstation form factor. The number everyone will quote is 576 GB of HBM3e memory. The number that actually matters is a different one, and I'll get to it below.
What happened
The system is liquid-cooled and built around a 96-core Threadripper PRO 9995WX — a CPU that already launched in 2025, so there's no new silicon here — with up to 2 TB of DDR5 and up to 2.6 TB of combined system memory.
The interesting part is the cards: up to four Instinct MI350P units hanging off PCIe. The MI350P arrived in May and is essentially half an MI350X squeezed into a standard card format. Each one carries 144 GB of HBM3e and shifts four terabytes per second; AMD claims up to 4.6 petaFLOPS working at four-bit precision, though that figure belongs to the 600-watt configuration. In a tower the sensible bet is that they run at 450, so any performance maths should be done lower.
Add the four together: 576 GB of HBM3e and 16 TB/s of aggregate bandwidth. That's enough to keep a model of more than a trillion parameters resident in GPU memory, quantised to four bits. Leaning on system memory as well, even the biggest open-weights model around fits — something the size of Moonshot.AI's Kimi K3, at roughly 2.8 trillion parameters.
AMD hasn't given a price. The Register estimates somewhere between one hundred thousand and one hundred and fifty thousand dollars depending on configuration, and compares it with Nvidia's DGX Station — announced at GTC 2025, with a 252 GB B300 GPU, a 72-core Grace CPU and 496 GB of LPDDR5x — which sells for around one hundred thousand dollars when you can find one. AMD boasts 3.4 times the total system memory and more than double the bandwidth of the Nvidia box.
And one telling detail: the units on the show floor had two cards fitted, not four.
Why it matters
The key number isn't petaFLOPS, it's bandwidth. When you run inference on a large model, producing each token means walking the active weights through memory. If you have compute to spare and bandwidth that's tight, the GPU sits waiting. That's why 16 TB/s weighs more in the decision than any FLOPS figure, and why the comparison with the DGX Station is fought over memory rather than raw power. If you're evaluating hardware for local inference, rank your options by gigabytes of fast memory and bandwidth per pound. The rest is noise.
The real niche is small, and it's worth saying so. This machine makes sense if you need the biggest open-weights models, on premises, with data that can't leave your network, and with enough sustained load to amortise six figures. That's research, healthcare, defence, banking and not much else. If what you want is a 20B or 70B model to classify tickets, pull data out of invoices or help your dev team, this is a sledgehammer for a drawing pin: there are far cheaper boxes that hold that range comfortably.
What does change for everyone is the mental frame. Until now, "this can only run in the cloud" was a reasonable answer for any frontier-class model. From next year, for open-weights models, it stops being one entirely: it becomes a budget decision, not a physics one. That hands you arguments in two awkward conversations. With the client who doesn't want their data touching a third-party API, there's now a technical alternative you can actually size. And with your cloud provider, because having a credible option on the table improves the negotiation quite a lot.
Before you order anything, look at the socket and the air. Four cards at 600 watts will overrun a standard North American domestic circuit unless AMD drops them to 300 or mandates a 20-amp line; that's probably why the show units had two rather than four. Over here, with 230 volts, the electrical side is less of a problem, but the heat isn't negotiable: a liquid-cooled tower dissipating that much doesn't belong in a small office with a domestic split unit. And the kit won't be available in every market.
The cost nobody budgets for is software. This is my read, not the article's: AMD's hardware has been competitive on paper for years, and the problem is still how much of your people's time goes into making the stack work. If your team has spent years on CUDA, count the hours of porting, debugging kernels and chasing versions. That number usually decides the purchase more than the price of the box.
What doesn't change
There's no official price. The six-figure range is the outlet's estimate, not AMD's, and it grates on me to see it repeated as though it were a tariff. There's no date either: "next year" is a generous window, and until there are real units with four cards inside, the top configuration is a brochure promise.
The 3.4 times figure on total memory is AMD talking about itself. And there's an honest technical counterweight: when a model is split across four cards linked by PCIe, the bus and the CPU can become the bottleneck and eat a good chunk of that theoretical bandwidth advantage. We'll need independent measurements.
Nor is this new hardware: it's a well-judged recombination of parts that already existed. That's not a flaw — it's the fastest route to a new form factor — but it explains why there's no architectural surprise here.
And the usual point: owning capacity only pays off with sustained load. A few weeks ago I wrote about Anthropic's cache price cut; while cloud cost per token keeps falling at that pace, the break-even point for a six-figure machine drifts further away every quarter it sits idle.
If you're thinking seriously about local inference and want to run the numbers together — memory needed, power draw, and which parts are better left in the cloud — tell me and we'll look at it.
Best regards, Vicente.
