Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

1.2 TB/s bandwidth of M5 Ultra comes from two dies of M5 Max (each 614 GB/s) connected together using 4.4 TB/s inter-die fabric.

For a non-quantized Deepseek V4 flash on an ultra, I would estimate about 1000+ tokens per second prefill and 50+ tokens per second on generation. This is actually quite usable and near parity to cloud.

They mention "adds the GPU Neural Accelerators." which, if exploitable for LLM loads, would probably help the prefill a lot

 help



Yes, and they specifically mention "Up to 10.7x faster LLM prompt processing in LM Studio" which is probably using the neural accelerator for prefill.

How much would is the cost for that machine though, I'm pretty sure I could just buy tokens from a provider and never run out of money for 10 years, and get far better quality output because inference is being served by professionals on far better hardware and this machine would be obsolete long before that as well. Hosting local seems like a possibly the dumbest thing you could possibly do from an economics perspective. And don't hit me with the privacy argument because everyone saying they care about privacy uses fucking gmail, whatsapp and instagram all day long.

One way to gauge how close to viable local inference is the quality of the arguments against it are.

$100/mo for ten years is $12K. $200/mo is $24K. When you look at the valuations of the companies providing these tokens, you don't necessarily see these rates going down.

Your argument against privacy is so obviously bad I don't think it needs a response.

Today, most companies and individuals (including me) are still going to be buying tokens from the cloud. But for many it makes sense to go local today, and it seems pretty clear the day where we turn the corner isn't that far away. E.g., suppose the RAMpocalyse eases in 3 years and Apple releases the M7 ultra, which has actually been designed for local inference (and other players are making their moves as well). I suspect it will become something of a no-brainer to do most inference locally, perhaps supplemented by lighter usage of per-token plans from the cloud for access to specialized models.


Why is the privacy argument invalid?

There is definitely classes of data that you legally can't share over "fucking gmail, whatsapp and instagram" and you can't send to commodity token providers either.

Taking liberties to save money at the expense of regulatory duties is a good way to get hit with a massive fine if you work in financial or health services.

This is a product aimed at niche professionals doing paid work. If you don't need that just move on to something else.


The problem is when you don't want to send your data to any provider. Private documents, health information, all that is something you don't want to send anywhere, even if the provider makes a pinky promise that they won't use your data for training or store it.

And no, I do not use gmail, whatsapp or instagram :-)


You can’t run multiple workers 24/7 for 100 bucks a month

you might want to do that, but you underestimate how much some hobbyists/makers/hackers value being able to do things on their own tools in their own house.

for me, it’s significantly more about the process than it is about penny pinching. it’s why im not an accountant and why i often despise bean counters.

my dad has soooo much money in tools alone for working on cars, faaar more than he’d ever spend at a mechanics. add in the costs he spends buying hobby cars and real estate/upkeep for the garage to work on his cars and it adds up quick. people spend a lot on their hobbies.

look at how much woodworkers spend on tools. look at how much homelabbers spend. photographers, boat/sailing wonks, horse owners, and on and on and on.

add in the above, then the value of the privacy aspects, idealism aspects, potential career knowledge, etc… and it’s obvious that local models will do well.


Ok sure. But this machine you can haul in an Ice Road Truck to the North Pole and do inference there in an off grid shack. Good luck talking to cloud AI there.

Wow even the elves will get replaced by AI

Yeah, that's who Apple is making these for fucking Santa and his Elves lmfao



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: