Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Correction, I took the 4 TB/s from TFA, but the NextPlatform article clarifies that it is 4 TB/s per chiplet, but 8 TB/s per socket, so more impressive.

The only US-designed CPU "exotics" are the Intel Xeon Max CPU series, which use HBM like the Chinese CPUs, but which have a theoretical maximum throughput of only 1.6 TB/s per socket, i.e. 5 times slower than the new Chinese CPUs.

Moreover, the users of Intel Xeon Max complained that they cannot reach the theoretical memory bandwidth. I do not know if that was due to some bug that might have been solved later by Intel with a microcode update or a new mask set stepping.

The server CPUs with standard DIMMs, which will be launched by AMD and Intel next year, will have a memory bandwidth of around 1 TB/s per socket.

The AMD MI300 GPU used in the fastest US supercomputer has a memory throughput of 5.2 TB/s per socket, so lower than the 8 TB/s per socket of the Chinese CPU, which explains why the advantage of the Chinese system increases in the benchmarks more dependent on memory performance.

The latest AMD Instinct GPU, MI355X, increases the memory bandwidth to 8 TB/s, so equal to the Chinese CPU.

However, it may pass some time until someone will build such a big system with MI355X, though perhaps the existence of this new contender might prompt the US labs to upgrade their systems by replacing the older AMD GPUs with newer AMD GPUs.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: