Rendered at 20:56:07 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
librasteve 1 days ago [-]
As the video points out, there is a hardware/software feedback loop at play here. Since the hardware is deeply pipelined SIMD FPU datapath, the software is hand tuned machine coded transformations. In addition to limiting flexibility (every variation has to be ground out in CUDA), this prevents sparse matrix type optimisations. I predict that a set of general purpose CPUs - non shared memory at this scale - would be a much better use of transistor/power. And you can code that at high level give a CSP style approach such as https://bil-lang.org
feffe 23 hours ago [-]
I think tenstorrent architecture is more like this. A grid of RISC-V cores with local SRAM and vector units.
tolugenius 1 days ago [-]
I wonder what would "compute" mean if cpus were more efficient at matrix multiplication say 15 years ago. And on the flipside, what it would take to say train a frontier model entirely on cpus in the future.
rhdunn 1 days ago [-]
On the compute side there's the issue of scalability. CPUs are designed to perform a handful of operations at one time. They typically have a small number of dedicated integer, float, and other ALU configurations. Having dedicated matrix multiplication instructions would still lock that to how many matrix-capable ALUs there are in the CPU.
GPUs are designed to process a large number of calculations at once (so they can process triangles in 3D graphics). This makes them good at ML applications as they can process many of the matrix calculations at once. A 4090 has 16,384 CUDA cores (general compute ALUs) and 512 Tensor cores (dedicated matrix compute ALUs); a 5090 has 21,760 CUDA and 680 Tensor cores.
The other issue when training models (and running larger models) is the amount of VRAM (or RAM for CPUs) available. GPUs are limited in this aspect, whereas CPUs can have a lot higher memory. This affects things like batch size and the size of model that can be trained or fine-tuned.
https://unsloth.ai has guides for how to fine-tune existing models like Qwen 3.8 27B, memory requirements, etc.
https://medium.com/@kailaspsudheer/the-transformers-arithmet... has some information on training a base model. A 7B llama model is estimated at taking ~34GB memory for inference at F32, but was observed requiring 96GB memory when training (for the model weights, gradients, activations, and optimizer states).
Note: you can reduce the memory required for training by recomputing the gradients, at a cost of performance/time. You can also do other tricks like performing a QLoRA/LoRA pass on the model then merging that into the model to create a checkpoint.
I don't know what sized model you could train on 64GB/128GB RAM via a CPU.
pjmlp 23 hours ago [-]
That was the whole point of Project Larrabee, whose ideas survive in AVX.
The ideas survive in AVX-512 a.k.a. AVX10, not in AVX, which was a project parallel to Larrabee and resulting in an inferior ISA, which was adopted in the mainline Intel CPUs due to internal politics, not due to technical superiority.
When brought to the Intel server CPUs, the Larrabee New Instructions were rebranded as "AVX-512", despite having no relationship with the AVX ISA extension.
AVX was the creation of the Intel A-team, while the Larrabee New Instructions were designed by a C-level or D-level Intel team, but the latter have benefited from the contribution of a few consultants hired from outside Intel, who had experience in programming graphic applications.
AVX, which included only minimal and obvious improvements over SSE, i.e. double width and 3-address instructions, has slowed down considerably the improvement of the computational performance of CPUs in comparison with an alternate time line where Intel Sandy Bridge would have implemented a variant of the Larrabee New Instructions instead of AVX. This could have been done in a manner that would not have required any significant cost increase over the Sandy Bridge with AVX, because in AVX-512 it is not the width that is important but the architecture of the vector instruction set (e.g. with masked operations).
pjmlp 9 hours ago [-]
Thanks for the correction, I assumed that is what people would understand by only saying AVX, my bad.
21 hours ago [-]
Dwedit 22 hours ago [-]
16 years ago, Intel CPUs finally got GPUs integrated inside of them. So they did get better at matrix multiplication 15 years ago.
alfiedotwtf 20 hours ago [-]
The tv ads for MMX made it feel like it was going to change the world
Dwedit 19 hours ago [-]
15 years ago was Sandy Bridge, which added the Intel HD Graphics GPU inside of the processor. MMX was almost 30 years ago.
actionfromafar 1 days ago [-]
I wonder what "compute" would mean if CPUs were more efficient at matrix multiplication and vendors had the balls to pair each core to its own dedicated DDR and a star interconnect between.
pjmlp 23 hours ago [-]
It could be like the connection machine or something like that.
There is plenty of matrix multiplication in SIMD, but it isn't widely explored.
aeve890 17 hours ago [-]
>There is plenty of matrix multiplication in SIMD, but it isn't widely explored.
Like in the existing SIMD technology?
pjmlp 16 hours ago [-]
Yes, like in the existing SIMD technology, ask your AI friend for code samples.
ekjhgkejhgk 7 hours ago [-]
"I think we know how to solve power, dyson spheres"
closed
vivzkestrel 18 hours ago [-]
- stupid question
- how exactly does apple silicon s unified gpu + ram thingy work?
- how come intel and nvidia cannot do the same?
- is there an actual difference in terms of hardware architecture or something or is it pure apple marketing hype?
aeve890 17 hours ago [-]
Yes. there is real hardware architecture behind Apple’s “Unified Memory”, but the underlying idea is not uniquely Apple. What Apple did was design the CPU, GPU, memory controller, cache hierarchy, interconnect, package, OS, and graphics APIs together around the architecture. I mean, they can do shit like that because they own the entire product pipeline. They can fine tune the hardware in ways other OEMs can't.
In a Intel+Nvidia CPU+GPU pipeline you have the CPU, that load something in ram, then for the GPU to process it you need to move it from RAM to VRAM through a slow PCI connect then when the GPU is done you have to move it again to RAM for the CPU to handle it again. These copies cost bandwidth, latency, energy, extra memory and low level programing complexity. In the Apple unified architecture both de CPU and GPU use the same memory and avoid copying data.
You may say "ah but don't you can handle shared memory with just the DDR controller?" and that's what Intel integrated graphics do but on top of that Apple: build the CPU and GPU in the same SOC, gives the SOC a massive memory subsystem and a large system level cache.
Nvidia can do the same. Eg the Grace Hopper has the Grace CPU - nvlink C2C - Hopper/Blackwell GPU. But the trade-off is modularity. You can mix processors, ram, Nvidia GPU and all parts must work at their best capacity, but isn't even close to the fine tuning of an Apple system.
ranguna 12 hours ago [-]
TLDR: Apple owns the whole stack. Whilst other build individual components that work together through slow interconnects.
Why not make a faster interconnect with an open communication standard?
GPUs are designed to process a large number of calculations at once (so they can process triangles in 3D graphics). This makes them good at ML applications as they can process many of the matrix calculations at once. A 4090 has 16,384 CUDA cores (general compute ALUs) and 512 Tensor cores (dedicated matrix compute ALUs); a 5090 has 21,760 CUDA and 680 Tensor cores.
The other issue when training models (and running larger models) is the amount of VRAM (or RAM for CPUs) available. GPUs are limited in this aspect, whereas CPUs can have a lot higher memory. This affects things like batch size and the size of model that can be trained or fine-tuned.
https://unsloth.ai has guides for how to fine-tune existing models like Qwen 3.8 27B, memory requirements, etc.
https://medium.com/@kailaspsudheer/the-transformers-arithmet... has some information on training a base model. A 7B llama model is estimated at taking ~34GB memory for inference at F32, but was observed requiring 96GB memory when training (for the model weights, gradients, activations, and optimizer states).
Note: you can reduce the memory required for training by recomputing the gradients, at a cost of performance/time. You can also do other tricks like performing a QLoRA/LoRA pass on the model then merging that into the model to create a checkpoint.
I don't know what sized model you could train on 64GB/128GB RAM via a CPU.
https://pages.cs.wisc.edu/~markhill/restricted/siggraph08_la...
When brought to the Intel server CPUs, the Larrabee New Instructions were rebranded as "AVX-512", despite having no relationship with the AVX ISA extension.
AVX was the creation of the Intel A-team, while the Larrabee New Instructions were designed by a C-level or D-level Intel team, but the latter have benefited from the contribution of a few consultants hired from outside Intel, who had experience in programming graphic applications.
AVX, which included only minimal and obvious improvements over SSE, i.e. double width and 3-address instructions, has slowed down considerably the improvement of the computational performance of CPUs in comparison with an alternate time line where Intel Sandy Bridge would have implemented a variant of the Larrabee New Instructions instead of AVX. This could have been done in a manner that would not have required any significant cost increase over the Sandy Bridge with AVX, because in AVX-512 it is not the width that is important but the architecture of the vector instruction set (e.g. with masked operations).
There is plenty of matrix multiplication in SIMD, but it isn't widely explored.
Like in the existing SIMD technology?
closed
- how exactly does apple silicon s unified gpu + ram thingy work?
- how come intel and nvidia cannot do the same?
- is there an actual difference in terms of hardware architecture or something or is it pure apple marketing hype?
In a Intel+Nvidia CPU+GPU pipeline you have the CPU, that load something in ram, then for the GPU to process it you need to move it from RAM to VRAM through a slow PCI connect then when the GPU is done you have to move it again to RAM for the CPU to handle it again. These copies cost bandwidth, latency, energy, extra memory and low level programing complexity. In the Apple unified architecture both de CPU and GPU use the same memory and avoid copying data. You may say "ah but don't you can handle shared memory with just the DDR controller?" and that's what Intel integrated graphics do but on top of that Apple: build the CPU and GPU in the same SOC, gives the SOC a massive memory subsystem and a large system level cache.
Nvidia can do the same. Eg the Grace Hopper has the Grace CPU - nvlink C2C - Hopper/Blackwell GPU. But the trade-off is modularity. You can mix processors, ram, Nvidia GPU and all parts must work at their best capacity, but isn't even close to the fine tuning of an Apple system.
Why not make a faster interconnect with an open communication standard?