Aether: Decoding, CPU-Resident Compute & GPU Scheduling, Running a 120B MoE on a Single Consumer…

Estimated read time 1 min read

In Part 1, we introduced Aether and the core ideas behind running a 120B MoE model on a single consumer GPU.

 

​ In Part 1, we introduced Aether and the core ideas behind running a 120B MoE model on a single consumer GPU.Continue reading on Medium »   Read More AI on Medium 

#AI

You May Also Like

More From Author