L10 · Smaller, Faster, Cheaper
Cheap to Run, Costly to Hold
Say what a mixture of experts saves and what it does not, on a busy server and on a phone.
Mixtral, a mixture of experts, holds 46.7 billion weights and runs 12.9 billion for each token, as chapter L3 showed. So it runs like a 13 billion weight model. It still has to be held like a 47 billion one, every byte, all the time.