L4 · Training a Language Model
Splitting the Work
Tell apart, by picture, the three ways a run is cut across many chips.
Training a 70 billion parameter model takes about a terabyte of memory. One H100 chip holds 80 gigabytes. So the job is cut into pieces, and the pieces spend their lives talking to each other.