Meta has improved the training efficiency of its Generative Ads Recommendation Model (GEM) by doubling end-to-end training efficiency to 20–25% Model FLOPs Utilization (MFU) while scaling training FLOPs 4x over 12 months.
From the source
Meta’s Generative Ads Recommendation Model (GEM), the foundation model behind ads recommendations across Instagram and Facebook, now trains at LLM scale on several thousand of the latest-generation GPUs. This post goes into the details on how we achieved: doubling end-to-end (E2E) training efficiency to 20–25% Model FLOPs Utilization (MFU) while scaling training FLOPs 4x in 12 months, by co-designing kernels, precision, parallelism, networking, and memory together.
engineering.fb.com