GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model
Meta's Generative Ads Recommendation Model (GEM) now trains at LLM scale, achieving doubled training efficiency and quadrupled FLOPs using thousands of advanced GPUs.
MAIN POINTS
- GEM is the foundation model for ads recommendations on Instagram and Facebook.
- The model now operates at LLM scale with thousands of latest-generation GPUs.
- Training efficiency has doubled to 20–25% Model FLOPs Utilization (MFU).
- Training FLOPs have been scaled up by four times.
TAKEAWAYS
- Meta has significantly enhanced GEM's training efficiency and scale.
- The advancements in GEM are crucial for improving ads recommendations.
- Utilizing thousands of GPUs is key to GEM's improved performance.
- The post provides detailed insights into the technical achievements behind GEM's scaling.