MediaFM: The Multimodal AI Foundation for Media Understanding at Netflix
Netflix's MediaFM is a multimodal AI model that enhances media understanding by integrating audio, video, and text to improve content recommendations, ad relevancy, and clip analysis, leveraging a self-supervised learning approach on its vast catalog.
MAIN POINTS
- MediaFM is a tri-modal model using audio, video, and text for content embedding.
- It employs a Transformer-based encoder to generate contextual embeddings.
- The model improves tasks like ad relevancy and clip popularity ranking.
- Evaluation shows MediaFM outperforms existing models in narrative understanding tasks.
TAKEAWAYS
- MediaFM enhances Netflix's ability to understand and recommend content by fusing multiple media modalities.
- The model's architecture allows for robust content analysis, aiding in promotional asset optimization.
- Contextualization of shot representations significantly boosts model performance.
- Future work includes leveraging pretrained multimodal LLMs for further model advancements.