Mustafa Shukor
Founding Research Scientist ( UMA). CS PhD (
Sorbonne).
Prev: FAIR ( Meta), LeRobot (
Hugging Face), MLR (
Apple).
Updates
-
2026-03: I defended my PhD at Sorbonne: Efficient and scalable multimodal learning.
-
2026-02: LeRobot is accepted at ICLR 2026.
-
2026-01: We released Action100M, a large-scale video action dataset with ~100M temporally localized segments.
-
2025-12: Introducing VL-JEPA, a vision-language model trained in the latent space.
-
2025-12: We are launching UMA, a company that is building general-purpose humanoids.
-
2025-11: VLAb is out! Our codebase for VLAs pretraining, including SmolVLA.
-
2025-09: l3m is released. Our codebase for large-scale pretraining of AIMv2, CLIP, Native VLMs and LLMs.
-
2025-09: Scaling laws for optimal data mixtures is accepted at NeurIPS 2025.
-
2025-09: Learning to Steer is accepted at NeurIPS 2025.
-
2025-07: Scaling laws for native multimodal models is accepted as an Oral (~top 2%) at ICCV 2025.
-
2025-07: Multimodal Steering is accepted at ICCV 2025.
-
2025-06: We released SmolVLA, an efficient foundation model for robotics: code and blogpost.
-
2025-02: AIMv2 is accepted as a Spotlight paper (~top 5%) at CVPR 2025. Released code.
-
2024-09: Implicit Multimodal Alignment, CoX-LMM, DiffCut and Skipping Computations are accepted at NeurIPS 2024.
-
2024-04: Beyond task performance is accepted at ICLR 2024.
-
2023-12: UnIVAL is accepted at TMLR 2023: code.
-
2023-09: Rewarded soups is accepted at NeurIPS 2023.