Mirgahney Mohamed on Modelling How People Dance | FAI CDT

Mirgahney Mohamed on Modelling How People Dance | FAI CDT

🎙 UCL Centre for Artificial Intelligence 👥 3K 📅 September 5, 2025 ⏱ 34 min 👁 123 📄 interview 🧭 2026-08-15
Available in: English (current) Français

Keywords

diffusion modelshuman motiontext-to-motionGaussian splatting4D reconstruction

Summary

In this interview, Mirgahney Mohamed, a final-year PhD student at UCL’s Foundational AI CDT, discusses her research on modeling human motion in 3D and 4D. She explains three aspects: modeling, reconstructing, and generating motion. The conversation covers her internship at Google DeepMind, where she worked on 4D scene reconstruction using Gaussian splatting, a technique that represents a 3D scene as a collection of colored blobs (Gaussians) and can render novel views and dynamic scenes. She also mentions a startup she co-founded that used NFTs to verify celebrity signatures on digital goods. The main focus is her paper on recurrent diffusion models for human motion generation, which addresses the task of generating human motion from text descriptions. She contrasts three approaches: volume diffusion (generating the entire sequence at once), autoregressive diffusion (generating chunks sequentially), and her proposed recurrent diffusion, which conditions on intermediate noisy outputs to speed up generation. The recurrent method achieves a 5x speedup over volume diffusion and 10x over autoregressive diffusion, while also improving quality. The interview highlights potential applications in entertainment, virtual reality, and robotic surgery.

179 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides valuable insights into the state-of-the-art in human motion generation, particularly the use of diffusion models. The explanation of Gaussian splatting and its extension to 4D is clear and accessible. The argumentation for the recurrent diffusion model is well-structured, with a logical progression from volume diffusion to autoregressive to recurrent, and the claimed speedups are impressive. However, the discussion lacks quantitative details and comparisons with other methods, and the interviewer’s questions, while helpful, sometimes lead the discussion away from technical depth. The value lies in the conceptual overview and the researcher’s perspective on the field.

Scientific Rigor, Source Quality, Title Accuracy

The video is an informal interview, so it does not cite formal sources. However, the speaker references her own research and her internship at Google DeepMind, which adds credibility. The title accurately reflects the content. The description mentions potential applications in robot surgery and computer animation, which are discussed in the video. No external sources are provided in the description, so the evaluation relies on the speaker’s expertise. The content is consistent with known research directions in the field, but without formal citations, the scientific rigor is moderate.

200 words

Title / Content Match

The title accurately reflects the content, which focuses on Mirgahney Mohamed's research on modeling human motion, including dance.

Quality & Reliability

7/10

The video is an interview with a PhD student discussing her research on human motion generation using diffusion models. The content is technically accurate and well-explained, but it is a high-level overview without deep technical details or peer-reviewed sources. The speaker is a researcher in the field, which adds credibility, but the lack of formal citations and the informal setting limit the score.

Key Moments

Contribution & Novelties

The video presents the researcher’s novel approach to human motion generation using recurrent diffusion models, which condition on intermediate noisy outputs to speed up generation while improving quality. This is an original contribution to the field. The discussion also covers Gaussian splatting for 4D reconstruction, which is a recent technique. The video provides a clear explanation of these concepts, making them accessible to a broader audience.

Pour aller plus loin :

109 words

Radar Profile

The radar profile shows balanced scores across all dimensions, with slightly higher scores in quantity and quality of information, reflecting the informative nature of the interview. The technical level is moderate, suitable for a general audience with some background in AI.

Reliability 7/10