Despite remarkable progress in image and video generation, translating user intent into precise and consistent visual outputs remains a challenge. This talk explores how understanding the representations within generative models can enable finer control over what they create.
It connects semantic image editing with compositional generation, examining how visual concepts can be isolated, manipulated, and combined while preserving their identity and surrounding content. Building on these insights, structured visual inputs provide a way to express complex intent through subject references, poses, and spatial layouts.
The discussion then extends from images to video, where representations must evolve to preserve scene continuity while accommodating motion and change. Together, these directions establish a unified perspective on how visual representations can support controllable editing, composition, and coherent generation across space and time.