Collected sources and patterns will appear here. Add from search or the patterns library.
MultimodalRepresentations -> Trajectory<Action>
Predict a chunk of continuous robotic actions by conditioning a Diffusion Transformer (DiT) on multimodal representations extracted from a Vision-Language Model backbone.
Problem it solves
Autoregressive VLMs struggle with modeling multi-modal, continuous action spaces directly.
Consumes
Emits
The real projects this mechanism was found in. Attribution is the point — this is how the best teams actually do it.