LoadingЗагрузка

Xiaomi-Robotics-U0 is an autoregressive world foundation model that... - Vibeus

Nora Bennett ·

Xiaomi-Robotics-U0 is an autoregressive world foundation model that combines image generation with robot-centered spatial and interaction modeling in one multimodal token space. Built on EMU3.5 with an IBQ image tokenizer, it uses one next-token objective across text, images, visual sequences, and robot observations. Its six listed tasks cover multi-view scene generation, editing across five axes (workspace, background, target objects, irrelevant foreground objects, and lighting), video rollout from an observation plus action context, interleaved subtask text and predictions, text-to-image synthesis, and image editing. The announcement reports first place on WorldArena among 100+ models for video rollouts and higher multi-view embodied scene-transfer win rates than GPT-Image-2. FlashAR+ combines anti-diagonal visual-token decoding with custom vLLM patches and cuts the reported 1024×1024 generation latency from 450 seconds to 5.44 seconds per sample — about 83×. Источник Источник