Reka has released a research preview of Rho-1, a 19B omni-reasoning model trained from scratch. A single neural network understands and generates text, images and video, reasons over them, and outputs robot actions. Reka frames it as a direct replacement for agentic pipelines that pass work between modality-specific models.
What Rho-1 Changes
Most multimodal systems today are pipelines. A central model plans, then hands jobs to specialists for images, video or detection. Each handoff adds latency, and each specialist sees only a narrow request.
Rho-1 removes those handoffs. Text, vision and robotic actions become tokens inside one context window. According to Reka’s research, one unedited session shows the full loop. The model draws a lighthouse, boxes it, animates it, edits the video into a snowstorm, and explains the difference. That all happens in 5 turns, with no tool call and no second model.
Architecture: Two Streams, One KV Cache
Every input and output uses one of two native formats:
- Discrete tokens carry text, symbolic reasoning and high-level commands.
- Continuous tokens carry image latents, video frames, robot actions and proprioception.
Each transformer block holds two expert weight streams. The understanding stream handles language and visual parsing. The generation stream denoises latents into images and video. Both streams share attention and operate over the same KV cache.
When a reply needs pixels, the understanding stream emits a discrete handoff token. The generation stream then renders from the full accumulated state. Training combines next-token prediction for discrete sequences with flow matching for continuous outputs.
This design has practical effects. Bounding boxes come out as coordinate tokens, not from a separate detector. A video’s first frame reuses the in-context image representation instead of a re-encoded copy.
Speed: Base vs Distilled
The base model generates video at 0.79x real-time (median), with a watchable stream starting in roughly 6 seconds. Reka team measured 7.0 seconds to a first clip, against an illustrative 13.8 seconds for a multi-agent pipeline.
A distilled variant cuts denoising from 99 steps to 8, with minimal quality loss reported. It returned a 5.3-second clip in about a second. In Reka’s internal tests, it matched the fastest dedicated image models. It was also the quickest model tested to the first text token. These are vendor-run tests, not independent benchmarks.
World Model and Robotics
Rho-1 streams continuously, clip after clip. New instructions enter through the understanding stream and update state mid-rollout. Reka demonstrates one opening forked into ‘bank left’ and ‘bank right’ continuations.
For robotics, actions and future frames decode from the same latent state. A LIBERO simulation episode shows Rho-1 emitting 7 action channels. To scale past scarce teleoperation logs, Reka pairs Rho-1 with its Inverse Dynamics Model, which infers control signals from raw video.
How Rho-1 Compares
Data verified on October 5, 2026 from official sources.
| Feature | Reka Rho-1 | ByteDance BAGEL | BAAI Emu3.5 | Google Genie 3 |
|---|---|---|---|---|
| Developer | Reka | ByteDance Seed | BAAI | Google DeepMind |
| Parameters | 19B | 14B total, 7B active (MoT) | 34B | Not disclosed |
| Inputs | Text, image, video, actions, proprioception | Text, image | Interleaved text and image | Text prompt, navigation inputs |
| Outputs | Text, image, video, actions, proprioception | Text, image | Interleaved text and image | Interactive video world |
| Native video generation | Yes, capped at 672×384 | No | No (image frames, not native clips) | Yes, 720p at 24 fps |
| Real-time steering | Yes, continuous rollouts | No | No | Yes, promptable world events |
| Robot actions | Native continuous action tokens | No | Embodied manipulation demos | Takes navigation actions, does not emit them |
| Open weights | No | Yes, Apache 2.0 | Yes, Apache 2.0 | No |
| Access today | Research preview via contact@reka.ai | Hugging Face, GitHub | Hugging Face, emu.world app | Project Genie for Google AI Ultra subscribers |
| Source/Resources | Reka blog | GitHub | Hugging Face | DeepMind blog |
Genie 3 access per Project Genie; Emu3.5 parameter count per its Hugging Face model card.
Key Takeaways
- Rho-1 is a 19B model that reads and writes text, image, video and robot actions.
- Two expert streams share attention and one KV cache in every block.
- Distillation cuts denoising from 99 to 8 steps, about 1 second per 5.3 s clip.
- Trained on 320 H100s for 3 months; video is capped at 672×384.
- Research preview only: no public weights, API or pricing yet.
Check out the Technical details and the announcement on X. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us
The post Reka Releases Rho-1: A 19B Omni-Reasoning Model That Understands, Generates Video and Outputs Robot Actions in One appeared first on MarkTechPost.
