Realtime-Venus-Omni
Understand visual events, respond proactively, and recall relevant moments from long videos.
Introducing Realtime-Venus
Realtime-Venus keeps listening while it responds, connects sound with visual context, and handles delegated tasks in the background.
Three everyday moments: a follow-up question, a microwave beep, and a flight search.
Hear a road-trip conversation shift from planning the route to arranging daily breaks and meals.
Around 15 s, you ask a follow-up while Realtime-Venus is still speaking. Its next response addresses breaks and meals.
One loop handles live perception and speech. The other runs delegated tasks and prepares replies, returning results to the same conversation.

Understand visual events, respond proactively, and recall relevant moments from long videos.
Understand sound and speech, respond by voice, and distinguish interruptions from brief acknowledgments.
Preserve each request’s context, run delegated tasks, and prepare replies for the conversation.
Save the context available when a request arrives and link it to the conversation it came from.
Route each task to a multimodal model, a general model, or a registered skill while perception and conversation continue.
The harness prepares a reply, and the frontend decides when to speak. Playback tracking confirms that the reply was delivered.

For longer audio–visual sessions, Realtime-Venus-Omni stores selected video frames and retrieves them based on relevance and novelty. This recalled context is combined with recent observations.
The memory module requires no additional training and operates separately from the fine-tuned dialogue policy.

Evaluations cover video and audio understanding, full-duplex conversation, and delegation. Scores shown here come from the September 9, 2026 manuscript; each benchmark uses its own evaluation protocol.
Video benchmarks where Realtime-Venus-Omni leads the online models evaluated in the report.
StreamingBench · Omni
MMAU · Audio
Full-Duplex-Bench v1.5 · Audio
C_RESUME continuation rate
Omni and Audio refer to Realtime-Venus-Omni and Realtime-Venus-Audio.
| Benchmark | Model | Score |
|---|---|---|
| StreamingBench | Omni | 70.2% |
| OVO-Bench | Omni | 64.7% |
| Daily-Omni | Omni | 81.3% |
| MMAU | Audio | 78.0% |
| MMAU-Pro | Audio | 63.2% |
| Llama Questions | Audio | 83.8% |
| Speech CMMLU | Audio | 67.8% |
See the technical report for datasets, prompts, and the full model comparisons.
Full-Duplex-Bench v1.5 measures whether Realtime-Venus-Audio continues speaking when a cue does not call for interruption (C_RESUME).
Response rate to user interruptions
C_RESPOND · measured separately

Realtime-Venus: A full-duplex interaction system with asynchronous delegation
arXiv:2609.13814 · September 12, 2026
Read the technical report@misc{zhao2026realtimevenus,
title = {{Realtime-Venus}: A full-duplex interaction
system with asynchronous delegation},
author = {Ruixiang Zhao and Hualei Wang and
Renhe Sun and Enzhi Zhou and
Jincenzi Wu and Xujie Song and
Kexin Shi and Zihang Liu and
Pengcheng Zhu and Jiayi Zhou and
Baoyue Zhang and Changhao Zhang and
Zitong Wang and Jinhong Wang and
Tong Niu and Jingjing Liu and
Junan Lin and Haolin He and
Hengshuo Chu and Yuhui Chen and
Jian Liu and Yuge Huang and
Junliang Xing and Yuntao Wang and
Weiqiang Wang and Chun Yu and
Yuanchun Shi},
year = {2026},
eprint = {2609.13814},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2609.13814}
}Video could not play. Try again or open the video.
Video could not play. Try again or open the video.
Video could not play. Try again or open the video.