What changed
WITA-Omni uses a thinker-Talker-Actor architecture. Thinker is multi-model inference core of text, picture, sound & video. Talker generates natural language & actor makes decision in physical motion.
This model is trained 10s of millions of hours of open sourced + proprietary multi-modal data. — via @tphuang
Why it matters
Agibot's WITA-Omni Preview leads DailyOmni leaderboard over Qwen, Gemini, Doubao, Nemotron.