AGIBOT Demonstrates WITA-Omni in Live Human-Robot Interaction at The Greater Bay Area Film Concert

【MACAU, September 21, 2026】At "A Full Moon Rising Above the Greater Bay Area - The Greater Bay Area Film Concert", an AGIBOT robot interacted face-to-face with a celebrity in a live setting, following conversations, identifying who to respond to, and coordinating speech, movement and expression in real time.


Rather than performing a fixed sequence of gestures or delivering pre-set lines, the robots followed the interaction as it developed. They identified who was speaking, turned toward the person they were responding to, waited for the right moment to speak and coordinated their movements and expressions with their replies.


1789969506951848.gif

AGIBOT robot interacting face-to-face with a celebrity at The Greater Bay Area Film Concert


The demonstration offered a glimpse of a different kind of human-robot interaction, one that depends less on scripted commands and more on a robot’s ability to understand what is happening around it and respond appropriately as the situation changes.

Behind the interactions was WITA-Omni, AGIBOT’s embodied-native omni-modal model designed for real-world human-robot interaction. The model is built to help robots not only see and hear what is happening, but also understand who is interacting with whom, when to respond and how to turn that understanding into coordinated speech and physical behavior.

 

Beyond Multimodal Understanding: Built for Real-World Interaction
Most multimodal AI systems are built to process combinations of text, images and audio. For a robot operating in the physical world, that is only part of the challenge.

Real-world interactions are continuous and often unpredictable. A robot needs to understand who is speaking, what happened first, who a comment or gesture is directed at, and whether it is the right moment to respond.

WITA-Omni brings visual, audio, language and temporal information into a unified framework, allowing the model to follow an interaction over time and understand who said or did what, when it happened, and how different signals relate to one another.

This capability is also reflected in benchmark performance. WITA-Omni Preview ranked first overall on the DailyOmni benchmark, achieving an average accuracy of 85.21% and ranking first in six of eight evaluation dimensions.

 

From Understanding to Interaction
At the core of WITA-Omni is a Thinker-Talker-Actor architecture, which extends the Thinker-Talker framework with an Actor designed specifically for embodied interaction.

The Thinker interprets what is happening and makes high-level decisions about the interaction. The Talker generates speech based on that understanding, while the Actor controls movement and expression.

Rather than treating speech and movement as separate steps, WITA-Omni aligns them along the same timeline. This allows a robot to speak, move and express itself as part of the same response, with gestures and expressions adjusting to elements such as speech rhythm, pauses, emphasis and emotional cues.


1789955501891685.png

WITA-Omni Thinker-Talker-Actor architecture


WITA-Omni is trained on large-scale multimodal data centered on real-world human interaction. The training data preserves the timing and relationships between speech, visual information, language, movement and expression, giving the model the context needed to understand how an interaction develops over time.

A multi-stage training process further improves interaction decision-making, helping the model determine whether to respond, when to respond and whom to respond to in dynamic situations.

 

From Model Capability to Live Interaction
At The Greater Bay Area Film Concert, these capabilities were put to work in a live environment.

As conversations developed, the robots followed both the dialogue and the people involved, identified the intended interaction target, and responded at the appropriate moment instead of simply moving through a predetermined sequence.

  

When a robot responded, speech was only one part of the exchange. Voice, body movement and expression formed part of the same response, allowing the robot to communicate in a more coordinated and natural way.


The significance goes beyond making robot conversations feel more natural. As robots move into environments shared with people, they need to do more than recognize speech or follow instructions. They need to understand how an interaction is developing, decide when it is appropriate to take part, and respond in a way that fits the situation.

These capabilities could support robots in environments where interactions are dynamic and difficult to script in advance, including entertainment venues, retail spaces, hotels, reception areas and other service settings.

AGIBOT will continue developing WITA-Omni and extending its capabilities across its robotic platforms, with the goal of enabling robots to move beyond simply receiving instructions and take part more naturally in ongoing human interactions.


For more information, please visit AGIBOT.com and follow AGIBOT on:
X: https://x.com/AGIBOTofficial
LinkedIn: https://www.linkedin.com/company/agibot/
YouTube: https://www.youtube.com/@AGIBOTofficial
Instagram: https://www.instagram.com/agibotofficial/
TikTok: https://www.tiktok.com/@agibotofficial
Facebook: https://www.facebook.com/AGIBOTofficial/

 


About AGIBOT
AGIBOT is a pioneer in the global general-purpose AI robotics industry, with its core technology focused on embodied AI. AGIBOT's "Three Intelligences in One" architecture integrates Locomotion Intelligence, Interaction Intelligence, and Manipulation Intelligence into a unified embodied system. Its portfolio spans humanoid robots, quadrupeds, dexterous systems, and commercial cleaning solutions. In June 2026, AGIBOT announced that its 15,000th robot had rolled off the production line.