Simplex P0: from seeing people to understanding interaction intent
Combining audio, gaze, pose and expression for continuous interaction perception, with a 10 fps deployment on the D-Robotics S100.
Understanding human intent is a central problem in human–computer interaction. Is someone speaking to the AI or to another person? Are they walking past, or preparing to engage?
Simplex P0 is our perception model for next-generation interaction systems. It aims to interpret the scene continuously, before a spoken command arrives.
Beyond a wake word
Wake words create an explicit entry point and help reduce accidental activations. A familiar loop is wake, question, answer. In a moving, multi-person scene, that loop does not establish who is speaking, whom they are addressing, or whether they are still engaged.
P0 treats participation as a continuously updated judgment.
Bringing signals together
The system direction combines voice characteristics, microphone-array direction and speech semantics with gestures, head orientation, gaze, facial expressions, facial features and speaking activity. Signals need to be aligned with people and time to become useful evidence of intent.
Looking at a robot does not necessarily request an answer. Speech does not necessarily address it. Combining these signals helps distinguish preparation to engage, listening, departure and simple action intentions. We also explore the broader scene-understanding capabilities of video VLMs.
Performance and continuous decisions
Based on the team's current evaluations, P0 matches or exceeds the compared frontier specialist models across gaze tracking, speaker detection, gesture, expression and face-recognition tasks, reaching SOTA-level results within the evaluated scope.
This is a team-reported, preliminary conclusion, not a claim of leadership across every dataset or setting. Task-specific comparators, splits and metrics will accompany subsequent reports.
We target continuous decisions above 10 Hz. With support from our partner D-Robotics, our current S100 deployment achieves 10 fps continuous decisions. This is the processing frequency of that configuration, not a guarantee that arbitrary video-VLM requests finish within 100 ms or a substitute for end-to-end latency measurements.
Participating naturally
Continuous perception can help a robot anticipate engagement and adjust when someone leaves or turns to another person. It provides a foundation for natural interaction without a wake word.
Perception still works alongside dialogue state, action rules and uncertainty handling. When evidence is ambiguous, continued observation may be more appropriate than a response.
Demo in preparation: continuous perception and wake-word-free interaction on S100.