Multimodal Full-duplex Real-time Interaction System for Humanoid Robots
Project Description
Developing a multimodal full-duplex real-time interaction system for humanoid robots is a critical step toward advancing human-robot interaction. However, current humanoid robots suffer from weak multimodal interaction capabilities—including speech, facial expressions, and actions—which limits their ability to engage in natural and expressive communication. Furthermore, most systems still rely on half-duplex interaction modes that cannot support natural interruptions or real-time feedback, thereby undermining the naturalness and immersion of human-robot interaction. While end-to-end speech-language models such as MiniCPM-o 4.5 natively support full-duplex dialogue and visual understanding, they are limited to audio output and lack both multimodal interaction capabilities and domain-specific knowledge for scenarios such as 4S store customer reception. This project aims to address these gaps by (1) designing an end-to-end full-duplex multimodal dialogue model architecture that outputs audio, facial expressions, and actions through a unified network while supporting user-initiated interruptions for more natural and fluid conversations; and (2) enabling real-world on-device deployment on humanoid robots in practical human-robot interaction scenarios such as 4S store customer reception and beyond.
Supervisor
OUYANG, Xiaomin
Quota
2
Course type
UROP1100
UROP2100
UROP3100
UROP3200
UROP4100
Applicant's Roles
(1) designing an end-to-end full-duplex multimodal dialogue model architecture that outputs audio, facial expressions, and actions through a unified network while supporting user-initiated interruptions for more natural and fluid conversations; and (2) enabling real-world on-device deployment on humanoid robots in practical human-robot interaction scenarios such as 4S store customer reception and beyond.
Applicant's Learning Objectives
1. Gain a solid foundation in efficient inference techniques for both large language models and mobile GUI agents.
2. Develop hands-on skills with model compression and acceleration techniques, specifically for mobile deployment.
3. Learn to balance trade-offs among accuracy, latency, and resource consumption in resource-constrained environments.
4. Gain experience in prototyping intelligent mobile applications and integrating multimodal systems for enhanced real-time interaction.
2. Develop hands-on skills with model compression and acceleration techniques, specifically for mobile deployment.
3. Learn to balance trade-offs among accuracy, latency, and resource consumption in resource-constrained environments.
4. Gain experience in prototyping intelligent mobile applications and integrating multimodal systems for enhanced real-time interaction.
Complexity of the project
Challenging