Unitree G1 voice commands

Unitree G1 Obeys Voice Commands and Improvises Moves in Real Time

A robot that jumps, drops into a plank, shakes its hips, and then busts out a Gangnam Style routine, all on spoken command and with no choreography loaded in advance. That is exactly what Unitree G1 pulled off in a video that Unitree Robotics released on May 19. The clip shows a female host giving the robot verbal instructions one after another, and the G1 responding in real time, with every movement generated on the fly by AI rather than recalled from a stored library.

One Take, No Edits, No Stored Moves

The footage was filmed in a single take, with the original on-set audio intact. There were no cuts, no split-shot tricks, and no motion corrections added in post. As observers noted, the robot is not replaying pre-recorded animations; the actions appear to be created live while the machine balances and reacts on the spot. The sequence escalates steadily. First a jump, then a plank, then hands on hips with a hip-shake, then walking forward and backward with arms raised. In the second half of the video the difficulty climbs further, with the G1 walking in a crouched posture, spinning right, dropping to one knee in a proposal gesture, and finally performing the full Gangnam Style dance.

Unitree was candid about the limits of what viewers are watching. The company acknowledged that slight delays and reduced smoothness during transitions may be noticeable, precisely because the movements are being generated autonomously rather than played back. That honesty is actually a point in the demo’s favor. The imperfections confirm the technology is working live, not hiding behind hidden edits.

What the Robot Has to Do in Seconds

To pull off any single command, the G1 must understand speech, process the request, calculate body movement, maintain balance, and execute the action, all within a few seconds. That is a significant technical challenge for a bipedal machine standing 1.32 meters tall and weighing around 35 kilograms. The base model ships with 23 degrees of freedom, while the EDU version reaches up to 43, giving it a joint range that goes well beyond what most humans can manage. On-board sensors include 3D LiDAR and a depth camera for spatial awareness, and the EDU configuration adds an NVIDIA Jetson Orin module for heavier AI workloads.

The G1 starts at around $16,000 for the base model, which makes it one of the most accessible production humanoids available. That price point has drawn interest from university labs, research startups, and developers who want a real platform for testing embodied AI without the six-figure budgets that earlier humanoids required.

Voice Control Opens a Wider Door

What this demo points toward is a shift in how people might interact with robots. No dedicated controller, no custom app, no pre-written script. A person speaks, and the machine moves. That kind of natural-language interface has obvious appeal across service, entertainment, and industrial settings, anywhere a robot needs to respond to instructions from someone who is not a robotics engineer.

Unitree has been building toward this. In March 2026, the company open-sourced UnifoLM-VLA-0, a Vision-Language-Action model that lets the G1 carry out household tasks through natural language commands. The voice-command demo released this week sits in the same direction of travel, showing that the gap between telling a robot what to do and watching it actually do it is getting narrower. Unitree says it continues to work on improving motion generation accuracy and response speed, so the slight lag visible in the current footage is likely to shrink with future updates.

Leave a Comment