Skip to main content

Tsinghua's Script-Free Demo: Robot Dog Commands Humans to Weigh Objects as Physical AGI Steps Out of the Fog

While the industry debated competing technical roadmaps for Physical AGI, a team at Tsinghua University chose the most direct — and the boldest — approach: they brought a robot dog in front of dozens of live spectators, with no script, no pre-programmed tasks, and no rehearsal, and ran a live demo.

In July 2026, the Unisonmind team held an extraordinary demonstration at Tsinghua. No polished demo videos. No pre-prepared benchmarks. Audience members were free to propose tasks on the spot. This was not merely a technology validation — it felt more like a public cross-examination of whether Physical AGI actually exists.

A Live Trio: Reading Drawings, Commanding, and Estimating

Powered by the Unisonmind on-device brain, the robot dog "Xiaotian" demonstrated impressive generalization across three entirely different tasks.

The first task: decipher a maze from a drawing. The host casually sketched a simple picture on the blackboard. The robot dog understood it represented the three-dimensional maze behind it. When the host circled a location on the drawing, the dog walked over to find the corresponding object in the physical maze. The highlight came when the host drew a figure-eight — Xiaotian first asked, "Can I take a shortcut?" and, upon receiving a yes, cut across the maze, then immediately resumed following the drawn path for the remainder.

The second task: command a human to weigh objects. The robot dog needed to instruct a human assistant to use a balance scale to weigh several randomly selected items. This required simultaneous deployment of physical intelligence, verbal communication, action planning, and social coordination — not a single-dimensional capability, but a complete behavioral sequence in which multiple intelligence facets worked together.

The third task: estimate the remaining water in a bottle. An audience member spontaneously produced two partially consumed water bottles and asked how much water remained. This was not a prepared question. Xiaotian glanced at them and said, "First, peel off the labels." It understood that a transparent bottle was necessary for visual estimation. It then gave a reasonable estimate.

One Brain, Multiple Bodies

Even more noteworthy: the same Unisonmind on-device brain was installed in different physical forms — another robot dog, a humanoid robot, and an electric wheelchair. The bodies changed, but the cognitive foundation remained the same. All devices operated in real-time within real-world environments, forming a continuous "perception-action-feedback" loop.

This means Unisonmind is not a perception module customized for a specific robot but a general cognitive system capable of cross-embodiment transfer. For the embodied AI industry, long constrained by the "one robot, one brain" paradigm, this is a significant directional signal.

The Technical Foundation: "3+1" Unified World Token Space

Behind the live demonstrations lies the Unisonmind team's "Unified World Token Space" architecture. The team summarizes its core features as "3+1":

First, full-modal Any-to-Any. Video, images, audio, text, and actions all flow in and out of a single model, without modal partitioning. Second, unified understanding and generation. A route is seen, a command is understood, a state is predicted, an action is generated, and a voice is output — all happening within the same model, with no need for chaining multiple models together. Third, full-chain Runtime. The model runs continuously, realigning with the world state every 18 milliseconds, self-correcting through constant feedback. Finally, on-device deployment. The entire model runs directly on the edge hardware of the robot dog, humanoid, and wheelchair — no cloud dependency.

This "3+1" explains why the robot dog could seamlessly switch between three radically different tasks: language, spatial reasoning, physics, action, and coordination all combine and collaborate dynamically within a single continuously running model, rather than switching between separate "skill modules."

The Threshold of Physical AGI

This demonstration matters because it provides a verifiable standard for Physical AGI. Historically, industry discussions have oscillated between two extremes — either drifting into boundless speculation at the conceptual level, or retreating into arguments over virtual benchmark scores. The Unisonmind team chose a more primitive and more reliable benchmark: human intelligence itself.

Drawing from the theory of multiple intelligences, the core definition of Physical AGI is a single general cognitive system that enters the real world through a physical embodiment, continuously organizes multi-dimensional intelligence in open environments, and iterates autonomously through environmental feedback. It is not a "language model plus a robotic arm" glued together, but a cognitive system whose capabilities naturally differentiate across multiple dimensions when confronted with different objects.

Everything in the Tsinghua live demo validates this definition: a single brain, a real-world closed loop, unscripted tasks, and multi-faceted intelligence working in concert. In this sense, Physical AGI is no longer a vague technological aspiration — it has stepped out of the fog of definitions and onto the solid ground of the real world.