Research Radar/Robotics/USA · China
Robot trained only in a simulator crossed 10 of 15 messy rooms
TANGO sets all 29 joints of a humanoid at once, so it can duck, squat and sidestep while it walks. Trained only in simulation, it finished 10 of 15 real cluttered runs against 6 for the rival.
TANGO is a model that turns a spoken instruction and camera images into all 29 joint angles of a humanoid robot at once. Trained only in simulation, it crossed cluttered real rooms in 10 of 15 tries, against 6 of 15 for the rival method. In simulation it cut the share of runs with a collision from 15.81 percent to 9.90 percent.
What happened
Picture yourself carrying a laundry basket down a narrow hallway. A chair sticks out. A coat hangs low. You do not stop and redraw your route. You turn your shoulders, tuck your elbows and duck a little. Your whole body solves the problem while you keep walking.
Robots are bad at this. Most walking robots treat a room as a flat map. They pick a line on the floor and follow it. That works in an empty hall. It fails when a chair leg blocks the way or a shelf hangs at head height.
A team of researchers built a system called TANGO. You give it a plain instruction, such as walk to the kitchen. It looks through two cameras. It then sets the angle of all 29 joints of a humanoid robot at once. Arms, torso and legs move as one. The robot can sidestep, squat or step over things while it keeps walking.
The test
The paper is called “TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model”. Anqi Li, Yuxin Chen, Zhaobo Li and colleagues wrote it. Their labs are the University of California, Berkeley, Peking University, Tsinghua University, the University of Hong Kong and Princeton University. It went online as a preprint on 8 September 2026. No journal or conference is named.
The robot never trained in a real room. All the training happened inside a simulator. The team took 578 indoor scenes and scattered extra obstacles through them. Software then planned safe paths, bent the walking motion around each obstacle, and checked that a real body could do it. That produced 64,633 practice runs. Making them took 211 hours on high-end graphics cards.
Then came three checks. The first was a navigation benchmark called VLNVerse, with 825 test routes in rooms the model had not seen. The second used those same rooms with extra clutter added. The third used a real Unitree G1 humanoid in a real office, with no extra training. The real test had three settings. Each setting had three scenes and five tries per scene. That is 15 tries per setting.
The result
In the hardest real test, a cluttered room with a blocking obstacle, TANGO finished 10 tries out of 15. The rival method finished 6 out of 15. TANGO also bumped into things less. It averaged 0.73 collisions per try. The rival averaged 1.93.
In simulation the pattern held. On cluttered scenes TANGO succeeded 43.75 percent of the time. The strongest baseline reached 41.88 percent. The bigger gap was safety. The share of runs with at least one collision fell from 15.81 percent to 9.90 percent. TANGO did that with ordinary colour cameras. The rival also had a laser scanner.
On the plain navigation benchmark TANGO reached 52.89 percent success on unseen routes. The best baseline there reached 48.60 percent. One detail matters. Every rival was tested by teleporting it along the route. Only TANGO had to really walk.
What it means
There are two lessons here. The first is about bodies. A robot that plans only on the floor plan will keep failing in real homes and shops. Real rooms are full of low shelves, cables and half-open doors. Planning the whole body at once is the fix this paper argues for.
The second is about cost. The training data took 211 graphics card hours. No person drove the robot. No real room was needed. If that keeps working, teaching robots gets much cheaper. The hard part moves from filming demonstrations to building good simulators.
There is a warning too. The team switched off one small trick that keeps motion smooth between decisions. Success then fell from 43.75 percent to 10.94 percent. In other words, much of the gain sits in fiddly engineering, not in the big model.
Business ideas from this paper
- A robot passability report. You walk a building with a phone, film the tight spots, and hand back a short list of what a walking robot would hit and what to move. Who buys it: warehouse, hotel and care home managers who are about to trial a robot. A price to test: 400 dollars per site. A one-week test: offer it to five local sites that already run a robot pilot, and see whether two of them pay a deposit.
- A clutter pack for robot trainers. You sell bundles of 3D indoor scenes with messy, realistic obstacles, ready to drop into a simulator. This paper had to build its own because standard scenes were too tidy. Who buys it: robotics teams and university labs that train walking robots. A price to test: 200 dollars a month per team. A one-week test: post 20 free sample scenes, then count the downloads and the emails asking for more.
- A bump log for robots already at work. A small tool records every knock a cleaning or delivery robot takes and ranks the worst spots in the building. Who buys it: facilities managers running floor cleaning robots. A price to test: 30 dollars per robot per month. A one-week test: log bumps by hand for one week in one building, then show the manager the top five trouble spots and ask what that is worth.
How sure can you be?
Not very sure yet. Start with the size of the real test. Fifteen tries per setting is small. Luck can move 10 out of 15 by quite a lot. The paper does not report error bars for those trials.
Next, look at the comparison. Only one rival method was run on the real robot. The simulation gaps are thin as well. Success rose by 1.87 percentage points over the strongest baseline. That is close.
The authors flag two limits themselves. The low-level controller that drives the joints is the main constraint, and it cannot yet handle stairs. The robot also sees only colour images. The authors expect depth cameras and laser scanners would help in dark or confusing rooms.
The paper says the team will release the data pipeline, the dataset, the model and the deployment code. Right now that is a promise, not a link. HERE IS WHAT WOULD SETTLE IT. Another group runs the released code on its own robot, in its own building, with many more tries.
Do this today
Walk through your workspace and look for what a robot would hit. Low shelves, trailing cables, chairs left out. Write down the five worst spots. That short list is the real cost of any robot you buy.
Source: TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model, September 2026. arXiv:2609.09158 · arxiv.org (preprint · not yet peer reviewed).
Just Out Tech explains new research in plain language. This article was drafted with AI assistance and checked by a human against the original source.
- TANGO completed 10 of 15 real-world runs through cluttered rooms on a Unitree G1 humanoid, against 6 of 15 for the fine-tuned InternVLA-N1 baseline.
- In cluttered simulated scenes TANGO cut the share of runs containing at least one collision from 15.81 percent to 9.90 percent, using only colour cameras while the baseline also had a laser scanner.
- The whole training set of 64,633 robot trajectories was generated in simulation for 211 graphics card hours, with no real-world navigation data at all.
Questions people ask
what is a vision-language-action model for robots?
It is one model that takes in pictures and a written instruction and puts out robot movements directly. There is no separate planner in the middle. In TANGO the output is the angle of all 29 joints of a humanoid, predicted in short chunks. A lower-level controller then drives the motors to match.
why does whole-body control matter for a walking robot?
A room is not flat. Shelves hang low and chairs stick out. If a robot only plans a line on the floor, its arms and head can still hit things. TANGO plans the arms, torso and legs together, so it can duck under or step over an obstacle instead of only going around it.
did the robot ever practise in a real room?
No. The paper says the model was trained entirely in simulation, on 64,633 generated trajectories, with no real-world navigation data. It was then run on a Unitree G1 humanoid straight away. The authors call this zero-shot transfer.
is the code available?
Not yet, as far as the paper shows. The authors write that they will open-source the data pipeline, the dataset, the model checkpoint and the deployment system. The paper lists a project website but does not link a public repository. Treat it as a promise until you can download it.