News Ticker-
Business News · Robot24.com Original

Stanford, Caltech Give Humanoid Spatial Memory for Household Tasks

HomeBody research system uses a digital twin, persistent spatial memory and reusable motor skills to let a Unitree G1 perform multistep tasks in an unfamiliar kitchen.

Diagram comparing a humanoid VLA pipeline with HomeBody, where GPT Astra selects navigation, picking and drawer-opening skills from a reusable robot skill library.

STANFORD, Calif. — Stanford University and Caltech researchers developed a system that gives a humanoid robot persistent spatial memory, allowing it to find objects that have left its camera view and carry out multistep tasks in an unfamiliar kitchen.

The research system, called HomeBody, tests whether stronger vision-language models can directly trigger specialized robot skills using stored spatial maps to plan multi-step tasks. The researchers tested it with a Unitree G1 humanoid guided by OpenAI's GPT Astra.

That architecture differs from approaches that use a learned vision-language-action model to translate high-level instructions into robot actions. HomeBody instead tests whether increasingly capable vision-language models can coordinate specialized robot skills directly while relying on persistent spatial information to plan longer tasks.

In demonstrations, the robot explored an unfamiliar kitchen and recorded camera images, LiDAR scans, joint poses and waypoints. HomeBody uses that data to build a digital twin in Nvidia Isaac Sim, giving the system a spatial model it can use after objects leave the robot's view.
The researchers showed the G1 gathering coffee bags on a kitchen island, throwing away certain milk and orange juice cartons and retrieving a medicine bottle from a drawer. For the medicine task, the bottle wasn't visible when the instruction arrived, so the system had to remember which drawer it came from.
GPT Astra handles high-level decisions, including selecting a skill and its target. Physical actions are performed through a library covering navigation, picking, placing, opening drawers and retrieving objects from them.

The researchers designed the system to handle some failures rather than assuming every commanded action will work on the first attempt. Individual skills retry failed actions and report errors to the vision-language model, which then selects an alternative action or repositions the robot.

HomeBody's spatial memory reduces repeated data transfer by storing object and location details and retrieving them on demand. Instead of continually supplying raw observations from the robot's exploration, the system stores information about objects and locations in its world representation and retrieves relevant details when needed.

The project is a research demonstration, not proof the system is ready for home use. The published examples cover selected tasks in a single kitchen environment and may not generalize to other homes or tasks.

The researchers also report practical limitations. Calls to GPT Astra can introduce reasoning delays, while the system's Real2Sim adds setup time and API fees. Its local perception and motion-planning software runs on a laptop with an Nvidia RTX 4090 GPU, and the researchers said the G1's finger servos overheated during long runs.
The project's GitHub repository says its code will be available soon meaning the complete system is not yet available there for outside researchers to reproduce.

 

Free Weekly

BUSINESS NEWS WEEKLY LETTER

The Weekly Letter for Robotics Professionals, Summarizing the Most Important Industry Moves, Launches, Deals and Signals.