Text-based Hierarchical Motion Planning and 3D Scene-Aware Interaction Synthesis
Implementing Organization
Indian Institute Of Technology Bombay
Principal Investigator
Mr. Shashikant Verma
Indian Institute Of Technology Bombay
shashikant.verma@iitgn.ac.in
Project Overview
Generating realistic human motion has long been a challenge in computer graphics and vision, due to its importance in areas like gaming, virtual reality, and film. While recent advances have improved motion quality, isolated realism alone is not enough. Synthesized motion must also be coherent with the surrounding 3D scene context. This includes accounting for physical constraints, object placement, and how people typically interact with their surroundings. As a result, there is growing interest in Human-Scene Interaction (HSI) and Human-Object Interaction (HOI), which aim to model how people move in response to their environments.
Recent progress in generative AI has opened new directions in this area, especially through the use of multiple input signals such as language commands, task goals, and 3D scene information. However, current models still face key limitations: they are trained on limited data, can only handle simple interactions, and often produce short motion sequences. While tasks like 'sit on a chair' are feasible, more complex ones, like 'find a chair and place it beside the sofa', require deeper semantic scene understanding and multi-step motion planning, which remain underexplored.
Our goal is to develop models that can handle such long-horizon tasks by learning to break them down into meaningful steps or milestones. These systems will be designed to translate high-level goals into structured sub-tasks while ensuring smooth and consistent motion throughout. To enable this, we will build hierarchical systems for both scene understanding and motion planning.
A major bottleneck is the lack of diverse training data. Capturing realistic motion in varied 3D environments is difficult and expensive. To address this, we propose to leverage a physics-based motion acquisition pipeline using synthetic environments and rigged avatars, inspired by advancements in the gaming domain. The gaming world has long relied on manually created, rigged avatars capable of performing a wide range of human motions with a high degree of control. Building on this, we aim to develop an interactive, game-like simulation environment where participants perform high-level tasks in virtual 3D scenes. This setup will enable us to collect goal-oriented motion trajectories at scale, forming the foundation for training our models.