Skip to content

About

Learning Spatial Reasoning through Interaction with the Observable Physical World

Topics

Resources

Stars

83 stars

Watchers

0 watching

Forks

Repository files navigation

Spatial-Interactor icon  Spatial-Interactor

Learning Spatial Reasoning through Interaction with the Observable Physical World

arXiv paper Project page Models and data on Hugging Face Code

Python 3.10+ Apache-2.0 license

Spatial-Interactor project introduction in English

🎞️ Visual Presentation

The full 20-slide English presentation accompanies the English project video above.

Spatial-Interactor 20-slide English presentation

📖 Overview

Spatial reasoning is not only a matter of recognizing static relations. An agent must also understand how the observable world changes after an action, how its own viewpoint changes during movement, and how successive transitions compose into a complete trajectory.

Spatial-Interactor learns these capabilities from observable physical interaction. It organizes interaction records into a progressive curriculum, uses supervised fine-tuning to learn local state transitions, and applies On-Policy Distillation (OPD) to integrate spatial evidence over long horizons. The released model receives only the original visual input and question at inference time.

Overview of Spatial-Interactor

🧭 Spatial Interaction Curriculum

The curriculum follows the increasing reasoning horizon of physical interaction:

  • L1, passive world-state transitions: infer how objects and scenes change under an external action.
  • L2, active self-state transitions: reason about ego-motion and its effect on position, distance, overlap, and visibility.
  • L3, long-horizon interaction trajectories: integrate ordered local changes into key-motion events and global paths.

Three-level spatial interaction curriculum

🧠 On-Policy Distillation

OPD combines verifiable answer rewards with training-only process supervision. For each student rollout, a privileged branch reads an ordered state-transition trace while the student branch sees only the video and question. Both branches score the same student-generated prefixes, and distillation is applied only to reasoning tokens. The privileged trace, teacher branch, reward functions, and reference policy are removed after training.

On-Policy Distillation pipeline

📦 Models and Data

The complete release is collected on Hugging Face.

Resource Description Access
LSI-108K 107,518 interaction-derived spatial QA pairs across L1, L2, and L3 🤗 Dataset
Spatial-Interactor-3B Qwen2.5-VL-3B checkpoint 🤗 Model
Spatial-Interactor-7B Qwen2.5-VL-7B checkpoint 🤗 Model
Spatial-Interactor-4B Qwen3-VL-4B checkpoint 🤗 Model
Spatial-Interactor-8B Qwen3-VL-8B checkpoint 🤗 Model

Source media are governed by their upstream licenses. When redistribution is not permitted, LSI-108K provides source and episode identifiers for obtaining the corresponding media from the original dataset.

🚀 Quick Start

1. Clone and validate the release

git clone https://github.com/ZJU-OmniAI/Spatial-Interactor.git
cd Spatial-Interactor
bash scripts/check_release.sh

2. Download LSI-108K

hf download kagakouko/LSI-108K \
  --repo-type dataset \
  --local-dir data/LSI-108K

The annotation schema, curriculum assignment, source-media manifest, and media access conditions are documented in Data.

3. Prepare the environment

The data tools, SFT stage, and OPD stage use separate environments. Follow the tested setup in Environment before launching training.

⚙️ Training

SFT learns local state transitions from L1/L2 and public spatial QA. OPD starts from the SFT checkpoint and learns long-horizon integration from L3 and VSTI trajectories.

# SFT
MODEL_PATH=/models/qwen-vl DATA_DIR=/data/sft OUTPUT_DIR=/outputs/sft \
  MODEL_FAMILY=qwen25vl bash training/sft/train_sft.sh

# OPD, initialized from the SFT checkpoint
MODEL_PATH=/outputs/sft DATA_DIR=/data/opd MEDIA_ROOT=/data/media \
  OUTPUT_ROOT=/outputs/opd PYTHON_BIN=/path/to/opd-env/bin/python \
  bash training/opd/scripts/train_opd.sh

See Training for data preparation and multi-model commands, On-Policy Distillation for the objective, and the paper-to-code map for implementation ownership.

🧪 Evaluation

The evaluation entrypoint delegates prompting and scoring to the corresponding upstream benchmark implementation while recording the run configuration and raw predictions.

CUDA_VISIBLE_DEVICES=0 python evaluation/run.py \
  --bench vsi --family qwen25vl \
  --model kagakouko/Spatial-Interactor-Qwen2.5-VL-7B \
  --toolkit ./VLMEvalKit --data-root /data/benchmarks \
  --output ./outputs/qwen25vl7b/vsi

Evaluation instructions cover VSI, VSTI, MindCube, SPBench-MV, MMSI, ViewSpatial, SAT-Real, and SAT-Syn. Use --dry-run to inspect a resolved command before loading model weights.

🏗️ Repository Layout

Spatial-Interactor/
├── data_generation/       interaction QA and privileged-trace construction
├── data_preparation/      curriculum, SFT-mixture, and OPD-data builders
├── training/
│   ├── sft/               LLaMA-Factory-based supervised fine-tuning
│   └── opd/               EasyR1/verl-based OPD, GRPO, and rewards
├── evaluation/            benchmark launchers and run manifests
├── docs/                  method, data, training, and evaluation guides
├── tests/                 lightweight data and objective tests
└── assets/                README figures and project media

🙏 Acknowledgements

We sincerely thank the authors of AI2-THOR, ProcTHOR, SIMS-V, HSSD, Replica, ScanNet, ScanNet++, ARKitScenes, MultiScan, RoomTour3D, and BridgeData V2 for their public environments, trajectories, and annotations. We also appreciate VSI-590K, MindCube, and STI-Bench for the public spatial QA used in the supervised training mixture.

The training implementation builds on LLaMA-Factory, EasyR1, and verl. We are grateful to their authors and maintainers for making these frameworks publicly available.

📝 Citation

@misc{yao2026spatialinteractor,
  title  = {Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World},
  author = {Yao, Kaixiang and Wang, Xu and Pan, Miao and Hu, Xiyue and Wang, Weishi and Dahlmeier, Daniel and Chen, Jintao and Shen, Yongliang and Zhang, Xuhong and Zhang, Wenqi},
  year   = {2026},
  eprint = {2609.23038},
  archivePrefix = {arXiv},
  primaryClass = {cs.AI},
  url    = {https://arxiv.org/abs/2609.23038}
}

⚖️ License

Spatial-Interactor source code is released under the Apache License 2.0. Third-party datasets, model checkpoints, and media remain subject to their original licenses and terms. See Third-Party Notices for the vendored training frameworks.

About

Learning Spatial Reasoning through Interaction with the Observable Physical World

Topics

Resources

Stars

83 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages