Learning Spatial Reasoning through Interaction with the Observable Physical World
The full 20-slide English presentation accompanies the English project video above.
Spatial reasoning is not only a matter of recognizing static relations. An agent must also understand how the observable world changes after an action, how its own viewpoint changes during movement, and how successive transitions compose into a complete trajectory.
Spatial-Interactor learns these capabilities from observable physical interaction. It organizes interaction records into a progressive curriculum, uses supervised fine-tuning to learn local state transitions, and applies On-Policy Distillation (OPD) to integrate spatial evidence over long horizons. The released model receives only the original visual input and question at inference time.
The curriculum follows the increasing reasoning horizon of physical interaction:
- L1, passive world-state transitions: infer how objects and scenes change under an external action.
- L2, active self-state transitions: reason about ego-motion and its effect on position, distance, overlap, and visibility.
- L3, long-horizon interaction trajectories: integrate ordered local changes into key-motion events and global paths.
OPD combines verifiable answer rewards with training-only process supervision. For each student rollout, a privileged branch reads an ordered state-transition trace while the student branch sees only the video and question. Both branches score the same student-generated prefixes, and distillation is applied only to reasoning tokens. The privileged trace, teacher branch, reward functions, and reference policy are removed after training.
The complete release is collected on Hugging Face.
| Resource | Description | Access |
|---|---|---|
| LSI-108K | 107,518 interaction-derived spatial QA pairs across L1, L2, and L3 | 🤗 Dataset |
| Spatial-Interactor-3B | Qwen2.5-VL-3B checkpoint | 🤗 Model |
| Spatial-Interactor-7B | Qwen2.5-VL-7B checkpoint | 🤗 Model |
| Spatial-Interactor-4B | Qwen3-VL-4B checkpoint | 🤗 Model |
| Spatial-Interactor-8B | Qwen3-VL-8B checkpoint | 🤗 Model |
Source media are governed by their upstream licenses. When redistribution is not permitted, LSI-108K provides source and episode identifiers for obtaining the corresponding media from the original dataset.
git clone https://github.com/ZJU-OmniAI/Spatial-Interactor.git
cd Spatial-Interactor
bash scripts/check_release.shhf download kagakouko/LSI-108K \
--repo-type dataset \
--local-dir data/LSI-108KThe annotation schema, curriculum assignment, source-media manifest, and media access conditions are documented in Data.
The data tools, SFT stage, and OPD stage use separate environments. Follow the tested setup in Environment before launching training.
SFT learns local state transitions from L1/L2 and public spatial QA. OPD starts from the SFT checkpoint and learns long-horizon integration from L3 and VSTI trajectories.
# SFT
MODEL_PATH=/models/qwen-vl DATA_DIR=/data/sft OUTPUT_DIR=/outputs/sft \
MODEL_FAMILY=qwen25vl bash training/sft/train_sft.sh
# OPD, initialized from the SFT checkpoint
MODEL_PATH=/outputs/sft DATA_DIR=/data/opd MEDIA_ROOT=/data/media \
OUTPUT_ROOT=/outputs/opd PYTHON_BIN=/path/to/opd-env/bin/python \
bash training/opd/scripts/train_opd.shSee Training for data preparation and multi-model commands, On-Policy Distillation for the objective, and the paper-to-code map for implementation ownership.
The evaluation entrypoint delegates prompting and scoring to the corresponding upstream benchmark implementation while recording the run configuration and raw predictions.
CUDA_VISIBLE_DEVICES=0 python evaluation/run.py \
--bench vsi --family qwen25vl \
--model kagakouko/Spatial-Interactor-Qwen2.5-VL-7B \
--toolkit ./VLMEvalKit --data-root /data/benchmarks \
--output ./outputs/qwen25vl7b/vsiEvaluation instructions cover VSI, VSTI, MindCube,
SPBench-MV, MMSI, ViewSpatial, SAT-Real, and SAT-Syn. Use --dry-run to inspect
a resolved command before loading model weights.
Spatial-Interactor/
├── data_generation/ interaction QA and privileged-trace construction
├── data_preparation/ curriculum, SFT-mixture, and OPD-data builders
├── training/
│ ├── sft/ LLaMA-Factory-based supervised fine-tuning
│ └── opd/ EasyR1/verl-based OPD, GRPO, and rewards
├── evaluation/ benchmark launchers and run manifests
├── docs/ method, data, training, and evaluation guides
├── tests/ lightweight data and objective tests
└── assets/ README figures and project media
We sincerely thank the authors of AI2-THOR, ProcTHOR, SIMS-V, HSSD, Replica, ScanNet, ScanNet++, ARKitScenes, MultiScan, RoomTour3D, and BridgeData V2 for their public environments, trajectories, and annotations. We also appreciate VSI-590K, MindCube, and STI-Bench for the public spatial QA used in the supervised training mixture.
The training implementation builds on LLaMA-Factory, EasyR1, and verl. We are grateful to their authors and maintainers for making these frameworks publicly available.
@misc{yao2026spatialinteractor,
title = {Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World},
author = {Yao, Kaixiang and Wang, Xu and Pan, Miao and Hu, Xiyue and Wang, Weishi and Dahlmeier, Daniel and Chen, Jintao and Shen, Yongliang and Zhang, Xuhong and Zhang, Wenqi},
year = {2026},
eprint = {2609.23038},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
url = {https://arxiv.org/abs/2609.23038}
}Spatial-Interactor source code is released under the Apache License 2.0. Third-party datasets, model checkpoints, and media remain subject to their original licenses and terms. See Third-Party Notices for the vendored training frameworks.


