SuperNav Learned Executor
Learned local-navigation checkpoint for SuperNav: An Agentic Navigation System for Any Task in Any Scene.
- Code: zju3dv/SuperNav.
- Checkpoint:
ckpt_latest125.pt(249,157,807 bytes). - Role: the learned local executor, served through SuperNav's
nomadbackend. The high-level agent, simulator, scene assets, and task manifests are separate requirements. - Format: PyTorch checkpoint containing
state_dictandtrain_config. The SuperNav loader reads the checkpoint's architecture settings.
Architecture
The checkpoint metadata selects a fine-tuned DINOv2 ViT-S/14 encoder (dino-s-finetune) with 224-pixel input images, GCA context (context_size=4, memory_budget=12), and a flow-matching action head with four inference steps. It predicts an eight-waypoint action horizon with waypoint scale 0.25 meters. The SuperNav loader reads these settings from the checkpoint.
Download
Install the Hugging Face CLI with python -m pip install -U huggingface_hub, then run:
hf download the0xka1/SuperNav-Learned-Executor \
ckpt_latest125.pt CHECKSUMS.sha256 --local-dir ./checkpoints
(cd checkpoints && sha256sum -c CHECKSUMS.sha256)
The checkpoint's SHA256 is:
efd3f50def5f65d0aa57da1d2bbe1b159cf7b14e870d82495d4b4167c0aa62dd
Start the policy service
Use Python 3.12 and install compatible PyTorch and torchvision builds for your device. For CUDA 12.1, the following example uses official PyTorch wheels:
python -m pip install torch==2.5.1 torchvision==0.20.1 \
--index-url https://download.pytorch.org/whl/cu121
Install the published SuperNav code and its runtime dependencies in the policy-service environment:
git clone https://github.com/zju3dv/SuperNav.git
cd SuperNav
python -m pip install -e '.[agents,evaluation]'
export TORCH_HOME=/absolute/path/to/torch-cache
export NAV_LOCALNAV_CKPT=/absolute/path/to/checkpoints/ckpt_latest125.pt
python -m supernav.methods.localnav.server \
--backend nomad --checkpoint "$NAV_LOCALNAV_CKPT" \
--host 127.0.0.1 --port 18914 --device cuda
Set TORCH_HOME to a writable cache directory. The DINOv2 loader uses Torch Hub to obtain the upstream repository and initial ViT-S/14 backbone weights (about 84.2 MiB) before loading this checkpoint. The first startup requires access to GitHub and dl.fbaipublicfiles.com, or an already populated Torch Hub cache. xFormers is optional.
For CPU inference, use --device cpu. In another terminal, check:
curl --fail http://127.0.0.1:18914/healthz
The response should report backend: nomad and the downloaded checkpoint path. Keep the service running for the experiment. If using an HTTP proxy, include 127.0.0.1,localhost in NO_PROXY and no_proxy so local policy and bridge requests bypass it.
Run a navigation task
Follow the published Habitat setup guide to prepare Habitat-GS, scene assets/NavMeshes, the agent, and an external task manifest. In the agent environment:
export SUPERNAV_HABITAT_PYTHON=/absolute/path/to/habitat-env/bin/python
export SUPERNAV_SCENE_DATASET_CONFIG=/absolute/path/to/scene_dataset_config.json
export NAV_LOCALNAV_URL=http://127.0.0.1:18914
supernav run --experiment habitat-learned-executor \
--instructions /absolute/path/to/tasks.json \
--task-ids example-task --arms default \
--model gpt-6-astra --output-dir data/runs/learned-executor
Replace the paths and task ID with your own assets and manifest. If the Habitat bridge runs on a different host, set NAV_LOCALNAV_URL to an address it can reach.
License
This checkpoint is distributed under the Project Registration License (PRL) v1.0. The DINOv2 backbone retains its Apache-2.0 terms. External dependencies, simulator assets, and datasets retain their respective licenses; see third-party notices.