VI3

Grounding Pretrained 3D Foundation Models with Inertial Cues

Ernesto Lozano, Alberto Jaenal, Javier Civera

Universidad de Zaragoza

arXiv 2026 (preprint)
Scroll down

VI3

Grounding Pretrained 3D Foundation Models with Inertial Cues

Ernesto Lozano, Alberto Jaenal, Javier Civera

Universidad de Zaragoza

arXiv 2026 (preprint)
Figure

Abstract

3D foundation models (3DFMs) excel at predicting camera poses and dense depth from multiple views of a scene, showcasing strong zero-shot generalization. However, as metric scale is not observable from monocular images, their absolute scale predictions are typically inaccurate. Inertial measurement units (IMUs), present in most devices, naturally complement monocular cameras by observing scaled motion. We introduce VI3, a model-agnostic framework that metrically anchors a pretrained 3DFM using only IMU readings. VI3 initializes and preintegrates the IMU to obtain a metric motion reference, which is then used to recover the scale of the 3DFM outputs. Our method includes adaptable anchoring strategies tailored to diverse 3DFM architectures. Experiments on synthetic and real aerial datasets demonstrate that VI3 recovers metric scale without ground-truth supervision while preserving geometric consistency, acting as a fine refinement under well-conditioned motion and as a strong prior when motion is less informative.

Overview

Method block diagram

Interactive Results

Drag to rotate · scroll to zoom · shift+drag to pan
Loading point cloud…
3DFM IMU Ours

Citation

@misc{lozano2026vi3groundingpretrained3d,
      title={VI3: Grounding Pretrained 3D Foundation Models with Inertial Cues}, 
      author={Ernesto Lozano and Alberto Jaenal and Javier Civera},
      year={2026},
      eprint={2609.03824},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2609.03824},
}