A high-performance Direct Sparse Visual Odometry (DSO) implementation written in modern C++. The project estimates the motion of a monocular camera directly from image intensities rather than extracted feature descriptors, making it particularly suitable for high-accuracy visual odometry under photometrically calibrated conditions.
Originally developed at the Technical University of Munich (TUM), this implementation closely follows the paper:
Direct Sparse Odometry
Jakob Engel, Vladlen Koltun, Daniel Cremers
DSO is a real-time sparse direct SLAM / Visual Odometry framework that tracks camera motion by minimizing photometric error over carefully selected image pixels.
Unlike traditional feature-based methods (ORB, SIFT, SURF, FAST, etc.), DSO operates directly on image brightness values and therefore avoids descriptor extraction and feature matching.
The repository contains:
- complete visual odometry pipeline
- nonlinear optimization backend
- coarse-to-fine tracking
- sparse map generation
- point activation / marginalization
- photometric calibration support
- dataset loader
- visualization framework
- reusable static library
Instead of matching image features, DSO minimizes
- intensity differences
- photometric residuals
- image gradients
between successive frames.
Advantages include
- sub-pixel precision
- improved accuracy
- better utilization of image information
- no descriptor computation
Rather than optimizing every pixel, DSO intelligently selects informative pixels with sufficient image gradient.
This dramatically reduces computation while preserving accuracy.
Components include
- adaptive pixel selection
- gradient-based sampling
- multi-scale image pyramids
- sparse residual graph
Supports motion estimation using only
- one camera
- no IMU
- no stereo camera
- no depth sensor
Camera pose is estimated continuously as images arrive.
Unlike many visual odometry systems, DSO explicitly models
- camera response function
- exposure changes
- vignette correction
- affine brightness changes
This significantly improves robustness under varying illumination.
The system maintains a fixed-size optimization window consisting of recent keyframes.
Old frames are
- marginalized
- compressed into priors
- removed from active optimization
This allows constant computational complexity while preserving historical information.
The optimization backend jointly estimates
- camera poses
- inverse depths
- affine brightness parameters
using nonlinear least squares optimization.
The system automatically decides when new keyframes should be inserted based on
- camera motion
- scene overlap
- optimization quality
Tracking is performed from
- coarse
- medium
- fine
image resolutions.
Benefits include
- larger convergence radius
- robustness
- stable optimization
The coarse tracker aligns incoming frames against existing keyframes without feature extraction.
Tracking includes
- pose prediction
- coarse optimization
- fine optimization
- photometric refinement
The implementation is heavily optimized using
- SSE vectorization
- Eigen linear algebra
- sparse matrix operations
- efficient memory layouts
- incremental optimization
Real-time operation is one of the primary design goals.
src/
│
├── FullSystem/
│ Complete visual odometry pipeline
│
├── OptimizationBackend/
│ Nonlinear optimization
│
├── IOWrapper/
│ Image loading
│ visualization
│ dataset handling
│
├── Util/
│ Math utilities
│
└── main_dso_pangolin.cpp
Dataset application
The central component implementing the complete visual odometry pipeline.
Contains:
- tracking
- mapping
- optimization
- keyframe creation
- marginalization
- point activation
Important classes include
FullSystemCoarseTrackerCoarseInitializerImmaturePointResidualsHessianBlocks
Responsible for solving the nonlinear optimization problem.
Includes
- Hessian construction
- Schur complement
- sparse solvers
- Gauss-Newton optimization
- Levenberg-Marquardt style iterations
Uses
- Eigen
- SuiteSparse
Initializes the entire odometry pipeline from the first frames.
Responsibilities include
- initial depth estimation
- pose initialization
- first keyframe creation
Performs rapid frame-to-frame alignment using image pyramids before full optimization.
Provides
- initial pose estimate
- coarse photometric optimization
- robust tracking
Candidate map points that have not yet converged.
Lifecycle:
Candidate
↓
Tracked
↓
Depth converged
↓
Active Point
Photometric residuals are stored and evaluated efficiently.
Residuals include
- intensity differences
- affine brightness correction
- robust weighting
Builds sparse normal equations used during optimization.
Supports
- block matrices
- efficient updates
- marginalization
Visualization is optional.
When Pangolin is available the project provides
- live trajectory
- camera visualization
- sparse point cloud
- keyframe display
Without Pangolin the library still functions.
OpenCV support is optional.
When available:
- image loading
- image saving
- image display
Otherwise dummy implementations are compiled, making it straightforward to replace OpenCV with another image backend.
Supports calibrated monocular datasets.
Can load
- image sequences
- ZIP-compressed datasets (via libzip)
- calibration files
- photometric calibration files
Designed for datasets such as the TUM MonoVO benchmark.
Uses CMake.
Primary dependencies:
- Eigen3
- SuiteSparse
- Boost
Optional:
- OpenCV
- Pangolin
- libzip
ARM builds are supported through sse2neon, allowing SSE intrinsics to be translated to ARM NEON instructions.
The implementation employs numerous optimization strategies:
- SSE SIMD vectorization
- image pyramids
- sparse Jacobians
- sparse Hessians
- incremental optimization
- sliding window bundle adjustment
- gradient-based pixel selection
- inverse-depth parameterization
- robust loss functions
- efficient memory reuse
The codebase follows a modular architecture:
Camera Images
│
▼
Image Pyramid
│
▼
Coarse Tracking
│
▼
Keyframe Decision
│
▼
Sparse Point Selection
│
▼
Photometric Residual Construction
│
▼
Bundle Adjustment
│
▼
Marginalization
│
▼
Camera Pose Output
- Direct sparse visual odometry
- Monocular camera support
- Photometric bundle adjustment
- Sliding-window optimization
- Sparse inverse-depth map
- Multi-scale tracking
- Adaptive pixel selection
- Robust photometric calibration
- Nonlinear least-squares optimization
- Real-time performance
- SSE acceleration
- ARM NEON compatibility
- Modular visualization backend
- Optional OpenCV dependency
- Optional Pangolin visualization
- Reusable C++ library
The project demonstrates a number of advanced computer vision and optimization techniques valuable to developers:
- Photometric error minimization instead of descriptor matching.
- Inverse-depth parameterization, improving numerical stability for distant scene points.
- Sliding-window marginalization, keeping runtime bounded while preserving information from old keyframes.
- Sparse Hessian assembly using block structures for efficient optimization.
- Coarse-to-fine image alignment with Gaussian pyramids to enlarge the optimization basin.
- Gradient-driven pixel selection, focusing computation on informative image regions.
- Affine brightness modeling, compensating for exposure and illumination changes.
- Highly optimized memory layout and SIMD vectorization (SSE, with ARM support via sse2neon) to achieve real-time performance.
- Monocular Visual Odometry
- Robotics
- Autonomous Navigation
- Drone Localization
- Augmented Reality
- Mobile Robotics
- Camera Motion Tracking
- Research in Direct SLAM
- Academic Computer Vision
- Real-Time State Estimation
DSO employs a generic omnidirectional camera model built around a flexible pixel-to-ray projection framework. Rather than hard-coding a single projection equation (such as the standard pinhole model), the implementation abstracts camera geometry behind a common interface, allowing multiple projection models to be used without modifying the tracking or optimization pipeline.
The camera model is responsible for converting between
- 2D image pixels
- normalized camera rays
- 3D camera coordinates
and provides the Jacobians required by the nonlinear optimizer.
Instead of embedding projection mathematics throughout the codebase, DSO isolates all camera-specific operations inside dedicated classes.
The remainder of the visual odometry pipeline operates only on
- image coordinates
- normalized rays
- inverse depths
- SE(3) camera poses
making the optimization backend independent of the actual lens model.
Image Pixel
│
▼
Camera Model
│
▼
Normalized Ray
│
▼
3D Geometry
│
▼
Optimization
This separation significantly simplifies the addition of new camera models.
The implementation generally supports two projection families.
The standard perspective projection model.
Projection:
u = fx * X / Z + cx
v = fy * Y / Z + cy
Intrinsic parameters:
- fx
- fy
- cx
- cy
Characteristics:
- linear projection
- inexpensive computation
- ideal for narrow field-of-view cameras
- common in robotics datasets
DSO also supports the FOV distortion model introduced by Devernay and Faugeras.
Instead of the classical radial distortion
r' = r (1 + k1 r² + ...)
the FOV model introduces a single distortion parameter
ω
that naturally handles
- wide-angle lenses
- moderate fisheye optics
Advantages:
- only one distortion coefficient
- numerically stable
- invertible
- efficient Jacobians
The code defines an abstract camera interface exposing operations such as
pixel
↓
unproject()
↓
3D ray
3D point
↓
project()
↓
pixel
The optimizer therefore never depends on
- distortion equations
- calibration representation
- projection mathematics
Only the camera model performs these operations.
Every observation follows approximately the same sequence.
Inverse Depth Point
│
▼
3D Camera Coordinates
│
▼
Projection Model
│
▼
Lens Distortion
│
▼
Pixel Coordinates
During optimization the inverse operation is also required.
Pixel
│
▼
Undistortion
│
▼
Normalized Ray
│
▼
Bundle Adjustment
Calibration files typically contain
- image width
- image height
- focal lengths
- principal point
- distortion parameter(s)
Depending on the selected model, the calibration parser interprets the values differently.
For the pinhole model
fx fy cx cy
are sufficient.
For the FOV model an additional
omega
parameter defines lens distortion.
Before optimization, pixels are transformed into normalized camera coordinates.
pixel
↓
remove distortion
↓
subtract principal point
↓
divide by focal length
↓
normalized image plane
This allows the optimization backend to work in a camera-independent coordinate system.
When estimating new map points, DSO converts image pixels into viewing rays.
Pixel
│
▼
Camera Model
│
▼
Unit Ray
│
▼
Inverse Depth
│
▼
3D Point
Inverse-depth parameterization allows points at very large distances to be represented stably.
The nonlinear optimizer requires derivatives of the projection function.
The camera model therefore computes Jacobians such as
- ∂u / ∂X
- ∂v / ∂Y
- ∂projection / ∂pose
- ∂projection / ∂inverse-depth
- ∂projection / ∂camera parameters (where applicable)
These derivatives are used during
- Gauss-Newton optimization
- Hessian construction
- residual linearization
Unlike feature-based systems, DSO minimizes photometric error directly.
For every optimization iteration:
3D Point
│
▼
Projection
│
▼
Pixel Position
│
▼
Image Intensity
│
▼
Photometric Residual
Accurate projection is therefore essential, as even small geometric errors directly affect intensity residuals.
The camera model is compatible with image pyramids.
For each pyramid level:
- focal lengths are scaled
- principal point is adjusted
- projection remains mathematically consistent
This enables coarse-to-fine optimization without requiring separate calibration files for each level.
The abstraction provides several benefits:
- camera-independent optimization
- support for multiple lens models
- reusable projection interface
- straightforward extension to additional camera models
- consistent Jacobian computation
- efficient projection and unprojection
- clean separation between geometry and optimization
Because projection is isolated behind an interface, adding a new lens model typically requires implementing only:
- projection (
project) - inverse projection (
unproject) - distortion and undistortion
- projection Jacobians
- calibration parsing
The remainder of the tracking, mapping, and optimization pipeline can remain unchanged.
Potential extensions include:
- Kannala–Brandt fisheye
- Double Sphere
- Unified Omnidirectional
- Extended Unified Camera Model (EUCM)
- Mei omnidirectional model
- Rational polynomial distortion models
- Abstract camera model interface
- Pinhole perspective projection
- FOV (Field-of-View) distortion model
- Projection and inverse projection support
- Normalized camera ray representation
- Camera-independent optimization backend
- Efficient analytical Jacobians
- Image pyramid compatibility
- Calibration file parsing
- Inverse-depth integration with bundle adjustment
This repository is a production-quality implementation of Direct Sparse Odometry, combining advanced photometric optimization, sparse nonlinear least-squares estimation, and real-time systems engineering. The architecture cleanly separates tracking, optimization, visualization, and I/O, making it valuable both as a research reference and as a reusable C++ library for robotics, SLAM, and computer vision applications. It showcases sophisticated optimization techniques, efficient memory management, modular design, and high-performance numerical computing while remaining extensible through optional visualization and image-processing backends.
For more information see https://vision.in.tum.de/dso
- Direct Sparse Odometry, J. Engel, V. Koltun, D. Cremers, In arXiv:1607.02565, 2016
- A Photometrically Calibrated Benchmark For Monocular Visual Odometry, J. Engel, V. Usenko, D. Cremers, In arXiv:1607.02555, 2016
Get some datasets from https://vision.in.tum.de/mono-dataset .
git clone https://github.com/JakobEngel/dso.git
Required. Install with
sudo apt-get install libsuitesparse-dev libeigen3-dev libboost-all-dev
Used to read / write / display images.
OpenCV is only used in IOWrapper/OpenCV/*. Without OpenCV, respective
dummy functions from IOWrapper/*_dummy.cpp will be compiled into the library, which do nothing.
The main binary will not be created, since it is useless if it can't read the datasets from disk.
Feel free to implement your own version of these functions with your prefered library,
if you want to stay away from OpenCV.
Install with
sudo apt-get install libopencv-dev
Used for 3D visualization & the GUI.
Pangolin is only used in IOWrapper/Pangolin/*. You can compile without Pangolin,
however then there is not going to be any visualization / GUI capability.
Feel free to implement your own version of Output3DWrapper with your preferred library,
and use it instead of PangolinDSOViewer
Install from https://github.com/stevenlovegrove/Pangolin
Used to read datasets with images as .zip, as e.g. in the TUM monoVO dataset. You can compile without this, however then you can only read images directly (i.e., have to unzip the dataset image archives before loading them).
sudo apt-get install zlib1g-dev
cd dso/thirdparty
tar -zxvf libzip-1.1.1.tar.gz
cd libzip-1.1.1/
./configure
make
sudo make install
sudo cp lib/zipconf.h /usr/local/include/zipconf.h # (no idea why that is needed).
After cloning, just run git submodule update --init to include this. It translates Intel-native SSE functions to ARM-native NEON functions during the compilation process.
cd dso
mkdir build
cd build
cmake ..
make -j4
this will compile a library libdso.a, which can be linked from external projects.
It will also build a binary dso_dataset, to run DSO on datasets. However, for this
OpenCV and Pangolin need to be installed.
Run on a dataset from https://vision.in.tum.de/mono-dataset using
bin/dso_dataset \
files=XXXXX/sequence_XX/images.zip \
calib=XXXXX/sequence_XX/camera.txt \
gamma=XXXXX/sequence_XX/pcalib.txt \
vignette=XXXXX/sequence_XX/vignette.png \
preset=0 \
mode=0
See https://github.com/JakobEngel/dso_ros for a minimal example on how the library can be used from another project. It should be straight forward to implement extentions for other camera drivers, to use DSO interactively without ROS.
The format assumed is that of https://vision.in.tum.de/mono-dataset. However, it should be easy to adapt it to your needs, if required. The binary is run with:
-
files=XXXwhere XXX is either a folder or .zip archive containing images. They are sorted alphabetically. for .zip to work, need to comiple with ziplib support. -
gamma=XXXwhere XXX is a gamma calibration file, containing a single row with 256 values, mapping [0..255] to the respective irradiance value, i.e. containing the discretized inverse response function. See TUM monoVO dataset for an example. -
vignette=XXXwhere XXX is a monochrome 16bit or 8bit image containing the vignette as pixelwise attenuation factors. See TUM monoVO dataset for an example. -
calib=XXXwhere XXX is a geometric camera calibration file. See below.
Pinhole fx fy cx cy 0
in_width in_height
"crop" / "full" / "none" / "fx fy cx cy 0"
out_width out_height
FOV fx fy cx cy omega
in_width in_height
"crop" / "full" / "fx fy cx cy 0"
out_width out_height
RadTan fx fy cx cy k1 k2 r1 r2
in_width in_height
"crop" / "full" / "fx fy cx cy 0"
out_width out_height
EquiDistant fx fy cx cy k1 k2 k3 k4
in_width in_height
"crop" / "full" / "fx fy cx cy 0"
out_width out_height
(note: for backwards-compatibility, "Pinhole", "FOV" and "RadTan" can be omitted). See the respective
::distortCoordinates implementation in Undistorter.cpp for the exact corresponding projection function.
Furthermore, it should be straight-forward to implement other camera models.
Explanation:
Across all models fx fy cx cy denotes the focal length / principal point relative to the image width / height,
i.e., DSO computes the camera matrix K as
K(0,0) = width * fx
K(1,1) = height * fy
K(0,2) = width * cx - 0.5
K(1,2) = height * cy - 0.5
For backwards-compatibility, if the given cx and cy are larger than 1, DSO assumes all four parameters to directly be the entries of K,
and ommits the above computation.
That strange "0.5" offset:
Internally, DSO uses the convention that the pixel at integer position (1,1) in the image, i.e. the pixel in the second row and second column,
contains the integral over the continuous image function from (0.5,0.5) to (1.5,1.5), i.e., approximates a "point-sample" of the
continuous image functions at (1.0, 1.0).
In turn, there seems to be no unifying convention across calibration toolboxes whether the pixel at integer position (1,1)
contains the integral over (0.5,0.5) to (1.5,1.5), or the integral over (1,1) to (2,2). The above conversion assumes that
the given calibration in the calibration file uses the latter convention, and thus applies the -0.5 correction.
Note that this also is taken into account when creating the scale-pyramid (see globalCalib.cpp).
Rectification modes:
For image rectification, DSO either supports rectification to a user-defined pinhole model (fx fy cx cy 0),
or has an option to automatically crop the image to the maximal rectangular, well-defined region (crop).
full will preserve the full original field of view and is mainly meant for debugging - it will create black
borders in undefined image regions, which DSO does NOT ignore (i.e., this option will generate additional
outliers along those borders, and corrupt the scale-pyramid).
there are many command line options available, see main_dso_pangolin.cpp. some examples include
-
mode=X:mode=0use iff a photometric calibration exists (e.g. TUM monoVO dataset).mode=1use iff NO photometric calibration exists (e.g. ETH EuRoC MAV dataset).mode=2use iff images are not photometrically distorted (e.g. synthetic datasets).
-
preset=Xpreset=0: default settings (2k pts etc.), not enforcing real-time executionpreset=1: default settings (2k pts etc.), enforcing 1x real-time executionpreset=2: fast settings (800 pts etc.), not enforcing real-time execution. WARNING: overwrites image resolution with 424 x 320.preset=3: fast settings (800 pts etc.), enforcing 5x real-time execution. WARNING: overwrites image resolution with 424 x 320.
-
nolog=1: disable logging of eigenvalues etc. (good for performance) -
reverse=1: play sequence in reverse -
nogui=1: disable gui (good for performance) -
nomt=1: single-threaded execution -
prefetch=1: load into memory & rectify all images before running DSO. -
start=X: start at frame X -
end=X: end at frame X -
speed=X: force execution at X times real-time speed (0 = not enforcing real-time) -
save=1: save lots of images for video creation -
quiet=1: disable most console output (good for performance) -
sampleoutput=1: register a "SampleOutputWrapper", printing some sample output data to the commandline. meant as example.
Some parameters can be reconfigured from the Pangolin GUI at runtime. Feel free to add more.
The easiest way to access the Data (poses, pointclouds, etc.) computed by DSO (in real-time)
is to create your own Output3DWrapper, and add it to the system, i.e., to FullSystem.outputWrapper.
The respective member functions will be called on various occations (e.g., when a new KF is created,
when a new frame is tracked, etc.), exposing the relevant data.
See IOWrapper/Output3DWrapper.h for a description of the different callbacks available,
and some basic notes on where to find which data in the used classes.
See IOWrapper/OutputWrapper/SampleOutputWrapper.h for an example implementation, which just prints
some example data to the commandline (use the options sampleoutput=1 quiet=1 to see the result).
Note that these callbacks block the respective DSO thread, thus expensive computations should not be performed in the callbacks, a better practice is to just copy over / publish / output the data you need.
Per default, dso_dataset writes all keyframe poses to a file result.txt at the end of a sequence,
using the TUM RGB-D / TUM monoVO format ([timestamp x y z qx qy qz qw] of the cameraToWorld transformation).
- the initializer is very slow, and does not work very reliably. Maybe replace by your own way to get an initialization.
- see https://github.com/JakobEngel/dso_ros for a minimal example project on how to use the library with your own input / output procedures.
- see
settings.cppfor a LOT of settings parameters. Most of which you shouldn't touch. setGlobalCalib(...)needs to be called once before anything is initialized, and globally sets the camera intrinsics and video resolution for convenience. probably not the most portable way of doing this though.
-
Please have a look at Chapter 4.3 from the DSO paper, in particular Figure 20 (Geometric Noise). Direct approaches suffer a LOT from bad geometric calibrations: Geometric distortions of 1.5 pixel already reduce the accuracy by factor 10.
-
Do not use a rolling shutter camera, the geometric distortions from a rolling shutter camera are huge. Even for high frame-rates (over 60fps).
-
Note that the reprojection RMSE reported by most calibration tools is the reprojection RMSE on the "training data", i.e., overfitted to the the images you used for calibration. If it is low, that does not imply that your calibration is good, you may just have used insufficient images.
-
try different camera / distortion models, not all lenses can be modelled by all models.
Use a photometric calibration (e.g. using https://github.com/tum-vision/mono_dataset_code ).
DSO cannot do magic: if you rotate the camera too much without translation, it will fail. Since it is a pure visual odometry, it cannot recover by re-localizing, or track through strong rotations by using previously triangulated geometry.... everything that leaves the field of view is marginalized immediately.
If your computer is slow, try to use "fast" settings. Or run DSO on a dataset, without enforcing real-time.
The current initializer is not very good... it is very slow and occasionally fails. Make sure, the initial camera motion is slow and "nice" (i.e., a lot of translation and little rotation) during initialization. Possibly replace by your own initializer.
DSO was developed at the Technical University of Munich and Intel. The open-source version is licensed under the GNU General Public License Version 3 (GPLv3). For commercial purposes, we also offer a professional version, see http://vision.in.tum.de/dso for details.