A deep learning project that predicts GPS coordinates from images using computer vision and geographic reasoning, trained on the CVL_SFBay-1.5K Test Split dataset. The contents within this repo are still WIP. You can wait for the main publication version which is coordinated for May of 2026 to November.
This project implements multiple neural network architectures to predict GPS coordinates from images, with a focus on the San Francisco Bay Area. The models combine visual feature extraction with geographic reasoning to estimate location coordinates and understand terrain/scene types.
CVL 1.1 - Nano (Lower than 10.0M) - Awaiting Release
CVL 2.0 - Base (19.0M Params) - Awaiting Release
CVL 2.0 - High (24.4M Params) - Main2_High_Improved.py
CVL 2.0 - Ultra (18.0M Params) - Awaiting Release
CVL 2.0 - Global Ultra - main2_ultra.py
CVL 3.0 - Dynamic - Awaiting Release
CVL 5.0 - Cluster Global - mainv5.py
You can download the models from: https://huggingface.co/Celsia/CvelsialArch_V1 or git clone https://huggingface.co/Celsia/CvelsialArch_V1
CVL_SFBay-1.5K Test Split
- 1,500+ images from the San Francisco Bay Area
- GPS coordinates: Latitude [37.301, 37.900], Longitude [-122.478, -121.902]
- Includes diverse terrain types: urban, water, forest, grassland, mountain, beach
- EXIF metadata with precise GPS coordinates for training
Models achieve meter-level accuracy on the CVL_SFBay-1.5K dataset:
- Typical errors: 800-2800 meters depending on terrain complexity
- Best performance on urban areas with distinct visual landmarks
- Challenges: Water areas, dense forests, similar-looking residential areas
Setup with pip install -r r.txt
If you have photos taken on any mobile device that contains EXIF you may do this:
python exif_extract.py
This code will automatically save the files to data/exif_data.json so that we can use it later for computations.
python main2_high_improved.py # Train high-end modelpython -c "
from main2_high_improved import predict_improved_high_end, ImprovedHighEndGPSNet
import torch
# Load trained model
model = ImprovedHighEndGPSNet()
checkpoint = torch.load('improved_high_end_gps_model.pth')
model.load_state_dict(checkpoint['model_state_dict'])
# Predict location
lat, lon, terrain, terrain_conf, coord_conf = predict_improved_high_end(model, 'path/to/image.jpg')
print(f'Predicted location: {lat:.6f}, {lon:.6f}')
print(f'Terrain: {terrain} (confidence: {terrain_conf:.3f})')
"streamlit run streamlit_app_high_improved.py├── main2_high_improved.py # High-end GPS model (29.5M params)
├── main2_ultra.py # Ultra GPS model (16.6M params)
├── streamlit_app_[version of the booter eg. high_improved].py # Web interface
├── data/
│ ├── images/ # Training images
│ ├── exif_data.json # GPS metadata
│ └── location_data.csv # Location dataset
├── *.pth # Trained model weights
└── README.md # This file
torch
torchvision
PIL
folium
numpy
streamlit
requests
tqdm
Trained on the CVL_SFBay-1.5K Test Split dataset, produced by Celsia Juilyn Fan for academic research purposes.
@misc{fan2025,
author = {Fan, Cheng-Jui},
title = {CvelsialArch for Computer Vision Locators},
year = {2025},
month = {Jul},
day = {17}
}