Paleobiologist, Professor of Biology & Scientific Data Engineer
I am a Paleobiologist and Professor of Biology who builds deployed AI/ML systems and production-ready data pipelines optimized for massive scientific databases. Backed by a PhD in geological sciences (micropaleontology) and two decades of hands-on domain experience with deep-time geological, paleontological, and spatial data, I bridge the gap between highly complex natural-science datasets and robust, verifiable computational engineering.
- Unified PostgreSQL/PostGIS Research Database β Integrated global geology, stratigraphy, paleobiology, and a 25GB biology dataset on a single PostgreSQL/PostGIS instance. Migrated via pgloader with row-count verification (which caught a silent schema-target failure), custom non-standard projection resolved to a user-defined SRID, and strict model-vs-observation provenance quarantine. The geospatial layers are engineered to act as a common spatial key across domains. π View Project
- Global Geology PostGIS Analysis β A PostgreSQL/PostGIS spatial database integrating USGS World Geology and State Geologic Map Compilation datasets (148.3M kmΒ² of mapped geology), using spatial SQL to compute continental-scale geologic surface exposure by period. π View Project
- Geostatistical Depositional-Environment Pipeline β An uncertainty-quantified kriging pipeline (reliability-weighted, declustered ordinary kriging via PyKrige / scikit-gstat) that estimates marine-vs-terrestrial depositional probability from integrated fossil-occurrence and stratigraphic data. Its uncertainty layer desaturates where data is sparse, making confidence visible rather than assumed. Reproducible method demonstration; full analysis in preparation for peer review. π View Project
- Spatial Biodiversity Gap Audit Platform β An automated ETL pipeline and live analytical dashboard cross-auditing global IUCN Red List conservation data against a 26-million-record GBIF occurrence dataset. π Launch Deployed App
- Semantic Socratic Tutor β A retrieval-augmented generation (RAG) system grounded in a specialized scientific corpus, using all-MiniLM-L6-v2 embeddings, NumPy cosine-similarity retrieval, and the Anthropic Claude API. π Launch Deployed App
- Languages & Databases: Python, SQL, PostgreSQL, PostGIS, SQLite, R
- Geospatial & Geostatistics: Spatial SQL, PostGIS, GeoPandas, kriging (PyKrige, scikit-gstat), uncertainty quantification, spatial indexing, equal-area projection, custom SRID/projection handling
- Data Engineering: ETL pipelines (pgloader), database normalization and migration, large-scale ingestion, row-count verification, model-vs-observation provenance discipline
- AI/ML & NLP: Retrieval-Augmented Generation (RAG), vector embeddings, semantic search, LLM API integration
- Deployment & UI: Streamlit, Streamlit Cloud, interactive dashboards
Alongside my computational and scientific work, I maintain a platform exploring the intersection of the historical sciences, philosophy, and faith at Faith and Reason-25.
Profiles: LinkedIn | ORCID | Google Scholar | ResearchGate