a lightweight, comprehensive solution for managing delta tables built on polars and deltalake
-
Updated
Jan 1, 2025 - Python
a lightweight, comprehensive solution for managing delta tables built on polars and deltalake
An open-source Python library for simplifying local testing of Databricks workflows that use PySpark and Delta tables.
A comprehensive ETL pipeline and sales analysis project leveraging Microsoft Azure and PySpark, designed to optimize e-commerce sales by providing actionable insights through detailed data analysis.
This is a prototype of a big-data system about cultural heritage data and metadata. We ingest, process and deduplicate cultural objects from Europeana, and then we make recommendations of similar content using CLIP
End-to-end E-Commerce Lakehouse built on Databricks using Medallion Architecture (Bronze → Silver → Gold). Ingests raw CRM and ERP data, applies PySpark transformations, and delivers a star schema via Spark SQL stored as Delta Tables in Unity Catalog.
Production-grade Azure Databricks ETL pipeline processing retail data into a Star Schema. Demonstrates idempotent streaming, advanced PySpark OOP transformations, and declarative data pipelines via Delta Live Tables.
This project is a real-time fleet analytics pipeline on Azure. It ingests delivery and telemetry data via Event Hubs, processes them with Stream Analytics, and stores results in Data Lake (Bronze, Silver, Gold). Synapse and Power BI provide dashboards for KPIs.
A modular data-driven framework leveraging Spark, Delta Lake, MLflow, Airflow, and Streamlit to build and orchestrate a full-stack data lake solution. Designed for scalable ETL, automated ML pipelines, and interactive dashboarding.
spark, databricks, kafka, batch and stream-processing
On-premise data lake architecture with Trino, Delta Tables and Hive Metastore
Azure data fundamentals portfolio with 3 labs on relational databases, non-relational storage, and large-scale analytics. Hands-on experience with SQL, PostgreSQL, MySQL, Azure Storage, Synapse Analytics, Spark, Delta tables, and KQL for batch and streaming data.
This repo contains details about Databricks Ecomm Event Driven Pipeline project execution, Thanks
End-to-end data engineering demo built on Databricks
Implementing Change Data Capture for Seamless Fintech Data Migration
This project builds a cloud-based pipeline to extract NYC taxi data from an API and store it in Azure Data Lake Storage (ADLS). Databricks and PySpark are used to transform the data through the medallion architecture (Bronze → Silver → Gold). Delta Lake ensures reliable storage, and Power BI provides visual insights for data-driven decision-making.
Databricks & Blueprint Hackathon - using databricks, spark structured streaming, delta, and azure devops to build automated deployment of notebooks and jobs.
Add a description, image, and links to the delta-tables topic page so that developers can more easily learn about it.
To associate your repository with the delta-tables topic, visit your repo's landing page and select "manage topics."