Enterprise-grade, production-hardened, serverless data lake on AWS
-
Updated
Oct 1, 2025 - Python
Enterprise-grade, production-hardened, serverless data lake on AWS
📄 Pipeline de ingestão e análise com AWS Glue, Athena e S3, utilizando arquivos Parquet e catalogação automatizada via Crawler. Permite consultas SQL rápidas e eficientes sobre dados armazenados na nuvem.
Hands-on AWS data lake and pipeline workshops for engineering teams. Glue, Redshift, Lake Formation, Databricks, LLMOps.
📄 Projeto educacional de pipeline de ingestão de dados utilizando AWS S3 e Lake Formation, estruturando a camada bronze de um Data Lake com Python.
A HIPAA-compliant healthcare data lake built on AWS using Medallion Architecture (Bronze/Silver/Gold). Features PySpark ETL pipelines via AWS Glue, column-level security through Lake Formation, serverless SQL with Athena, and a live Tableau dashboard — simulating a real hospital's cloud migration with full governance and monitoring.
Event-driven serverless lakehouse on AWS (Lambda x6 patterns, Step Functions, Glue PySpark, Iceberg, Lake Formation, Superset) — runs end-to-end on LocalStack at zero cost. Terraform IaC + CDKTF parity.
Kinesis Firehose → S3 → Glue → Lake Formation → Athena → QuickSight data platform
Production-grade financial data lake on AWS — S3, Glue, Lake Formation, Athena, DMS. Built for BFSI analytics use case with PII masking and column-level access control.
8 hands-on AWS Certified Data Engineer Associate labs covering S3 data lakes, ingestion, Glue, Athena, ETL, data quality, Redshift, DynamoDB, orchestration, Lake Formation, KMS, monitoring, cost optimization, and exam review.
Production-style AWS data lakehouse — Terraform, EMR Serverless, Apache Iceberg, and Lake Formation with idempotent Spark writes and governed access.
Add a description, image, and links to the lake-formation topic page so that developers can more easily learn about it.
To associate your repository with the lake-formation topic, visit your repo's landing page and select "manage topics."