Skip to content
View HarshSharma0007's full-sized avatar
🎧
Vibe till you die.
🎧
Vibe till you die.

Block or report HarshSharma0007

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
HarshSharma0007/README.md

🚀 AI/ML Engineer & Data Systems Architect

Bachelors in Artificial Intelligence & Machine Learning

LinkedIn


⚡ About Me

I build high-performance data pipelines, predictive models, and agentic AI systems. I specialize in moving beyond tutorials to architect production-grade, local-first applications that solve real engineering constraints—from processing 200M+ event streams on constrained hardware to building multi-track retrieval systems that eliminate LLM hallucinations.

  • 🔭 Currently Building: Enterprise-scale predictive analytics & Agentic RAG architectures
  • 💡 Core Focus: Multi-track RAG, Knowledge Graphs, Out-of-core Data Processing, NLP
  • 🛠️ Recent Obsessions: Polars, DuckDB, LangGraph, KuzuDB, LLM Evaluation (RAGAS)
  • 🎓 Education: B.Tech AIML @ Madhav Institute of Technology & Science (SGPA: 8.72)

🏆 Featured Work

🛡️ AegisLogic (PCRA) | Agentic AI & Cybersecurity

A local-first, multi-track RAG engine for Cyber Threat Intelligence. Dynamically routes queries across a Knowledge Graph (KuzuDB), Semantic Vector Store (Qdrant), and Temporal Incident Timeline (SQLite) using a LangGraph state machine and parallel asyncio fan-out. Achieved 1.0 Faithfulness on RAGAS evaluation while diagnosing a core library metric flaw.

Tech: Python, LangGraph, KuzuDB, Qdrant, Ollama, Llama 3, RAGAS, Docker

🕷️ AraneAI | Big Data & Predictive Analytics

An out-of-core analytics pipeline processing 232M+ eCommerce events (30GB+). Achieved 12× faster ingestion and 80% lower memory usage vs. Pandas using Polars lazy evaluation and Apache Parquet. Features a 20+ feature RFM store built on DuckDB to train an XGBoost churn model (AUC: 0.90), served via an async FastAPI backend with NL-to-SQL capabilities.

Tech: Python, Polars, DuckDB, XGBoost, FastAPI, Streamlit, Docker Compose


🛠️ Tech Stack & Arsenal

Languages:
Python C++ SQL Java R

AI / LLMs & Evaluation:
RAG LangChain LangGraph Prompt Engineering Ollama Groq API RAGAS MLflow

ML Frameworks:
TensorFlow Keras PyTorch Scikit-Learn XGBoost

Databases & Data Engineering:
Polars Pandas NumPy Apache Parquet PostgreSQL DuckDB Qdrant KuzuDB SQLite

DevOps, Tools & Soft Skills:
Docker Docker Compose FastAPI Streamlit Git Plotly Jupyter Notebook


Let's build something scalable. Feel free to reach out via LinkedIn.

Pinned Loading

  1. AraneAI AraneAI Public

    AraneAI: Enterprise Customer Intelligence & Predictive Analytics Platform processing 232M+ eCommerce events (30GB+) via out-of-core Polars lazy evaluation. Features a DuckDB star-schema warehouse, …

    Jupyter Notebook

  2. AegisLogic AegisLogic Public

    AegisLogic (PCRA): A novel, local-first retrieval architecture (PCRA) for zero-hallucination Cyber Threat Intelligence. Synthesizes Graph (KuzuDB), Vector (Qdrant), and Temporal data via a LangGrap…

    Python

  3. Hostel Hostel Public

    A secure, dynamic, and ticket-based Hostel Complaint Management System (HCMS). Features role-based dashboards, Google OAuth, dynamic forms, and ghost-routing security. Built with Django and Postgre…

    HTML

  4. Ragapplication Ragapplication Public

    Local Retrieval-Augmented Generation (RAG) pipeline for document QA. Built with LangChain, Llama2 (via Ollama), and an in-memory DocArray vector store, featuring an interactive Streamlit interface …

    Jupyter Notebook

  5. ml-internship ml-internship Public

    End-to-end Machine Learning internship projects covering Computer Vision, Clinical Prediction, and Financial Fraud Detection. Features modular pipelines, explainability (SHAP/Captum), interactive S…

    Jupyter Notebook

  6. CodeAlpha_Data_Science_Portfolio CodeAlpha_Data_Science_Portfolio Public

    A collection of full-stack Machine Learning applications built during the CodeAlpha Data Science Internship. Features predictive models with FastAPI backends and interactive React dashboards.

    Jupyter Notebook