Data & AI Engineering

Designing Resilient Pipelines
Engineering AI Insights

> |

I am a Master of Science in Information Technology graduate from Arizona State University (GPA: 4.0/4.0), specializing in scalable ETL/ELT pipelines, real-time IoT telematics, machine learning models, and automated data observability frameworks using Python, SQL, and advanced business intelligence analytics.

Explore Pipelines
Deshraj Jogiya

Deshraj Jogiya

Data & AI/ML Engineer

11

Production Pipelines

2300

Git Commits

98

ETL Reliability %

4

ASU MS IT GPA

Professional Experience & Technical Impact

3+ years of designing production-grade data pipelines, optimizing LLM/AI models, and building high-reliability cloud data architectures.

Robotics & Parallel Pipelines

Teleoperation Data Collection Associate

Objectways Technologies LLC | Tempe, AZ

May 2026 – Present Β· Part-Time
  • Collected and validated 10,000+ high-quality teleoperation sensor data samples for AI/ML model training, improving dataset accuracy and consistency by 20%.
  • Engineered scalable data pipelines using Python, Scala, PySpark, and Kubernetes to process massive robotics datasets, cutting data processing time by 30%.
  • Developed responsive web dashboards using React JS, JavaScript, and Java Spring Boot to monitor 10,000+ concurrent telemetry streams, improving system latency visualization by 30%.
  • Executed robotics dataset QA and statistical anomaly detection, lowering downstream ML model training error rates by 15%.
  • Leveraged AI-powered developer tools (Cursor, GitHub Copilot) to refactor backend microservices, write automated unit/integration tests, and maintain technical documentation.
Python PySpark Scala Kubernetes Docker React JS Spring Boot Robotics Data QA Linux
LLMs, Supabase & Vector Sync

Applied Machine Learning Engineer

Technoid LLC | Piscataway, NJ

Dec 2025 – May 2026 Β· Contract
  • Optimized GPT-4o mini models for candidate resume matching using OpenAI APIs, SQL, and PostgreSQL, raising recommendation accuracy by 25%.
  • Restructured Supabase data sync layers with PostgreSQL backend RLS fixes, cutting real-time vector synchronization latency by 65%.
  • Architected backend microservices and RESTful APIs using Core Java, Spring Boot, and Python endpoints, reducing data sync latency by 65%.
  • Built dynamic, accessible UI components in React JS connected to Spring Boot backend services, conducting peer code reviews to ensure Object-Oriented design compliance.
  • Established automated phased regression and UAT testing frameworks using PyTest, cutting model deployment errors across AWS and Azure cloud environments by 30%.
Python Core Java Spring Boot React JS GPT-4o mini OpenAI API PostgreSQL Supabase PyTest AWS/Azure
Cloud Migration & Data Quality

Data Analyst

Zifatech Solutions LLC | Milwaukee, WI

Jun 2025 – Dec 2025 Β· Contract
  • Transformed SQL/Python-based ETL pipelines to automate sales insights reporting, cutting manual reporting effort by 70%.
  • Migrated legacy database workflows to AWS Glue & S3, increasing data availability by 60% and streamlining integration with Snowflake and Power BI.
  • Modernized Star Schema dimensional models in Snowflake using Great Expectations QA validations, securing a 98%+ data reliability standard.
  • Built executive Power BI dashboards connected to Snowflake data warehouses, converting complex data streams into strategic business insights.
  • Engineered serverless RESTful API data pipelines on AWS (Glue, S3, Lambda) with Apache Airflow orchestration to aggregate multi-source application data, accelerating ingestion speed by 60%.
AWS Glue AWS S3 AWS Lambda Apache Airflow Snowflake Python SQL Great Expectations Power BI
Streaming & Recommendation Systems

Data Engineer & ML Research Assistant

Arizona State University | Tempe, AZ

Sep 2024 – Jun 2025
  • Architected a real-time streaming pipeline utilizing Node.js & MongoDB, sustaining 99.9% uptime for 5,000 concurrent users and optimizing data throughput by 35%.
  • Constructed a behavioral content recommendation engine, refining algorithms weekly to drive a 12% increase in click-through rates and lower user churn.
  • Developed an NLP chatbot analyzing user input trends, identifying top UI friction points and boosting positive user satisfaction by 15%.
  • Integrated transformer NLP classifiers (BERT) into web applications, boosting user sentiment feature extraction accuracy by 15%.
  • Centralized application telemetry into SQL databases, presenting architecture designs, schema maps, and visual decks directly to leadership.
Node.js MongoDB Python BERT / NLP Recommendation Engine Streaming ETL SQL
Robotics Pipeline Automation

AI-ML Analyst Apprentice

Jetson Infinity | Austin, TX

Jul 2024 – Aug 2024
  • Implemented Python-based data processing pipelines leveraging Pandas & NumPy (+3% robotic arm precision, +10% workflow efficiency).
  • Automated robotic arm motion ETL pipeline with Python & SQL, improving data processing speed by 17% and enabling real-time anomaly detection.
  • Developed modular Python, C#, and .NET data transformation scripts, improving application dataset cleaning efficiency by 17%.
  • Leveraged AI-powered developer tools (GitHub Copilot, Cursor) to write, test, debug, and refactor production-grade microservices and transformation pipelines.
  • Devised an internal knowledge base with Python tutorials and real-world case studies, accelerating new analyst onboarding by 25% across Unix/Linux environments.
Python Pandas NumPy SQL C# / .NET Cursor / Copilot Linux
Subledger Accounting & BI

Data Analyst

Kronic Keys | Junagadh, GJ

Aug 2021 – Mar 2022
  • Cleaned and prepared large financial datasets (500k+ rows) in PostgreSQL, reducing monthly reporting cycle times by 30%.
  • Engineered SQL queries and Tableau executive dashboards, contributing to a 12% increase in customer acquisition.
  • Conducted weekly data integrity audits across relational database tables, eliminating duplicate records and ensuring 100% data quality control.
  • Designed relational database schemas and indexed foreign keys in PostgreSQL to accelerate analytics query performance.
  • Documented data definitions, data dictionary assets, and reporting workflows for business stakeholders.
PostgreSQL SQL Tableau ETL Ingestion Data Modeling

Relentless Execution: Portfolio Repositories

Filter through active production-ready implementations built with end-to-end automation, modeling, and custom dashboards.

Job Search CRM & AI Application Tailoring Center Dashboard
Zoom Dashboard πŸ”
AI & Process Automation

Job Search CRM & AI Application Tailoring Center

Local job placement command center with FastAPI and SQLite. Features AI-powered profile matching, resume and cover letter tailoring, LinkedIn recruiter outreach generation, and Kanban application tracking.

FastAPI Python SQLite OpenAI API Jinja2 TailwindCSS GitHub Actions Docker
CurioSync: Serverless Tech News & LinkedIn Publisher Dashboard
Zoom Dashboard πŸ”
Cloud & Generative AI

CurioSync: Serverless Tech News & LinkedIn Publisher

Scheduled serverless news curation and publisher pipeline running on GitHub Actions cron workflows. Integrates Gemini LLM chains and LinkedIn APIs, secured by Fernet symmetric encryption.

FastAPI Python GitHub Actions Gemini LLM HTMX Cryptography
Member Messages QA System: Semantic Search & NER-Based Question Answering Dashboard
Zoom Dashboard πŸ”
NLP & Semantic Search

Member Messages QA System: Semantic Search & NER-Based Question Answering

Interactive QA API answering natural-language questions about member data, ranking unstructured messages via SentenceTransformer embeddings and identifying subjects with spaCy NER.

Python Sentence-Transformers spaCy Scikit-Learn Pandas GitHub Actions LlamaIndex Semantic Kernel LangChain
TalentVenue EventIntel: Enterprise Data Platform & Predictive Intelligence Dashboard
Zoom Dashboard πŸ”
Cloud & Enterprise ML

TalentVenue EventIntel: Enterprise Data Platform & Predictive Intelligence

Enterprise cloud data platform connecting SQL Server, Azure ADLS Gen2, Snowflake, and Streamlit. Deployed a Random Forest model predicting contract cancellation risk with 86.20% accuracy, alongside automated financial reporting and reconciliation workflows.

Python Snowflake Azure ADLS Streamlit Scikit-Learn SQL Server Terraform Parquet Delta Lake Apache Iceberg TypeScript Jenkins
Reactive Systems

Reactive Job Ingest Gateway

A non-blocking job-posting ingestion gateway built on Spring WebFlux and R2DBC, republishing ingested postings as a live Server-Sent-Events stream.

Java Spring Boot Spring WebFlux R2DBC Project Reactor
Data Warehousing

Job Search dbt Warehouse

A dimensional dbt project (staging views into dimension and fact tables) running against a live Databricks SQL warehouse, with real data-quality tests in CI.

Databricks dbt SQL Dimensional Modeling
Big Data & Orchestration

Job Market Lakehouse

A PySpark/Scala Spark SQL pipeline over 60k synthetic postings, gated by real Great Expectations checks and orchestrated end-to-end by a real Airflow DAG.

PySpark Spark SQL Scala Great Expectations Apache Airflow
Full Stack

Job Notes

A company-research notes app: Express + MongoDB backend, Angular frontend, both with real automated tests.

MongoDB Express Angular Node.js Mongoose
Full Stack

Interview Prep API

A real ASP.NET Core minimal API (C#) for interview-prep questions, backed by Entity Framework Core + SQLite.

C# .NET ASP.NET Core Entity Framework Core SQLite
Full Stack

Company Blocklist

A real PHP/CodeIgniter 4 app tracking companies to avoid in a job search, with real automated feature tests.

PHP CodeIgniter 4 SQLite
Full Stack

Interview Prep Mobile

A real React Native (Expo) app for browsing/practicing interview-prep questions against a companion ASP.NET Core API.

React Native Expo TypeScript AsyncStorage
Full Stack

Company Search (HTMX)

A live-search page using HTMX for partial-page updates -- FastAPI + Jinja2, no client-side JS framework.

HTMX FastAPI Jinja2
Geospatial

ArcGIS Commute Insights

A job-search tool geocoding addresses and pulling real ArcGIS basemap vector-tile data via the live ArcGIS APIs.

ArcGIS Geocoding Vector Tiles Geospatial Analysis
AI Model Observability & Fairness Audits Dashboard
Zoom Dashboard πŸ”
AI Governance

AI Model Observability & Fairness Audits

Audits machine learning models in production by running Kolmogorov-Smirnov (KS) tests to track input data drift and evaluating Disparate Impact for bias protection.

Python Drift Audit KS-Test PSI Tableau SQL GitHub Actions
FinTech Credit Risk & Fraud Command Center Dashboard
Zoom Dashboard πŸ”
FinTech & Risk

FinTech Credit Risk & Fraud Command Center

End-to-end transaction ingestion and credit evaluation pipeline. Integrates a Random Forest classifier predicting loan eligibility with real-time transactional fraud checks.

FastAPI Python SQL Scikit-Learn Chart.js XGBoost GitHub Actions MLflow GraphQL
Sentence Transformers, Multi-Task Learning & LLM Fine-Tuning Dashboard
Zoom Dashboard πŸ”
NLP & Transfer Learning

Sentence Transformers, Multi-Task Learning & LLM Fine-Tuning

Built a Sentence Transformer from a Small BERT backbone with TensorFlow Hub, extended it into a multi-task architecture with classification and sentiment heads, and later added real LoRA fine-tuning of a modern generative LLM on a free GPU.

TensorFlow TensorFlow Hub BERT Multi-Task Learning PyTorch LoRA Hugging Face NumPy GitHub Actions
Multi-State Land Use Emissions Analysis Dashboard
Zoom Dashboard πŸ”
Geospatial & Climate

Multi-State Land Use Emissions Analysis

Geospatial emissions intelligence pipeline processing daily land cover changes across 5 U.S. states. Computes COβ‚‚ inventory and forecasts emissions with ~90% accuracy using Linear Regression and Random Forest.

Python SQLite Random Forest ETL Ingestion ArcGIS GitHub Actions
Tax Anomaly Audit Compliance Engine Dashboard
Zoom Dashboard πŸ”
Finance & Auditing

Tax Anomaly Audit Compliance Engine

Relational general ledger audit pipeline modeling transaction values following Benford's Law distribution and executing Isolation Forest anomaly flags for audit risk.

Benford's Law Isolation Forest Python Power BI SQL GitHub Actions Dagster
Automated Daily Data Insights Dashboard
Zoom Dashboard πŸ”
Automation & ETL

Automated Daily Data Insights

Stateless serverless daily financial data ingestion pipeline calling yFinance, computing rolling Z-score anomalies, and committing reports automatically via GitHub Actions.

Python yFinance Matplotlib GitHub Actions SQLite
Sales Operations & Customer Segmentation Dashboard
Zoom Dashboard πŸ”
Sales & Marketing

Sales Operations & Customer Segmentation

Star schema database pipeline modeling Pareto 80/20 sales distributions. Automates RFM scaling and K-Means clustering to classify customers into 4 target personas.

Python K-Means PCA SQLite Power BI DAX Seaborn GitHub Actions
Clinical Trials Patient Outcomes Analysis Dashboard
Zoom Dashboard πŸ”
Healthcare & Statistics

Clinical Trials Patient Outcomes Analysis

Clinical database pipeline tracking patient survival and blood pressure across 3 treatment arms. Computes Log-Rank test statistics and Kaplan-Meier curves.

Python Kaplan-Meier Log-Rank Tableau scipy GitHub Actions
Real-Time IoT Telematics & Predictive Maintenance Dashboard
Zoom Dashboard πŸ”
Streaming & IoT

Real-Time IoT Telematics & Predictive Maintenance

Processes high-frequency EV fleet battery and motor sensor telemetry in streaming micro-batches, detecting Z-score anomalies and estimating Remaining Useful Life (RUL).

Z-Score Anomaly RUL Model Python Power BI SQL GitHub Actions Jax
AI-ML Data Science Simulation Dashboard
Zoom Dashboard πŸ”
Operations & Demand

AI-ML Data Science Simulation

Retail data pipeline automating daily sales ingestion from 5 U.S. branches. Centralizes operations in SQL and uses Linear Regression/Random Forest forecasting to reduce stockouts by 15%.

Python SQL Random Forest Sales ETL Tableau GitHub Actions
Extending STEM across ASL Dashboard
Zoom Dashboard πŸ”
Accessibility & Deep Learning

Extending STEM across ASL

Inclusive educational learning platform utilizing a Convolutional Neural Network (CNN) in TensorFlow/Keras to recognize American Sign Language gestures for 7 core STEM concepts.

Python TensorFlow Keras CNN Model Flask ASL Recognition GitHub Actions
Chef at Gathering Dashboard
Zoom Dashboard πŸ”
Web Development & Catering

Chef at Gathering

Collaborated on an event coordination and professional catering booking React application, contributing to a 50% reduction in coordination time.

React JavaScript HTML5 CSS3 Git
Stay Aware of Branch Dashboard
Zoom Dashboard πŸ”
Mobile & Education

Stay Aware of Branch

Collaborated on a parent-school engagement application using React Native and Android frameworks, contributing to progress-report and attendance-tracking features that boosted parent involvement by 40%.

React Native Android JavaScript Mobile Git
Get your Token Dashboard
Zoom Dashboard πŸ”
FinTech & Blockchain

Get your Token

Collaborated on a Web3 crypto marketplace frontend and administrative dashboard using Express.js and Angular for NFT and Penky token transactions, contributing to a system that handled 1,000+ transactions in the first month.

Node.js Express Angular Cryptocurrency API Git
City Forums Dashboard
Zoom Dashboard πŸ”
Web Development & Community

City Forums

Collaborated on an online community engagement and marketplace platform using CodeIgniter 3 and MySQL, contributing to $10,000 in classified ad revenue through targeted marketing strategies.

PHP CodeIgniter 3 MySQL Bootstrap Git
Make It Short Dashboard
Zoom Dashboard πŸ”
Web Services & Gamification

Make It Short

Collaborated on integrating a link shortening engine and analytics tracking into a CodeIgniter 4 app, contributing to 10,000+ shortened URLs generated through gamified, redeemable points.

PHP CodeIgniter 4 MySQL HTML5 CSS3 Git
Solid Object Detection & Identification using Image Processing Dashboard
Zoom Dashboard πŸ”
Computer Vision & Deep Learning

Solid Object Detection & Identification using Image Processing

Developed a shape detection and identification pipeline using PyTorch and OpenCV, achieving 98.97% classification accuracy by combining a custom CNN with traditional contour geometry approximations.

Python PyTorch OpenCV CNN Model Image Processing Git GitHub Actions AWS Lambda AWS S3 ONNX Runtime

Interactive Engineering Demos

Test and monitor pipeline diagnostics or execute analytical SQL queries directly in-browser to see how data decisions are made.

πŸ“Š SQL Analytics Playground

πŸ’‘ Project Value & Key Takeaways
Select a query above to see the business value of this data engineering task and the core skills it demonstrates.
Key Takeaway: -
πŸ“‹ Query Results Output
Click 'Execute Query' to fetch database records...

πŸ€– Live ETL Pipeline Observability Monitor

System Idle
πŸ’‘ What is this? This widget simulates a real-time data engineering pipeline monitoring system. In production, data pipelines process millions of rows daily and can fail silently due to bad data. Click "Run Pipeline Diagnostics" to run a live validation check. You'll see how data ingestion, quality checks, machine learning scoring, and database commits flow sequentially, demonstrating automated pipeline monitoring and observability skills.
πŸ“₯
Ingest
πŸ›‘οΈ
Validate
🧠
ML Score
πŸ’Ύ
Commit
bash - live_etl_pipeline.sh
[SYSTEM] Click 'Run Pipeline Diagnostics' to run mock pipeline validations...
0.00s
ETL Latency
98.2%
Reliability (30d)
59/60 Daily Runs Successful
0
Processed Rows
Stable
Drift Status

πŸ§ͺ A/B Testing & Hypothesis Testing Simulator

Compute real-time Z-scores & p-values for experimental conversions.

Control (A)
0.00% conv. rate
Conversions: 0 / 1000
Challenger (B)
0.00% conv. rate
Conversions: 0 / 1000
Uplift (Relative): 0.00%
Z-Score: 0.00
Two-tailed P-value: 0.0000
Status: Not Simulated
πŸ“Š Methodology & Business Context

The Problem: When launching new features or ML models, we must prove they actually improve user conversion rate (CR) rather than just being random noise.

How to Use: Drag the target sliders to set conversion rates for A and B. Click "Run Experiment Simulation" to simulate 1,000 visitors per group. The engine runs a live two-tailed Z-test on the proportions.

Recruiting Value: This simulates how a Data Engineer or Scientist conducts hypothesis testing to validate feature rolls. A result is statistically significant if the P-value is less than 0.05 (95% confidence).

πŸŽ›οΈ ML Model Threshold & User Conversion Optimizer

Adjust decision boundary to balance user friction and maximize conversion ROI.

Confusion Matrix
Pred Conversion Pred Passive
Act Conversion 0 0
Act Passive 0 0
Precision
0.0%
Recall
0.0%
Engaged Converter (+$100/TP): +$0
User Friction Cost (-$20/FP): -$0
Missed Conversions (-$150/FN): -$0
Net Conversion ROI Impact: $0
πŸ“Š Methodology & Business Context

The Problem: Machine Learning classifiers don't just output labels; they output probabilities. Choosing a threshold of 0.5 is rarely optimal when the business costs of errors are asymmetric.

How to Use: Drag the threshold slider. Observe how lower thresholds catch more converters (higher Recall) but increase false alarms (lower Precision, costing user friction). Watch the confusion matrix adjust in real time.

ROI Business Value: By modeling the costs (+$100 per conversion, -$20 per false alarm friction, -$150 per missed conversion), you can find the mathematical point that maximizes net returns, which occurs at the 0.35 threshold.

πŸ—ΊοΈ Data Pipeline Architecture Explorer

Interact with real-world pipeline system designs to see data flow, tool stacks, and Python execution modules.

Select a pipeline block above...

Click on any stage in the flow diagram above to view its technical specifications, tool stack, and Python code implementation.

Core Technical Expertise

Proficient across modern data engineering tools, machine learning modeling, database modeling, and visual analytics platforms.

Generative AI, LLMs & Agentic Systems

  • LangChain, LlamaIndex & Semantic Kernel
  • MCP (Model Context Protocol) & Tool Calling
  • ReAct Agent Workflows & Autonomous Agents
  • Retrieval-Augmented Generation (RAG) & Fine-Tuning (LoRA/PEFT)
  • Vector Databases (pgvector, ChromaDB, Supabase RLS)
  • OpenAI APIs (GPT-4o mini) & Gemini LLM Integration
  • LLMOps & Agent Evaluation (LangSmith, Opik, Langfuse)

Machine Learning & Deep Learning Frameworks

  • PyTorch, Jax, TensorFlow & Keras Deep Learning
  • Scikit-Learn Modeling (Random Forest, XGBoost, K-Means)
  • Computer Vision & Pattern Recognition (OpenCV, YOLO CNNs)
  • Statistical Drift Auditing (KS-Test, PSI) & MLflow MLOps
  • Responsible AI Governance (FTC & EU AI Act Bias Auditing)
  • Financial & Audit Anomaly Detection (Benford's Law, Isolation Forest)
  • Hypothesis Testing & Z-Test A/B Experimentation

Data Engineering, Platforms & Cloud

  • PySpark & Spark SQL Large-Scale Data Processing
  • Databricks Platform & dbt Dimensional Transformations
  • AWS Cloud Infrastructure (S3, Glue, Lambda, EC2)
  • GCP (Google Cloud Platform, BigQuery, Vertex AI) & Azure (ADLS Gen2)
  • Snowflake Cloud Data Warehousing & Star Schema Modeling
  • Apache Airflow & Dagster Pipeline Orchestration
  • Great Expectations Data Quality Rules & Governance
  • Table Formats (Delta Lake, Apache Iceberg, Parquet)

Full Stack & Software Engineering

  • Python Backend Microservices (FastAPI, Flask, Asyncio)
  • Core Java, Spring Boot & Spring Reactive Engineering
  • C# & .NET Framework Backend Services
  • JavaScript, TypeScript & Node.js/Express APIs
  • React JS, React Native Mobile & Angular Web Interfaces
  • PHP & CodeIgniter 3/4 Application Platforms
  • RESTful APIs, GraphQL & Microservices Architecture
  • Data Structures, Algorithms & Object-Oriented System Design

DevOps, Tooling & Analytics

  • Docker Containerization & Kubernetes Cluster Scaling
  • Git & GitHub Actions CI/CD Automated Pipelines
  • Terraform Infrastructure as Code (IaC) & Jenkins
  • AI Developer Productivity Tools (Cursor, Claude Code, GitHub Copilot)
  • Power BI (DAX Modeling, KPI Dashboards)
  • Tableau (LOD Expressions, Executive Storytelling)
  • ArcGIS Geospatial Emissions Analytics & Chart.js Visualizations
  • Relational Databases (PostgreSQL, SQLite, SQL Server, MySQL)

Live Learning Logs (TIL)

Real-time learning feed pulled dynamically from my `dev-root-affinity` learning log repository via the GitHub API, showcasing daily coding updates, concepts, and algorithms.

Fetching latest learning entries from GitHub...

Let's Connect

Looking for a technical Data Engineer / ML Engineer who combines mathematical modeling with production-grade data pipelines? Get in touch via these channels or send a message directly.

Send a Message