Skip to content
View Zafeeruddin's full-sized avatar

Block or report Zafeeruddin

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Zafeeruddin/README.md

Hey, I'm Zafeer πŸ‘‹

Platform / DevOps Engineer specializing in Kubernetes-based AI Infrastructure, MLOps, GPU Workloads, CI/CD, and On-Prem / Air-Gapped Deployments.

LinkedIn β€’ Email β€’ GitHub


πŸš€ About Me

I build and operate production-grade infrastructure for AI, computer vision, and MLOps platforms.

My work sits at the intersection of:

  • Kubernetes platform engineering
  • MLOps infrastructure
  • GPU-based model training workloads
  • CI/CD and GitOps
  • On-prem and air-gapped deployments
  • Computer vision systems at camera scale
  • Observability, secrets management, and infrastructure automation
  • Hybrid cloud and client-side enterprise deployments

I have worked on high-scale AI/CV platforms involving 50+ Kubernetes nodes, 2,000+ cameras, 50+ GPU machines, and 200+ computer vision use cases across government and enterprise environments.

My strongest fit is where infrastructure meets AI systems.


🧠 Core Engineering Focus

AI Platform Engineering
Kubernetes Infrastructure
MLOps / ML Training Orchestration
DevOps Automation
GPU Workload Scheduling
CI/CD and GitOps
On-Prem / Air-Gapped Deployments
Computer Vision Infrastructure
Cloud + Hybrid Infrastructure
Observability and Reliability
Infrastructure Automation
Edge / Device Onboarding

πŸ› οΈ Tech Stack

Platform & DevOps

Kubernetes Docker ArgoCD Jenkins GitHub Actions Azure DevOps Terraform Ansible

Cloud, Infra & Networking

AWS Alibaba Cloud NGINX HAProxy MetalLB Harbor Keepalived NFS

MLOps, AI & Data

Python PyTorch MLflow Kafka Qdrant Redis InfluxDB Postgres

Observability & Security

Prometheus Grafana HashiCorp Vault New Relic


πŸ—οΈ Highlight: AI / Computer Vision Platform Infrastructure

I worked on a low-code / no-code computer vision platform covering the full AI lifecycle:

Data Collection
↓
Labelling
↓
Model Training
↓
Custom Notebook Training
↓
Real-Time Inference
↓
Camera Management
↓
Device Onboarding
↓
Analytics Dashboards

The platform was used to support a large-scale government computer vision project covering:

  • 200+ AI/CV use cases
  • 2,000+ cameras
  • 50+ GPU machines
  • Multi-environment Kubernetes deployments
  • On-prem and air-gapped infrastructure
  • Real-time camera inference
  • Large-scale model training and deployment workflows

My role covered the DevOps, MLOps, infrastructure, deployment, and platform reliability side of the system.


⚑ Key Engineering Wins

Reduced model training time by ~90%

Built distributed training infrastructure using PyTorch DDP, allowing GPU pooling for computer vision model training.

Before: 55 hours
After:   5 hours

Operated 50+ node Kubernetes infrastructure

Worked as Kubernetes administrator across multiple environments, supporting:

  • platform services
  • backend services
  • frontend services
  • databases
  • Redis / Redis Stack
  • Node-RED
  • agents
  • health services
  • model training jobs
  • inference services
  • notebooks
  • observability components
  • ingress and routing layers

Built production CI/CD pipelines

Created Jenkins and GitHub Actions pipelines for:

  • linting changed services
  • selecting deployment environments
  • building Docker images
  • pushing images to Harbor registry
  • updating image tags
  • updating Kubernetes manifests
  • deploying across dev, prod, and client environments
  • managing environment variables through HashiCorp Vault
  • maintaining semantic versioning
  • enabling multi-environment deployments

Built HA NGINX architecture

Removed a single point of failure by designing a highly available NGINX setup using:

  • Keepalived
  • virtual IP
  • multiple NGINX machines
  • Git-backed NGINX configuration
  • automated config sync
  • failover routing

Built device onboarding infrastructure

Designed and migrated device onboarding from Bash-heavy scripts to a more robust system using:

  • Python microservices
  • Kubernetes Jobs
  • Ansible playbooks and roles
  • Kafka-based status communication
  • Prometheus-based device monitoring
  • persistent deploy manager service

The system supported:

  • single-device onboarding
  • bulk onboarding
  • device patching
  • device deboarding
  • destructive cleanup
  • unified health monitoring
  • multi-phase onboarding lifecycle
  • Kubernetes cluster creation or node joining
  • NVIDIA driver and container runtime setup
  • Docker setup
  • image pre-pulling through DaemonSets
  • persistent device services

Built self-hosted container registry

Hosted Harbor locally, mounted on NAS, proxied over NGINX, with team-based access and regular backups.

Impact:

Saved approximately $1,000 in cloud cost

Reduced platform setup time

Created Docker Compose-based setup for platform services.

Impact:

Before: 2 hours
After:  5 minutes

Saved 300k SAR yearly on client infrastructure

Architected a tightly scoped deployment environment for a Saudi government-related client, reducing yearly infrastructure cost.


πŸ”¬ MLOps Systems I Have Built

Model Training Orchestration

Built an on-demand model training system running as Kubernetes Jobs.

Supported:

  • detection models
  • classification models
  • segmentation models
  • multiple architectures
  • multiple pretrained weights
  • distributed GPU training
  • MLflow integration
  • real-time metric reporting
  • pause training
  • stop training
  • resume training
  • prioritize training
  • Kafka-based progress tracking
  • custom status lifecycle
  • logger and status microservices
  • resilience across interruptions and reboots

Jupyter Notebook Platform

Built notebook infrastructure for custom model training.

Features:

  • each notebook exposed through a live subdomain
  • ingress-based routing
  • ingress controller integration
  • persistent storage on NFS/NAS
  • StatefulSet-based deployment
  • automatic idle monitoring
  • warning before shutdown
  • automatic cleanup after inactivity
  • MetalLB load balancing for on-prem Kubernetes
  • persistent notebook code across restarts

Face Registration Service

Built an on-demand Kubernetes Job-based service for face registration.

Supported:

  • multiple face models
  • embedding generation
  • face registration workflows
  • Qdrant vector database integration
  • scalable execution through Kubernetes Jobs

πŸŽ₯ Computer Vision / Camera Infrastructure

Built and optimized camera management services for real-time streaming workloads.

Worked with:

  • RTSP streams
  • HLS conversion
  • MediaMTX
  • custom camera management APIs
  • multi-camera inference workflows
  • stream lifecycle management
  • camera uptime metrics
  • Prometheus/Grafana dashboards
  • add/remove camera APIs
  • stream expiration handling
  • real-time camera health monitoring

Built an HLS streaming service that converts RTSP streams into HLS and exposes APIs for backend systems to add, remove, and manage cameras.


πŸ§ͺ Early Computer Vision PoCs

Before moving deeply into platform and infrastructure, I worked on multiple computer vision PoCs.

Ajdan

Built a camera-stream-based crowd detection PoC.

Worked on:

  • camera stream capture
  • model training
  • live inference
  • real-time count extraction
  • male/female/children class detection

King Salman Military Base

Worked on computer vision use cases including:

  • employee availability
  • unauthorized access detection
  • child/person class detection
  • officer availability checks
  • alert generation

Gold Chain

Worked on multi-class computer vision classification and retail analytics.

Covered:

  • employee availability
  • people entering shops
  • people exiting shops
  • multi-camera inference

Dubai Airport

Worked on a high-stakes turnaround management system PoC for flight arrivals and departures across multiple gates.

Built an automated stream recovery system using Python, Selenium, and SMTP.

The system:

  • monitored camera stream health
  • detected expired streams
  • automated browser login flows
  • refreshed cookies
  • replaced expired cookies
  • restarted streams
  • sent email notifications

☁️ Client / Deployment Experience

Eastern Provincial Municipality β€” Saudi Arabia

Worked on infrastructure and MLOps execution for a high-scale computer vision project.

Scope:

  • 200+ use cases
  • 2,000+ cameras
  • 50+ GPU machines
  • on-prem Kubernetes
  • AI training and inference workloads
  • multi-environment deployment
  • camera streaming infrastructure
  • device onboarding and patching
  • monitoring and alerting

Ministry of Economy and Petroleum β€” Saudi Arabia

Worked on deployment architecture for AI platforms in air-gapped, on-prem environments.

AmplifAI β€” RAG Platform

Architected end-to-end platform deployment.

Worked on:

  • Dockerization
  • Kubernetes YAML manifests
  • dev, staging, and production environments
  • database connectivity
  • block storage connectivity
  • ingress routing
  • air-gapped deployment constraints

AmplifAI Meetings β€” Meeting Bot

Architected end-to-end platform deployment.

Worked on:

  • Dockerized services
  • Kubernetes deployment
  • dev/staging/prod setup
  • database connectivity
  • block storage
  • ingress routing
  • air-gapped on-prem deployment

Expro β€” Saudi Government-Related Expenditure Entity

Designed and deployed infrastructure for a client environment.

Implemented:

  • Alibaba ECS deployments
  • Alibaba Container Registry
  • Azure DevOps blue-green CI/CD pipelines
  • approval-based deployments
  • rollback support
  • multi-environment deployment flow
  • environment-variable updates
  • local HashiCorp Vault setup

Impact:

Saved approximately 300k SAR yearly

Monsha'at β€” Saudi Government Entity

Deployed AI and platform services using:

  • Alibaba ECS
  • Alibaba ACK Kubernetes
  • multiple VPCs
  • private connectivity between services
  • AI service deployment
  • platform service deployment

🧱 Infrastructure Work

I have worked across the full lifecycle of infrastructure setup, operation, monitoring, deployment, and automation.

Areas covered:

  • on-prem Kubernetes cluster architecture
  • high availability using HAProxy
  • Kubernetes administration
  • NGINX routing and reverse proxying
  • HA NGINX with Keepalived
  • Docker registry hosting with Harbor
  • NAS-mounted registry storage
  • registry backup strategy
  • team-based registry access
  • Vault-based secrets management
  • Terraform-based infrastructure automation
  • Ansible-based device provisioning
  • Docker image optimization
  • multi-stage Docker builds
  • Docker Compose platform packaging
  • Prometheus metrics scraping
  • Grafana dashboards
  • Alertmanager/open alerts
  • StatefulSets for persistent services
  • Redis / InfluxDB persistence
  • Kubeflow Trainer integration
  • GPU training persistence across reboots
  • air-gapped Kubernetes deployment

πŸ“Œ Featured Projects

CloudPad β€” Kubernetes Sandbox Environment

A DevOps-focused platform that lets users create live portfolio/code environments through chat.

Tech: React, FastAPI, Kubernetes, Docker, AWS EC2, Ingress NGINX, Claude MCP

Core idea:

Two pods share the same persistent volume:

1. Code writer / editor pod
2. Code serving pod

Features:

  • live code editing
  • chatbot-based code modification
  • VS Code web server
  • live deployable links
  • Kubernetes-based sandbox environments
  • persistent volume sharing between pods
  • real-time code reflection
  • isolated sandbox environments

Medium Clone β€” Full-Stack Blogging Platform

Tech: React, HonoJS, Cloudflare Workers, Cloudflare Pages, Redis, Postgres, Prisma, New Relic

Features:

  • Google authentication
  • OTP authentication
  • comments and replies
  • live notifications
  • interactive blog engagement
  • Cloudflare Workers backend
  • Cloudflare Pages deployment
  • Postgres on Aiven
  • Prisma Accelerate
  • Redis integration
  • GitHub Actions CI/CD
  • New Relic monitoring
  • frontend deployment using S3
  • CDN distribution
  • Cloudflare R2 image hosting

Live:

codesphere.live

ReminderBot β€” WhatsApp Reminder System

Tech: Python, Celery, Postgres, Twilio

Features:

  • WhatsApp reminders
  • recurring daily reminders
  • single-use reminders
  • Celery-based scheduling
  • Postgres persistence
  • Twilio integration

🧩 Architecture Areas I Care About

How do we deploy AI workloads safely?
How do we make Kubernetes usable for AI teams?
How do we reduce model training time?
How do we onboard edge/GPU devices reliably?
How do we deploy in air-gapped environments?
How do we make CI/CD safe, observable, and rollback-friendly?
How do we make infrastructure repeatable with Terraform and Ansible?
How do we monitor thousands of cameras and AI services?
How do we remove single points of failure?
How do we run GPU workloads reliably on Kubernetes?
How do we make on-prem AI infrastructure feel cloud-like?
How do we build platforms that developers actually enjoy using?

🎯 Roles I Am Best Aligned With

I am interested in roles around:

  • Platform Engineer
  • MLOps Engineer
  • AI Infrastructure Engineer
  • DevOps Engineer β€” Kubernetes
  • Cloud Infrastructure Engineer
  • Kubernetes Engineer
  • SRE β€” AI / Platform / Infrastructure
  • Infrastructure Engineer β€” On-Prem / Hybrid Cloud
  • DevOps Engineer β€” GPU / ML Workloads
  • Platform Engineer β€” Computer Vision / AI Products

My strongest fit is where infrastructure meets AI systems.


🌍 Markets I Am Open To

I am open to Platform, DevOps, MLOps, and AI Infrastructure roles across:

  • UAE
  • Saudi Arabia
  • Gulf region
  • Remote international teams
  • Hybrid cloud / on-prem AI infrastructure teams
  • AI platform teams
  • computer vision infrastructure teams
  • DevOps / platform teams serving ML engineers

πŸ“« Reach Me

Email: mohammed.xafeer@gmail.com 
LinkedIn: www.linkedin.com/in/mohammed-zafeer-3b5a82265
GitHub: https://github.com/Zafeeruddin

Building reliable infrastructure for AI systems that need to work outside demos β€” in real production environments.

Pinned Loading

  1. cloudDoc cloudDoc Public

    Python

  2. scribblr scribblr Public

    An extensive blogging website providing features like, comments, replies, notifications, likes and bookmarks

    TypeScript 2

  3. k8s-ops-toolkit k8s-ops-toolkit Public

    Python