Hi there, I'm Kanav Jeet Singh! β‘
π€ AI Engineer | Hackathon Winner | Latency Slayer
"I don't just call APIs; I optimize what happens behind them."
I'm an AI Engineer who gets a kick out of turning research papers into production-grade systems. While others are satisfied with "it runs," I'm obsessing over inference latency, GPU memory usage, and why the model decided to hallucinate a fictional law of physics.
I bridge the gap between Deep Learning research and real-world deployment. Whether it's fine-tuning LLaMA-3 to be less chatty and more factual, or teaching RL agents to navigate warehouses without crashing, I build systems that work.
π The Highlight Reel
π₯ Winner, Smart India Hackathon 2024: National Hardware Innovation.
π₯ Winner, MahaKumbh Digital Hackathon 2025: AI in Disaster Management.
Internship @ Eastern International University: Built multi-agent RL systems that improved logistics efficiency by 20%.
π οΈ My Arsenal (The Tech Stack)
The Brains π§
The Muscle πͺ
The Ops βοΈ
Generative AI: LLaMA-3, Qwen-VL, DPO, RAG
Core: PyTorch, Transformers, Ray (RL)
Deploy: Docker, Kubernetes (EKS), Helm
Optimization: vLLM, FlashAttention-2, QLORA
Langs: Python (Daily Driver), C++, SQL
Monitor: MLflow, Prometheus, Grafana
Vector DBs: Qdrant, FAISS
Backend: FastAPI
CI/CD: GitHub Actions
π Cool Stuff I've Built
π§ Neuro-Doc: The Enterprise Brain
The Problem: Standard LLMs hallucinate technical details and run slow.
The Fix: Fine-tuned LLaMA-3 with QLORA & Direct Preference Optimization (DPO).
The Impact: Reduced hallucinations by 35% and hit 60 tokens/sec throughput using vLLM on AWS.
π¦ DoubleDQFormer: Logistics Solved
The Problem: Warehouse robots are inefficient.
The Fix: A Transformer-augmented Reinforcement Learning agent.
The Impact: +25% retrieval speed and +60% space utilization compared to heuristics.
πͺοΈ DRISTI AI: Disaster Response
The Problem: Coordinating evacuation during disasters is chaotic.
The Fix: Sim-to-Real RL evacuation agents + YOLOv8 crowd detection.
The Impact: Reduced simulated response times by 30%.
β‘ Current Obsessions
Squeezing every last drop of performance out of vLLM.
Exploring Agentic Workflows that can reason, plan, and execute (not just chat).
Finding the perfect balance between coffee intake and code output. β
π« Let's Build Something Crazy
Email Me β’ LinkedIn β’ Portfolio
[
](https://linkedin.com/in/Kanav Jeet Singh)
[](https://paypal.me/Kanav Jeet Singh)