I make models do cool stuff for you at scale :)
100,000+ model downloads · 5.2M+ views of my technical writing · 1,000+ public GPU-template hours/month · Multi hackathon winner
- First author, ExecRetrieval: Measuring the Functional-Correctness Gap in Code-Embedding Retrieval - EMNLP 2026 Main Conference. Code retrieval that measures whether the code runs, not whether it looks right. (Pending Preprint) Dataset + harness
- Surface - a universal display for AI agents. Your agent renders live UI onto any screen you own; every tap, stroke and answer lands back in its context, even hours after the session ended. One CLI + one skill, shared by every agent on the machine.
npm i -g surface-display - 24hr-research-agent - decomposes a question into 200+ tasks and compiles book-length, fully cited reports from 1,000+ sources
- SOTA-Coder - autonomous coding-agent harness built from scratch: event-driven runtime, WebSocket streaming, dynamic MCP tool loading, no agent-SDK dependencies
- video-production-skill, d2l-cli / degreeworks-cli
- AK quant lines - custom per-tensor bit allocations, benchmarked head-to-head on KL divergence with 95% bootstrap CIs:
- Qwen3.5-9B: 31 wins / 3 ties / 0 losses across 34 comparisons vs Unsloth, bartowski, lmstudio, mradermacher, byteshape, AtomicChat - up to 49% lower KLD than Unsloth at the same size
- Meta Muse-Glimmer-30B: 26 wins / 6 ties / 0 losses vs Meta's own, Unsloth and bartowski
- Auto-Quant (private) - my end-to-end quantize-test-release pipeline behind huggingface.co/AaryanK: Day-0 GGUFs for major launches, 100k+ downloads
- TurboQuant KV-cache quantization in llama.cpp - 4.4× smaller KV cache, no throughput loss; ported to ik_llama.cpp by the community
- One of the first toggleable-reasoning LLMs (Feb 2025, DOI) - trained with GRPO on a single consumer GPU
- Also: upstream llama.cpp contributor (MiMo-V2-Flash, Kimi Linear)
Will work for GPUs · The Why: waiting for RSI so I can take a nap
Hugging Face • LinkedIn • OpenReview • Email




