Fast LLM speculative inference server for consumer hardware.
-
Updated
Jul 24, 2026 - C++
Fast LLM speculative inference server for consumer hardware.
Serve Poolside Laguna S 2.1 (NVFP4) on the NVIDIA DGX Spark (GB10) without hanging your box. Working stack, crash-safe configs, benchmarks, and the exact gotchas.
Poolside AI model provider for Pi with live model discovery and reasoning controls
the sunniest spot on your Macintosh™️
Use hosted Poolside coding models in GitHub Copilot Chat with your Poolside API key
OpenCode plugin for the Poolside AI model provider with live model discovery
poolside Laguna-S-2.1 INT4 + DFlash speculative decoding on 4x RTX 3090: 200K context, 282 tok/s peak decode, gate-proven with a 190K-token prompt. Full levers menu + failure catalog.
Experimental fork of the Séance scrolling terminal multiplexer that tracks your AI coding agents.
Poolside Laguna S 2.1: Run 1M Context Locally (Tested) - Complete overview, benchmarks, local setup guides (vLLM, SGLang, llama.cpp), and test suite for Laguna S 2.1 118B MoE.
vLLM serving poolside Laguna S 2.1 (118B MoE, 8B active, INT4) on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151, 128GB unified) via TheRock ROCm nightly. OpenAI-compatible, 256K context, DFlash speculative decoding.
Add a description, image, and links to the poolside topic page so that developers can more easily learn about it.
To associate your repository with the poolside topic, visit your repo's landing page and select "manage topics."