Discrete Kakeya cover for LLM KV cache: D4/E8 nested-lattice quantisation realising a Kakeya-style tube-cover over the direction sphere. 2.4x-2.8x compression at <1% perplexity loss on Qwen3, Llama-3, DeepSeek, GLM-4, Gemma. Drop-in transformers.DynamicCache. pip install kakeyalattice.
transformers quantization discrete-geometry kv-cache long-context vllm llm-inference kv-cache-compression qwen3 lattice-quantization e8-lattice d4-lattice kakeya kakeya-set
-
Updated
Jun 15, 2026 - Python