Self-improving long-horizon LLM agent — ChromaDB strategy memory + failure analysis, Grok-4 teacher labels → QLoRA-distilled LLaMA-3.2-1B student. 90% on Tau Bench, 95% inference cost reduction.
llama agents knowledge-distillation llm chromadb qlora terminal-bench long-horizon-tasks self-improving-agents agentbench
-
Updated
Jul 14, 2026 - Python