FastAPI-based In-Context RAG chatbot backend using LlamaIndex and Google Gemini. Features dynamic entity-matching context slicing, an automatic multi-model fallback chain for 429 quota resilience, thread-safe response caching, and adaptive temporal timeframe fallback.
-
Updated
Aug 20, 2026 - Python