Open-source runtime profiler & economic measurement layer for AI agents, execution DAGs, and multi-step LLM pipelines.
-
Updated
Aug 30, 2026 - Python
Open-source runtime profiler & economic measurement layer for AI agents, execution DAGs, and multi-step LLM pipelines.
Local RAG evaluation framework for compressed trace retrieval, grounding, citation support, multi-step traces, latency profiling, and agentic efficiency.
Autonomy ML project for driving data, failure mining, policy learning, safety metrics, and latency benchmarking.
Final Year Research & Engineering Project — A hardware-aware compression pipeline that profiles layer-by-layer CPU execution metrics. It balances model perplexity against physical latency, eliminates microarchitectural overhead, and compiles optimized mixed-precision graphs into deployable .gguf format.
High-performance C++20 order book engine with REST API, React web terminal, LOBSTER replay, and online ML pipeline.
Optimized multimodal video understanding pipeline using Silero VAD gatekeeping and Groq LPUs to reduce processing latency by over 70%.
Performance analysis, data cleaning and visualization for Digital Twin experiments running on industrial testbeds (Fischertechnik). The result is published in the scientific article "Hierarchical Digital Twin Ecosystem for Industrial Manufacturing Scenarios" (2024 50th Euromicro Conference on Software Engineering and Advanced Applications - SEAA).
To associate your repository with the latency-profiling topic, visit your repo's landing page and select "manage topics."