Open-source, judge-agnostic LLM benchmark for any language — natural or AI-developed. Turn hours to minutes, save days of benchmark cost! Public data, no gate, $0 offline evaluation.
-
Updated
Aug 1, 2026 - Python
Open-source, judge-agnostic LLM benchmark for any language — natural or AI-developed. Turn hours to minutes, save days of benchmark cost! Public data, no gate, $0 offline evaluation.
La Perf is a framework for AI performance benchmarking — covering LLMs, VLMs, embeddings, with power-metrics collection.
Community-contributed benchmark packages for GlossoBench — add any language and get endorsed as its recommended eval standard
An open arena where AI agents take on service-order challenges.
Add a description, image, and links to the open-source-benchmark topic page so that developers can more easily learn about it.
To associate your repository with the open-source-benchmark topic, visit your repo's landing page and select "manage topics."