Open-source benchmark for evaluating LLMs on 220 real professional tasks across 9 sectors and 44 occupations. Reproducible experiments, artifact validation, grading, and a live evidence dashboard.
-
Updated
Sep 1, 2026 - Python
Open-source benchmark for evaluating LLMs on 220 real professional tasks across 9 sectors and 44 occupations. Reproducible experiments, artifact validation, grading, and a live evidence dashboard.
GitHub Release asset quality gate and npm CLI: catch missing platforms, wrong architectures, empty installers, version drift, and missing checksums.
Python CLI for paper-reproduction workflows with PDF extraction, artifact manifests, opt-in provider readiness, human-gated experiment scaffolds, benchmark suites, run comparison, reports, and agent handoff.
Lightweight, agent-friendly inspection and contract validation for bioinformatics artifacts.
Worked as a project-based AI Agent Specialist with Handshake AI through the Dynamo project, working on structured software engineering tasks involving AI coding agents, GitHub repositories, task execution, testing, validation, and evaluation.
Mechanistic Workbench (mwb): Local-first mechanistic interpretability workbench for IPython research, agent-readable state, artifact validation, evidence graphs, claim-safe MechanismCards, provenance, run ledgers, SAE/TransformerLens workflows, and reproducible MI experiments
Download Trimble RealWorks for Windows 10/11 (64-bit) and install it as Administrator for a complete setup experience.
To associate your repository with the artifact-validation topic, visit your repo's landing page and select "manage topics."