Skip to content
#

rubric-based-evaluation

Here are 9 public repositories matching this topic...

Language: All
Filter by language
LLM_InSight

This is my personal home rig for serious LLM experimentation. I built it to test models head-to-head, create custom evaluation rubrics, automatically improve prompts based on the previous run’s results, and generate high-quality synthetic training data. Everything runs locally first (Ollama by default), with optional cloud support. logged locally.

  • Updated Jun 15, 2026
  • Python

Analyze Claude Code session logs and generate efficiency reports, cost diagnostics, and actionable recommendations. This project reads local JSONL session logs, computes deterministic efficiency signals, and can optionally add local LLM recommendations using Ollama.

  • Updated Jul 16, 2026
  • Python

A research tool that scores a human-AI conversation to show whether the person got more capable or just more dependent. Claude rates cognitive agency, prompt steering and critical discernment phase by phase; the code checks the arithmetic and the evidence, and a calibration loop lets a human correct the reviewer over time.

  • Updated Jul 22, 2026
  • Python

Add this topic to your repo

To associate your repository with the rubric-based-evaluation topic, visit your repo's landing page and select "manage topics."

Learn more