A research tool that scores a human-AI conversation to show whether the person got more capable or just more dependent. Claude rates cognitive agency, prompt steering and critical discernment phase by phase; the code checks the arithmetic and the evidence, and a calibration loop lets a human correct the reviewer over time.
python evaluation-framework ai-education streamlit-webapp claude-api rubric-based-evaluation ai-education-edtech learning-analytics-tools
-
Updated
Jul 22, 2026 - Python