A modular, extensible toolkit for red teaming large language models and NLP systems — covering adversarial text attacks (character, word, sentence, semantic), jailbreak evaluation via JailbreakBench, and prompt injection — with pluggable model targets, automated judges, and clean reporting. Built for researchers and AI safety practitioners.
-
Updated
Aug 8, 2026 - Python