Evaluating PAIR, GCG, and Prompt-RS jailbreak attacks against LLaMA-3.1, LLaMA-4, and Qwen3-32B. Two-stage defense pipeline (prompt sanitization + LlamaGuard) achieving 95% Defense Block Rate.
nlp jailbreak pytorch llama ai-safety adversarial-attacks huggingface llm llm-security llama-guard jailbreakbench
-
Updated
May 31, 2026 - Jupyter Notebook