#jailbreaks
-
Payload Splitting and Encoding Jailbreaks Explained
How payload splitting and encoding jailbreaks evade input filters: fragmentation, ciphers, low-resource translation, decomposition, and what still lands.
-
Best AI Red Teaming Tools for LLMs Compared
Compare the best AI red teaming tools for LLMs: Garak, PyRIT, Promptfoo, and Giskard, on probe coverage, CI/CD fit, and OWASP LLM Top 10 mapping.
-
How LLM Jailbreaks Work: Techniques and Success Rates
A practitioner's breakdown of how LLM jailbreaks work, from roleplay conditioning to multi-turn manipulation, with attack success rates from research.
-
DAN Prompt Jailbreak Explained: How 'Do Anything Now' Works
DAN (Do Anything Now) is the most replicated persona-injection jailbreak in LLM history. Here is the mechanism, why it worked, and how versions evolved.
-
LLM Defense Stack: Guardrails, Tool Scoping, and Egress
An LLM defense stack maps jailbreaks and prompt injection to guardrail models, scoped tools, least privilege, egress controls, and human approval.
-
Why Jailbreaks Work: Competing Objectives and Generalization
Jailbreaks are not a grab-bag of tricks. They exploit two structural failure modes of safety training: competing objectives and mismatched generalization.
-
ArtPrompt Post-Mortem: Why ASCII-Art Bypasses Worked
A walkthrough of the ArtPrompt ASCII-art jailbreak: where it slipped past safety training, how model families patched it, and which encodings still land.
-
Multi-Turn Role-Play Attacks: Why One Safe Turn Gets Unsafe
Crescendo, many-shot and gradual context manipulation: how multi-turn jailbreaks evade single-turn classifiers, and what still lands against defenses.
-
Multimodal Jailbreaks: Image and Audio Attack Surfaces
Vision and audio inputs are a separate attack channel from text. A survey of the multimodal jailbreaks that still land, from typographic to audio carriers.
-
System Prompt Extraction: Techniques and Defenses
How system prompts get exfiltrated from production LLM apps: direct extraction, indirect inference, behavioral fingerprinting, and what actually works.
-
Jailbreak Technique Catalog: Working as of 2026 Q2
This catalog tracks ten jailbreak technique classes across production LLMs, covering Q2 2026 status, attack surfaces, trends, and defender guidance.