Jailbreaks FYI

LLM Red Teaming Tools

The tooling landscape for testing LLM applications, as covered by Jailbreaks FYI: the interactive tracker built here, the open-source scanners worth installing, and the benchmarks that decide whether a reported success rate means anything.

No vendor sponsorship, no affiliate links, no paid placement. Every assessment below is drawn from published documentation, source repositories, and primary research.

Built here

Still Works? Tracker

A filterable status matrix of jailbreak technique classes against current frontier and open-weight model families. Filter by model, by attack surface, or to active techniques only. Free, no signup, runs in your browser.

Treat each cell as a directional field assessment for the quarter stated on the page, not a live guarantee. Re-verify against the linked source before you rely on it.

Open-source scanners and harnesses

Four tools carry most of the automated testing done against LLM applications today. They are not substitutes for each other: one scans a model endpoint, one orchestrates adaptive attacks, one runs in CI against an application, one tests a retrieval pipeline.

Full side-by-side on licence, interface, multi-turn support, and CI fit: the AI red teaming tools comparison.

Frameworks, benchmarks, and defences

What tooling does not cover

Every scanner above runs a fixed or semi-adaptive payload library. That catches what has already been categorised. It misses application-specific trust escalation, multi-agent lateral movement through poisoned tool output, and any encoding or decomposition attack newer than the corpus. Automated tooling sets the floor for an assessment; it does not set the ceiling.

Browse all articles · Browse by topic