Prompt Injection
59 items from fonduai/awesome-prompt-injection ★674
-
Garak github.com
Automate looking for hallucination, data leakage, prompt injection, misinformation, toxicity generation, jailbreaks, and many other weaknesses in LLM's.
-
Agent Threat Rules (ATR) github.com
An open MIT detection-rule standard for AI-agent and MCP attacks (prompt injection, tool poisoning, context exfiltration), like Sigma or YARA for the agent layer, with OWASP LLM/Agentic and MITRE ATLAS mappings on each rule.
-
OWASP Agent Memory Guard github.com
Open-source scanner for AI agent memory poisoning attacks (OWASP ASI06). Detects prompt injection payloads, memory manipulation patterns, and data exfiltration attempts in agent memory stores. Available as a Python package (pip install agent-memory-guard) and GitHub Action.
-
Damn Vulnerable LLM Agent github.com
A sample chatbot powered by a ReAct agent, implemented with Langchain. It's designed to be an educational tool for security researchers, developers, and enthusiasts to understand and experiment with prompt injection attacks in ReAct agents.
-
LLMVault github.com
Self-hosted CTF-style training range for the OWASP LLM Top 10, with 25 labs across three tiers. Play Mode uses scripted assistants so flags reproduce every time; Live Mode points the same attacks at a real model on your own machine with the secret regenerated per session, so…
-
InjecGuard github.com
Open-source prompt guard with published training data; achieves +30.8% over prior state-of-the-art on the NotInject benchmark, specifically addressing overdefense false positives that break legitimate use cases.
-
brood-box github.com
Hardware-isolated microVM sandbox for running coding agents (Claude Code, Codex, OpenCode) with workspace snapshot isolation, DNS-aware egress control, and MCP authorization profiles to contain damage from prompt injection attacks.
-
PIC Standard github.com
Protocol to block unauthorized or unproven agent actions via intent + provenance checks. Mitigates prompt injection & side-effect risks. Open-source (Apache 2.0).
-
ai-prompt-ctf by c-goosen github.com
One of the few CTFs that tests indirect injection against tool-calling agents, spanning RAG, function calling, and ReAct agent scenarios using LlamaIndex, ChromaDB, GPT-4o, and Llama 3.2.
-
prompt-shield github.com
Self-learning prompt injection detection engine with novel cross-domain techniques: Smith-Waterman sequence alignment (bioinformatics), stylometric discontinuity detection (forensic linguistics), and adversarial fatigue tracking (materials science). 27 detectors, 6 output…
-
Guard Bands github.com
Cryptographic data boundary for LLM applications: untrusted content is wrapped in HMAC-SHA256 or Ed25519 signed markers that bind provenance, lifetime and application context, and the verifier reconstructs that context from trusted state before a protected path proceeds…
-
How AI Prompt Injection Works | Hands-on with LLMs youtube.com
Jan 2026 AppSecEngineer tutorial with a code-level demo of injecting against a real LLM application and live testing of LLM Guard detection. One of the most practical end-to-end tutorials published to date.
-
MCP Prompt Injection: How AI Gets Hacked youtube.com
Nov 2025 hands-on walkthrough showing how prompt injection exploits tool metadata and trust boundaries in Model Context Protocol-integrated agents, which remain a primary attack surface as MCP adoption widens.
-
Prompt Injection in LLM Agents (ReAct, Langchain) youtube.com
Theory and hands-on lab on prompt injection against Langchain ReAct agents.
-
Indirect Prompt Injection Through MCP Tools: A Defense Guide stackone.com
Feb 2026 guide explaining why any MCP tool that reads data written outside your trust boundary (CRM notes, calendar invites, API responses) is an injection vector, with concrete mitigations per tool category.
-
ToxicSkills: Snyk Finds Malware and Prompt Injection in 36% of AI Agent Skills snyk.io
Feb 2026 Snyk research across the ClawHub AI agent skills registry: 36% of audited skills contained security flaws, 1,467 malicious payloads found, and 2.9% used curl | bash remote instruction loading to evade static analysis. Covers indirect injection via poisoned web content…
-
New Prompt Injection Papers: Agents Rule of Two and The Attacker Moves Second simonwillison.net
Simon Willison's Nov 2025 commentary on both landmark papers, including the finding that 12 published defenses were bypassed at >90% success rate using gradient descent and RL-based adaptive attacks.
-
Design Patterns for Securing LLM Agents against Prompt Injections simonwillison.net
Overview of various strategies to mitigate the risk of prompt injection.
-
Prompt injection explained simonwillison.net
Video, slides, and a transcript of an introduction to prompt injection and why it's important.
-
Prompt injection: What's the worst that can happen? simonwillison.net
General overview of Prompt Injection attacks, part of a series.
-
Simon Willison's Blog simonwillison.net
The most consistent independent tracker of real-world prompt injection incidents, new papers, and tooling across the field.
-
Google AI Red Team: Securing AI services.google.com
Google's red team walkthrough of how AI systems are attacked in practice. An early foundational report, useful for framing rather than for current technique.
-
r/llmsecurity reddit.com
The most active subreddit dedicated to LLM security research; a good early-warning channel for real-world incidents and new disclosures.
-
PromptTrace prompttrace.airedlab.com
Free AI security training platform with 7 hands-on prompt injection labs and a 15-level CTF (the Gauntlet) with progressively harder defenses, from prompt-level rules to code guards to LLM classifiers. Unique feature: Context Trace shows the full prompt stack (system prompt…
-
Adversarial Prompting promptingguide.ai
A guide on the various types of adversarial prompting and ways to mitigate them.
-
Augustus praetorian.com
Feb 2026 open-source tool from Praetorian. A single Go binary with 210+ vulnerability probes across 47 attack categories, 28 LLM providers, 90+ detectors, and 7 payload transformation buffs. Built for penetration testing workflows without Python/npm dependencies.
-
Continuously Hardening ChatGPT Atlas Against Prompt Injection Attacks openai.com
OpenAI's Dec 2025 disclosure of a real attack chain (malicious email → agent sends resignation letter) and the RL-trained automated attacker they built to find new injection classes before external adversaries do. OpenAI explicitly states deterministic guarantees are not…
-
How Microsoft Defends Against Indirect Prompt Injection Attacks microsoft.com
Microsoft MSRC's Jul 2025 post on FIDES, an information-flow control system enforcing privilege separation and prompt isolation to deterministically block IPI in Copilot-class agents.
-
-
Synthetic Recollections - A Case Study in Prompt Injection for ReAct LLM Agents labs.withsecure.com
A practical scenario showing how prompt injection can be used to hi-jack the ReAct loop used by LLM agents to inject forged thoughts and associated observations into the LLM context, thus altering the intended behavior.
- next page of items loading…