Skip to content
appsgit

Prompt Guard

Prompt Guard is an agent skill (a SKILL.md file) from Orchestra-Research/AI-Research-SKILLs. Meta's 86M prompt injection and jailbreak detector. It works with Claude Code, Codex, Cursor, Gemini CLI and OpenCode and has 13,315 GitHub stars across a repository of 5 listed skills.

github.com/Orchestra-Research/AI-Research-SKILLs/07-safety-alignment/prompt-guard (opens in a new tab)

  • Multi-skill repo
  • Plugin marketplace
  • Security
  • Maintained

Add this skill

Claude

This repository is a Claude Code plugin marketplace. In Claude Code:

/plugin marketplace add Orchestra-Research/AI-Research-SKILLs
/plugin   # browse ai-research-skills and install the plugin that contains prompt-guard

In the Claude apps, zip the prompt-guard folder and upload it under Customize > Skills > + > Upload a skill (code execution must be on).

ChatGPT / Codex

Codex reads skills from .agents/skills/ in a repo or ~/.agents/skills/ for every project:

git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git
cp -r AI-Research-SKILLs/07-safety-alignment/prompt-guard .agents/skills/prompt-guard   # repo; ~/.agents/skills for all projects

Standalone skills also load in the ChatGPT desktop app.

Cursor

Cursor loads skills from .cursor/skills/ (or ~/.cursor/skills/) and also reads .claude/skills/:

git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git
cp -r AI-Research-SKILLs/07-safety-alignment/prompt-guard .cursor/skills/prompt-guard   # project; ~/.cursor/skills for all projects

Source (checked Oct 7, 2026): code.claude.com/docs/en/skills (opens in a new tab), code.claude.com/docs/en/plugin-marketplaces (opens in a new tab), support.claude.com/en/articles/12512180-using-skills-in-claude (opens in a new tab), learn.chatgpt.com/docs/build-skills (opens in a new tab), cursor.com/docs/context/skills (opens in a new tab)

What this skill does

Meta's 86M prompt injection and jailbreak detector. Filters malicious prompts and third-party data for LLM apps. 99%+ TPR, <1% FPR. Fast (<2ms GPU). Multilingual (8 languages). Deploy with HuggingFace or batch processing for RAG security. Prompt Guard is an 86M parameter classifier that detects prompt injections and jailbreak attempts in LLM applications.

More skills in Orchestra-Research/AI-Research-SKILLs

5 skills are listed from this repository.

FAQ

Prompt Guard FAQ

Still curious? Email info@appsgit.com.

What is the Prompt Guard skill?

Prompt Guard is an agent skill (a SKILL.md file) from Orchestra-Research/AI-Research-SKILLs. Meta's 86M prompt injection and jailbreak detector. It works with Claude Code, Codex, Cursor, Gemini CLI and OpenCode and has 13,315 GitHub stars across a repository of 5 listed skills. Its SKILL.md lives at github.com/Orchestra-Research/AI-Research-SKILLs/07-safety-alignment/prompt-guard.

How do I install the Prompt Guard skill?

In Claude Code, run /plugin marketplace add Orchestra-Research/AI-Research-SKILLs, then install its plugin from /plugin. For Codex or Cursor, copy the prompt-guard folder into .agents/skills/ or .cursor/skills/.

Is the Prompt Guard skill free?

Yes. The repository is open source under the MIT license.

Is Prompt Guard maintained?

The repository's most recent commit was on Jun 16, 2026. Its latest release is v1.7.2. appsgit only lists skills from repositories with a commit in the last six months.