Skip to content
appsgit

Vision Skills

Vision Skills is an agent skill (a SKILL.md file) from Anionex/agent-vision-toolkit. Local vision CLIs: glance (describe/ask/OCR an image), ground (locate a target, pixel box), detect (element inventory), trace (image to SVG geometry), crop (cut a pixel box to a file), and. It works with Claude Code, Codex and OpenCode and has 1,216 GitHub stars.

github.com/Anionex/agent-vision-toolkit/skills/vision-skills (opens in a new tab)

Add this skill

Claude

Claude Code loads skills from ~/.claude/skills/ (all projects) or .claude/skills/ (one project):

git clone --depth 1 https://github.com/Anionex/agent-vision-toolkit.git
cp -r agent-vision-toolkit/skills/vision-skills ~/.claude/skills/vision-skills   # personal, or .claude/skills in a project

In the Claude apps, zip the vision-skills folder and upload it under Customize > Skills > + > Upload a skill (code execution must be on).

ChatGPT / Codex

Codex reads skills from .agents/skills/ in a repo or ~/.agents/skills/ for every project:

git clone --depth 1 https://github.com/Anionex/agent-vision-toolkit.git
cp -r agent-vision-toolkit/skills/vision-skills .agents/skills/vision-skills   # repo; ~/.agents/skills for all projects

Standalone skills also load in the ChatGPT desktop app.

Cursor

Cursor loads skills from .cursor/skills/ (or ~/.cursor/skills/) and also reads .claude/skills/:

git clone --depth 1 https://github.com/Anionex/agent-vision-toolkit.git
cp -r agent-vision-toolkit/skills/vision-skills .cursor/skills/vision-skills   # project; ~/.cursor/skills for all projects

Source (checked Oct 7, 2026): code.claude.com/docs/en/skills (opens in a new tab), support.claude.com/en/articles/12512180-using-skills-in-claude (opens in a new tab), learn.chatgpt.com/docs/build-skills (opens in a new tab), cursor.com/docs/context/skills (opens in a new tab)

What this skill does

Local vision CLIs: glance (describe/ask/OCR an image), ground (locate a target, pixel box), detect (element inventory), trace (image to SVG geometry), crop (cut a pixel box to a file), and scripts/html_shot.py (HTML file to image). Use for any task involving an image — questions, text, splitting and transcribing long screenshots or chat histories, locating elements, comparing, rebuilding as HTML/SVG, digitizing a sketch or diagram, reading values off a chart, operating a GUI from screenshots —…

When it triggers

  • Use for any task involving an image — questions, text, splitting and transcribing long screenshots or chat histories, locating elements, comparing, rebuilding as HTML/SVG, digitizing a sketch or diagram, reading values off a chart, operatin

FAQ

Vision Skills FAQ

Still curious? Email info@appsgit.com.

What is the Vision Skills skill?

Vision Skills is an agent skill (a SKILL.md file) from Anionex/agent-vision-toolkit. Local vision CLIs: glance (describe/ask/OCR an image), ground (locate a target, pixel box), detect (element inventory), trace (image to SVG geometry), crop (cut a pixel box to a file), and. It works with Claude Code, Codex and OpenCode and has 1,216 GitHub stars. Its SKILL.md lives at github.com/Anionex/agent-vision-toolkit/skills/vision-skills.

How do I install the Vision Skills skill?

Copy the vision-skills folder (the one containing SKILL.md) into ~/.claude/skills/ for Claude Code, .agents/skills/ for Codex or .cursor/skills/ for Cursor. The agent picks it up automatically when a task matches its description.

Is the Vision Skills skill free?

Yes. The repository is open source under the MIT license. The skill mentions an API key or token for an external service, which may need its own account.

Is Vision Skills maintained?

The repository's most recent commit was on Aug 27, 2026. Its latest release is v0.2.0. appsgit only lists skills from repositories with a commit in the last six months.