Vision Skills
Vision Skills is an agent skill (a SKILL.md file) from Anionex/agent-vision-toolkit. Local vision CLIs: glance (describe/ask/OCR an image), ground (locate a target, pixel box), detect (element inventory), trace (image to SVG geometry), crop (cut a pixel box to a file), and. It works with Claude Code, Codex and OpenCode and has 1,216 GitHub stars.
github.com/Anionex/agent-vision-toolkit/skills/vision-skills (opens in a new tab)
- Needs API key
- Data & analysis
- Actively maintained
Add this skill
Claude
Claude Code loads skills from ~/.claude/skills/ (all projects) or .claude/skills/ (one project):
git clone --depth 1 https://github.com/Anionex/agent-vision-toolkit.git
cp -r agent-vision-toolkit/skills/vision-skills ~/.claude/skills/vision-skills # personal, or .claude/skills in a projectIn the Claude apps, zip the vision-skills folder and upload it under Customize > Skills > + > Upload a skill (code execution must be on).
ChatGPT / Codex
Codex reads skills from .agents/skills/ in a repo or ~/.agents/skills/ for every project:
git clone --depth 1 https://github.com/Anionex/agent-vision-toolkit.git
cp -r agent-vision-toolkit/skills/vision-skills .agents/skills/vision-skills # repo; ~/.agents/skills for all projectsStandalone skills also load in the ChatGPT desktop app.
Cursor
Cursor loads skills from .cursor/skills/ (or ~/.cursor/skills/) and also reads .claude/skills/:
git clone --depth 1 https://github.com/Anionex/agent-vision-toolkit.git
cp -r agent-vision-toolkit/skills/vision-skills .cursor/skills/vision-skills # project; ~/.cursor/skills for all projectsSource (checked Oct 7, 2026): code.claude.com/docs/en/skills (opens in a new tab), support.claude.com/en/articles/12512180-using-skills-in-claude (opens in a new tab), learn.chatgpt.com/docs/build-skills (opens in a new tab), cursor.com/docs/context/skills (opens in a new tab)
What this skill does
Local vision CLIs: glance (describe/ask/OCR an image), ground (locate a target, pixel box), detect (element inventory), trace (image to SVG geometry), crop (cut a pixel box to a file), and scripts/html_shot.py (HTML file to image). Use for any task involving an image — questions, text, splitting and transcribing long screenshots or chat histories, locating elements, comparing, rebuilding as HTML/SVG, digitizing a sketch or diagram, reading values off a chart, operating a GUI from screenshots —…
When it triggers
- Use for any task involving an image — questions, text, splitting and transcribing long screenshots or chat histories, locating elements, comparing, rebuilding as HTML/SVG, digitizing a sketch or diagram, reading values off a chart, operatin
Similar skills
More data & analysis skills
Migration
JuliusBrussee/caveman
Implement reversible compatibility-safe transitions.
Data & analysisGoDating Web
nexu-io/open-design
A consumer-feeling dating / matchmaking dashboard — left rail navigation, ticker bar of community signals, headline KPIs, a 30-day mutual-matches bar chart, and a match-rate trend block.
Data & analysisTypeScriptDcf Valuation
nexu-io/open-design
Discounted cash flow valuation and intrinsic value analysis for public companies.
Data & analysisTypeScriptFlowai Live Dashboard Template
nexu-io/open-design
Team-management dashboard skill in the FlowAI aesthetic — three tabs (Team Members, Team Details, Activity Log), KPI stat row, member table, role distribution bar chart, online presence and activity…
Data & analysisTypeScriptHow It Works
thedotmack/claude-mem
Explain how claude-mem captures observations, when memory injection kicks in, and where data lives.
Data & analysisTypeScriptWeekly Digests
thedotmack/claude-mem
Generate a serial week-by-week narrative digest of a project's full claude-mem timeline.
Data & analysisTypeScript
What is the Vision Skills skill?
Vision Skills is an agent skill (a SKILL.md file) from Anionex/agent-vision-toolkit. Local vision CLIs: glance (describe/ask/OCR an image), ground (locate a target, pixel box), detect (element inventory), trace (image to SVG geometry), crop (cut a pixel box to a file), and. It works with Claude Code, Codex and OpenCode and has 1,216 GitHub stars. Its SKILL.md lives at github.com/Anionex/agent-vision-toolkit/skills/vision-skills.
How do I install the Vision Skills skill?
Copy the vision-skills folder (the one containing SKILL.md) into ~/.claude/skills/ for Claude Code, .agents/skills/ for Codex or .cursor/skills/ for Cursor. The agent picks it up automatically when a task matches its description.
Is the Vision Skills skill free?
Yes. The repository is open source under the MIT license. The skill mentions an API key or token for an external service, which may need its own account.
Is Vision Skills maintained?
The repository's most recent commit was on Aug 27, 2026. Its latest release is v0.2.0. appsgit only lists skills from repositories with a commit in the last six months.