Benchmark Checklist
Benchmark Checklist is an agent skill (a SKILL.md file) from michael-denyer/pstack-claude. Vet a perf measurement (limiter, tuning, limits, errors, repeatability, relevance, and whether the work happened) before you report or act on it. It works with Claude Code, Codex, Cursor, Gemini CLI, GitHub Copilot and OpenCode and has 1,554 GitHub stars across a repository of 2 listed skills.
- Multi-skill repo
- Plugin marketplace
- Testing & QA
- Actively maintained
Add this skill
Claude
This repository is a Claude Code plugin marketplace. In Claude Code:
/plugin marketplace add michael-denyer/pstack-claude
/plugin install pstack@pstack-claudeIn the Claude apps, zip the benchmark-checklist folder and upload it under Customize > Skills > + > Upload a skill (code execution must be on).
ChatGPT / Codex
Codex reads skills from .agents/skills/ in a repo or ~/.agents/skills/ for every project:
git clone --depth 1 https://github.com/michael-denyer/pstack-claude.git
cp -r pstack-claude/plugins/pstack/skills/benchmark-checklist .agents/skills/benchmark-checklist # repo; ~/.agents/skills for all projectsStandalone skills also load in the ChatGPT desktop app.
Cursor
Cursor loads skills from .cursor/skills/ (or ~/.cursor/skills/) and also reads .claude/skills/:
git clone --depth 1 https://github.com/michael-denyer/pstack-claude.git
cp -r pstack-claude/plugins/pstack/skills/benchmark-checklist .cursor/skills/benchmark-checklist # project; ~/.cursor/skills for all projectsSource (checked Oct 7, 2026): code.claude.com/docs/en/skills (opens in a new tab), code.claude.com/docs/en/plugin-marketplaces (opens in a new tab), support.claude.com/en/articles/12512180-using-skills-in-claude (opens in a new tab), learn.chatgpt.com/docs/build-skills (opens in a new tab), cursor.com/docs/context/skills (opens in a new tab)
What this skill does
Vet a perf measurement (limiter, tuning, limits, errors, repeatability, relevance, and whether the work happened) before you report or act on it. Use when you run a benchmark or report a speedup or regression you measured. Use this when you produce a performance number: a PR's before and after, a regression claim, a hillclimb harness, or a library or config choice. Explain the Number says why. Answer each question below with evidence from a run, not from a guess about the code.
When it triggers
- Use when you run a benchmark or report a speedup or regression you measured.
More skills in michael-denyer/pstack-claude
2 skills are listed from this repository.
Similar skills
More testing & qa skills
Systematic Debugging
obra/superpowers
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.
Testing & QAShellTest Driven Development
obra/superpowers
Use when implementing any feature or bugfix, before writing implementation code.
Testing & QAShellVerification Before Completion
obra/superpowers
Use when about to claim work is complete, fixed, or passing, before committing or creating PRs - requires running verification commands and confirming output before making any success claims…
Testing & QAShellFinishing A Development Branch
obra/superpowers
Use when implementation is complete, all tests pass, and you need to decide how to integrate the work.
Testing & QAShellDiagnosing Superpowers
obra/superpowers
Use when a superpowers session went wrong and your human partner wants to know why — repeated work, ignored plans, stumbles, poor results, a skill that didn't fire, "it took too long", "why is it so…
Testing & QAShellTDD
mattpocock/skills
Test-driven development. Use when the user wants to build features or fix bugs test-first, mentions "red-green-refactor", or wants integration tests.
Testing & QAShell
What is the Benchmark Checklist skill?
Benchmark Checklist is an agent skill (a SKILL.md file) from michael-denyer/pstack-claude. Vet a perf measurement (limiter, tuning, limits, errors, repeatability, relevance, and whether the work happened) before you report or act on it. It works with Claude Code, Codex, Cursor, Gemini CLI, GitHub Copilot and OpenCode and has 1,554 GitHub stars across a repository of 2 listed skills. Its SKILL.md lives at github.com/michael-denyer/pstack-claude/plugins/pstack/skills/benchmark-checklist.
How do I install the Benchmark Checklist skill?
In Claude Code, run /plugin marketplace add michael-denyer/pstack-claude and then /plugin install pstack@pstack-claude. For Codex or Cursor, copy the benchmark-checklist folder into .agents/skills/ or .cursor/skills/.
Is the Benchmark Checklist skill free?
Yes. The repository is open source under the MIT license.
Is Benchmark Checklist maintained?
The repository's most recent commit was on Oct 6, 2026. Its latest release is v0.9.74. appsgit only lists skills from repositories with a commit in the last six months.