PRE-PRODUCTION TESTING FOR AGENT SKILLS

Test Agent Skills
before they hit
your stack.

Run a skill across models and agent environments. Compare it with the base model. Inspect every run. Blind-test subjective outputs. Know what you're installing before it reaches production.

I found a skillImport/Test Skill →I need a skillBrowse/Search Skills →I built a skillTest My Skill →

For vibe coders, agent builders, and the people shipping their work.

Available now: text + bundled-reference lab environments.
Native runtime execution and image testing are not yet available.

SKILLRIG / EXECUTION MATRIXExample Run
RIG 001 / WRITING SKILL

Same task. Two conditions.

Skill loadedBaseline runningSkill runningComparingResult
MODEL / LAB ENVIRONMENTBASELINEWITH SKILL
Text model ABundled context
QueuedQueued
Text model AReference reader
QueuedQueued
Text model BBundled context
QueuedQueued
Text model BReference reader
QueuedQueued
OBSERVATIONSkill snapshot loaded. Baselines are ready.

Illustrative sequence only. No AI calls, measurements, or benchmark evidence.

FOR VIBE CODERS AND AGENT BUILDERS

Found a skill online?
Test it before you install it.

A promising SKILL.md can still be a poor fit for your model, tools or taste. Stop downloading skills and debugging blind. Compare a matched baseline, inspect requirements and traces, and see what actually changed before spending hours tweaking.

Import from GitHub →Use Discover Skills → Import skill in the workspace. Native runtimes are not executed by this build.

FOUR QUESTIONS. SEPARATE ANSWERS.

Installing successfully isn't the same
as working correctly.

01 / FIT

Compatibility

Does it work in your stack?

Separate declared requirements, static findings, and what the lab has actually tested.
02 / REPEAT

Reliability

Does it work repeatedly?

Repeat scenarios. Look for failed checks, execution errors, and inconsistent outputs.
03 / DIFFERENCE

Contribution

Does the skill actually improve the base model?

Compare matched conditions. Equal check scores can mean your checks missed the difference.
04 / JUDGEMENT

Taste

Would you actually use the output?

Make the human call separately. A compliant response can still be the wrong response for you.

THE TEST PATH

From GitHub to evidence in minutes.

Start with a small text comparison. Package size, model response times and test scope affect how long it takes.

SOURCE → SNAPSHOT

Bring the actual skill, not just its promise.

Import a public GitHub directory, ClawHub version, SKILL.md or ZIP. Keep the instructions, supporting files and immutable provenance together.

Testing availability ↗

PUBLIC SKILL LIBRARY

Don't have a skill yet?
Find one worth testing.

Start with a genuine indexed package.
Then decide whether it belongs in your rig.

Official means the publisher’s own repository; it is not a SkillRig endorsement. Empty filters mean no indexed match, not no such skills exist.

Current harness: text + bundled references. Ecosystem tags are source/static assessments, not native runtime test results.

GitHub / ImportedReady now · static

brand-guidelines

Applies Anthropic's official brand colors and typography to any sort of artifact that may benefit from having Anthropic's look-and-feel. Use it when brand colors or style guidelines, visual formatting, or company design standards apply.

Claude CodeUniversalNative runtime: not tested

Detected capabilities: text

VERSION sha256:ceabc11d075e
INDEXED 2026-10-03 · LICENCE Complete terms in LICENSE.txt
ClawHub / DocumentationReady now · static

writing

Use when writing or editing prose humans will read — documentation, commit messages, error messages, READMEs, reports, UI text, explanations. Triggers on clarity concerns, wordiness, passive voice, or vague language.

OpenClawUniversalNative runtime: not tested

Detected capabilities: text, references

VERSION 1.0.0
INDEXED 2026-10-03 · LICENCE MIT-0

HUMAN PREFERENCE / BLIND COMPARISON

Passing the test doesn't mean
you'd use the result.

Hide the identities. Read the outputs.
Choose what works for you.

Which would you actually use?Example outputs

TASK / Announce a tool for testing Agent Skills before installation. Keep it clear and concise.

A

Evaluate Agent Skills before installation. Compare outputs across models, review execution traces, and record your preferences to inform deployment decisions.

Identity hidden
B

Before you add a skill to your agent, put it on the rig. Compare what changes, inspect the run, and decide whether you'd actually use the output.

Identity hidden
Your preference

Try a choice. There is no universal winner.

Handwritten examples, not model results. This interaction saves no preference profile. Human preference remains separate from automated compliance.

OBSERVABILITY / RUN TRACE

Don't just get a score.
See what actually happened.

Follow the instructions supplied, reference reads, model responses and explicit checks. Missing usage stays unavailable. Errors stay visible.

Testing availability ↗

Current tool access is limited to imported text references. The lab does not run shell commands, browse the web, or execute imported code.

TRACE / WITH-SKILL CONDITIONExample trace
  1. 01SkillLoaded

    SKILL.mdExact package snapshot recorded

  2. 02ContextPrepared

    Matched task + reference manifestPrimary instructions added to this condition

  3. 03ReferenceRead

    references/voice.mdRead-only · imported package path

  4. 04ModelResponse

    Output capturedModel identity and reported usage retained when available

  5. 05Evaluation

    Configured checks evaluatedHuman preference assessed separately

Illustrative event sequence; no real invocation or timing measurements.

SKILL CONTRIBUTION / MATCHED CONDITIONS

Does the skill actually
make the model better?

Same task. Same model. Same references and tool access. The primary skill instructions are the variable.

TASK / ANNOUNCE A TEXT-TESTING TOOL. DO NOT INVENT FEATURES.Handwritten example
WITHOUT SKILL

Meet your all-in-one AI testing platform. Generate images, automate every workflow, and guarantee production-ready results.

× Unsupported feature claims
WITH SKILL

SkillRig helps you inspect Agent Skills and compare text outputs before installation. Review the evidence, then decide what to use.

✓ Stays within the supplied product facts

An illustration of a difference worth checking, not measured uplift. Real tests may show improvement, no difference, or regression. Literal checks alone do not establish factual accuracy.

REPRODUCIBLE INPUTS / HONEST LIMITS

Know exactly what was tested.

Keep the source version and test configuration with the result. A report is evidence for your decision, not security certification.

Evidence manifestIndexed source + example rig
Source
Loading indexed provenance…

Source provenance is from a real indexed package. Model and run settings below are illustrative; this public example has not been executed.

QUICK START / CURRENT BUILD

A clear first run.

For vibe coders and experienced builders alike: replace “will this skill work for me?” with a comparison you can inspect.

How do I run my first test?

Testing availability, choose a skill in Discover Skills, and select Test skill. Start with one model, one scenario and one repetition. Use DEMO to explore the workflow without AI calls.

What do live tests require?

Connect your own OpenRouter key in Settings, select LIVE, choose exact model IDs, and review the maximum budget before starting. API usage is billed to your OpenRouter account. A ChatGPT subscription does not fund these calls.

Does it run inside Claude Code, Codex or OpenClaw?

Not yet. These are skill ecosystem labels. Current tests use bundled-context and emulated reference-reader lab environments. Native runtime compatibility remains untested.

Can I run tests on this website?

This is the public catalogue and product demonstration. Hosted testing and sign-in are not open yet. All examples are illustrative; this site makes no model calls.

INSPECT FIRST. SHIP WITH EVIDENCE.

Put your next Agent Skill
on the rig.

Find out how it behaves before your production agent does.

The local lab supports free demos and live tests with your own OpenRouter key. Hosted testing is not open yet.

TRY LOCALLY

Hosted testing is not open yet.

Use your local SkillRig build to run tests. No tests or API keys are accepted on this public website. Local builds are currently shared privately; there is no public download or sign-up yet.

Already have the build? With Node.js 22 or later installed, double-click Start-Lab.cmd on Windows. Alternatively, run node server.mjs from the project folder and open the local address printed in the terminal. Start with Demo; Live requires your own OpenRouter key.

Browse Skills

TESTING ACCESS

Private beta coming later.

Hosted sign-in is not available. This website offers public skill discovery and examples only.