Agent skill testing

Test one skill across different agent stacks.

AgentSky runs a custom agent skill against the same real task across multiple harness and model combinations in a live cloud runtime. Teams compare what each configuration produces under consistent conditions, then save the setup that performs best and promote it to production through web, API, or CLI.

Testing a custom skill across agent stacks means standing up a separate runtime for each one.

How it works

One job. One complete cloud agent.

01

Skill runs live on a cloud agent

Give a live agent the instructions, tools, and data the skill needs without standing up another runtime.

02

Results compare across harness and model pairs

Run the same kind of work with supported harness and model combinations and inspect the result.

03

Best configuration is ready to promote

Save the agent configuration that works and continue through web, API, CLI, or a connected channel.

You choose

  • Skill and instructions
  • Test task
  • Harness and compatible model

AgentSky runs

  • Live agent runtime
  • Tools and app connections
  • Web, API, CLI, and channels

Agent Playground

AgentSky ranks agent harnesses and models by task category, so teams know which configuration to test a new skill on first — before committing to a stack.

Learn more

Pick the stack your skill performs best on.

Browse skills