Agent skill testing
Test one skill across different agent stacks.
AgentSky runs a custom agent skill against the same real task across multiple harness and model combinations in a live cloud runtime. Teams compare what each configuration produces under consistent conditions, then save the setup that performs best and promote it to production through web, API, or CLI.
Testing a custom skill across agent stacks means standing up a separate runtime for each one.
How it works
One job. One complete cloud agent.
Skill runs live on a cloud agent
Give a live agent the instructions, tools, and data the skill needs without standing up another runtime.
Results compare across harness and model pairs
Run the same kind of work with supported harness and model combinations and inspect the result.
Best configuration is ready to promote
Save the agent configuration that works and continue through web, API, CLI, or a connected channel.
You choose
- Skill and instructions
- Test task
- Harness and compatible model
AgentSky runs
- Live agent runtime
- Tools and app connections
- Web, API, CLI, and channels
Agent Playground
AgentSky ranks agent harnesses and models by task category, so teams know which configuration to test a new skill on first — before committing to a stack.
Learn more