Harness and model guide

Run Claude Code with Claude Sonnet 5.

This combination is currently selectable in AgentSky. Use the task guide below to prepare an evaluation; the launch link carries this harness and model into the workspace.

A useful first task

Start with clear inputs and an observable result.

Use an authorized repository and a disposable working branch. Include the failing command and expected behavior so the agent can reproduce the issue.

01

Open the selected combination

Use the launch button, sign in if needed, and confirm the harness and model before sending work. Configure the required repository, tools and application access.

02

Adapt this request

In this repository, explain the cause of the failing test, make the smallest relevant change, and run the affected checks. Summarize the diff and any remaining uncertainty.

03

Review the output

Review the actual diff, reproduce the test result and check for unrelated edits. Model selection alone does not establish correctness.

Evaluate the combination

Compare the work, time and total cost.

A supported combination is not a universal best choice. Use representative tasks and include failure recovery and human review in the evaluation.

  • Harness

    The harness controls the agent loop, tool use and working environment.

  • Model

    The selected model supplies the reasoning and generation within that harness. Only compatible, currently offered choices appear in the agent's model list.

  • Usage

    Review model, runtime and tool charges on the pricing page and inspect actual usage for your task. No per-task cost or speed is guaranteed by this guide.

Common questions

Claude Code with Claude Sonnet 5 questions

Will the launch button preserve this choice?

The launch URL specifies this harness and model and preserves them through sign-in. Confirm the selection in the workspace before running the task.

Can I substitute any model?

Use the supported model list on the harness page. The same compatibility rules used by the product determine whether this guide is available.

Is this a benchmark winner?

This is a setup and task guide, not a measured ranking. Consult the dated benchmark results for their measured tasks and scope, then evaluate your own workload.

Evaluate Claude Code with Claude Sonnet 5 on your task.

Try this combination