Harness and model guide
Run Codex with GPT-6 Astra.
This combination is currently selectable in AgentSky. Use the task guide below to prepare an evaluation; the launch link carries this harness and model into the workspace.
A useful first task
Start with clear inputs and an observable result.
Provide the repository, base and candidate revision, and the intended behavior. Give read access for a review-only task; specify separately if fixes are wanted.
Open the selected combination
Use the launch button, sign in if needed, and confirm the harness and model before sending work. Configure the required repository, tools and application access.
Adapt this request
Review this change against its stated requirement. Trace the affected callers, report reproducible defects with file references, and distinguish confirmed issues from questions. Do not modify the branch.
Review the output
Reproduce important findings against the named revision. A plausible review comment is not proof of a defect, and a clean review does not replace the relevant tests.
Evaluate the combination
Compare the work, time and total cost.
A supported combination is not a universal best choice. Use representative tasks and include failure recovery and human review in the evaluation.
Harness
The harness controls the agent loop, tool use and working environment.
Model
The selected model supplies the reasoning and generation within that harness. Only compatible, currently offered choices appear in the agent's model list.
Usage
Review model, runtime and tool charges on the pricing page and inspect actual usage for your task. No per-task cost or speed is guaranteed by this guide.
Common questions
Codex with GPT-6 Astra questions
Will the launch button preserve this choice?
The launch URL specifies this harness and model and preserves them through sign-in. Confirm the selection in the workspace before running the task.
Can I substitute any model?
Use the supported model list on the harness page. The same compatibility rules used by the product determine whether this guide is available.
Is this a benchmark winner?
This is a setup and task guide, not a measured ranking. Consult the dated benchmark results for their measured tasks and scope, then evaluate your own workload.
