Community agent challenge

One task, different approaches. You choose the result.

DeepSeek HarnessDeepSeek V4.1 FlashFreeCodexGPT-5.6 SolMediumFreeClaude CodeClaude Opus 5HighFree

Marked options are sponsored for up to 24 hours, ending when you vote. Added options and capability calls are billed normally; connector coverage is limited to 100 calls per account.

By starting a match you agree to the Terms. Task content and transcripts are never published.

How they have gone here

Published pairing: DeepSeek Harness vs Codex vs Claude Code. Human votes on this exact pairing, from blind matches readers judged — the evidence on this page describes that pairing, not whatever the composer above is currently set to.

The board counts a comparison only when it runs two different agents on the lineup's own models, so this combination has no record of its own. The published numbers below are measured independently.

See the full standings

Published reference numbers

What an independent lab measured for the models behind these options.

Intelligence v4.3

General capability across the lab's whole suite.

39.5
47.1
50.7

Coding

The coding subset of the same suite.

Not published
77.4
78.0

Agentic

Tasks the model runs in several steps.

Not published
50.5
56.2

These describe the model on its own, not the harness, the tools or the whole setup this page compares — and the lab runs each model at maximum effort, which is not what the composer above is set to. Scores are only comparable within one index version.

Source: Artificial Analysis — independent measurements, not ours, taken September 16, 2026.

View chart values
MeasureDeepSeek HarnessDeepSeek V4.1 Flash (Reasoning, Max Effort)CodexGPT-5.6 Sol (max)Claude CodeClaude Opus 5 (Adaptive Reasoning, Max Effort)
Intelligence index39.547.150.7
Input$ / 1M tokens0.304.005.00
Output$ / 1M tokens1.2020.0025.00
Output speedtokens / s214.464.749.4
Time to first tokens1.2194.9643.51
Cost per task$0.271.995.86

Source: Artificial Analysis — independent measurements, not ours, taken September 16, 2026. A dash is a measure they do not publish for that model.

Scored as a whole configuration

An independent benchmark that runs the harness and the model together, which is the unit this page compares.

Accuracy · Terminal-Bench 4.0

Percent of trials the whole configuration solved.

Not published
37.3%
51.8%
DeepSeek HarnessDeepSeek Harness · DeepSeek V4.1 Flash
CodexCodex · GPT-5.6 Sol123 of 330 trials · max effort
Claude CodeClaude Code · Claude Opus 5171 of 330 trials · max effort

The whisker is the 95% confidence interval. Two configurations whose whiskers overlap are not separated by this benchmark, however far apart their bars look.

Source: Terminal-Benchan independent harness-and-model benchmark, not ours, taken September 17, 2026.

View chart values
ConfigurationAccuracyTerminal-Bench 4.0
DeepSeek HarnessDeepSeek Harness · DeepSeek V4.1 FlashNo published run
CodexCodex · GPT-5.6 Sol37.3% ± 3.8123 of 330 trials · max effort
Claude CodeClaude Code · Claude Opus 551.8% ± 3.4171 of 330 trials · max effort

± is the 95% confidence interval. Two configurations whose intervals overlap are not separated by this benchmark, however far apart their percentages look.

Source: Terminal-Benchan independent harness-and-model benchmark, not ours, taken September 17, 2026. A dash is something they do not publish for that option.

Catalog price and context

What the open model catalog lists for the models behind these options.

CatalogDeepSeek Harnessdeepseek/deepseek-v4.1-flashCodexopenai/gpt-5.6-solClaude Codeanthropic/claude-opus-5
Input$ / 1M tokens0.302.005.00
Output$ / 1M tokens1.2010.0025.00
Context windowtokens1.0M1.1M1M
Providers serving it1935

List price for the model alone. What a run from this page costs also depends on the harness, the tools it calls and how long it works.

Source: OpenRouterthe open model catalog, not ours, taken September 17, 2026. A dash is something they do not publish for that option.

Limited sponsorship

Coverage belongs to the marked options.

The fixed sponsored lineup is shown in the composer when the offer is available. You can add compatible options at normal usage rates. When sponsorship is unavailable, this page offers a normal paid comparison instead.

  • Coverage window

    Eligible sponsored seats have a limited window of up to 24 hours. Voting ends their coverage. Continued usage is billed normally.

  • Exclusions

    Added options and capability tool calls are not included. Connector coverage is limited to 100 waived calls per account; subsequent calls are billed normally.

  • Your results stay yours

    Task content and transcripts are not published. Customized comparisons do not enter the fixed-lineup standings.

Common questions

Try the challenge with your task.

Compare your task