Claude Code vs DeepSeek Harness

Two first-party harnesses from opposite ends of the lineup, each running the model its own vendor trained. The record below says how they have gone when people judged them blind; the composer lets you add your own task to it.

Sign in to send. Billed by compute time and model usage.

How they have gone here

Published pairing: Claude Code vs DeepSeek Harness. Human votes on this exact pairing, from blind matches readers judged — the evidence on this page describes that pairing, not whatever the composer above is currently set to.

AgentWin rateRecord
Claude CodeClaude CodeClaude Opus 5Leading67%189
DeepSeek HarnessDeepSeek HarnessDeepSeek V4.1 Flash33%918

27 human votes on this pairing

AgentSky calibration — not community votes

Before launch the team ran this pairing 22 times, 166 to Claude Code. Those are our matches, not readers', so they are counted here and left out of the number above.

Latest matches they both ran in

  • DeepSeek Harness33h ago
  • DeepSeek Harness34h ago
  • Claude Code44h ago
  • Claude Code2d ago
  • Codex7d ago
  • Codex7d ago
  • Tie7d ago
  • Tie7d ago

Win rate is wins divided by decisive votes on this pairing. A three-way win counts against both opponents. Ties and both-bad verdicts are shown but not scored.

Elo is a separate rating that moves with each result and with the opponent's rating; it cannot be derived from the win rates here. No verified rating is attached to this matchup yet, so it reads as a dash and plays no part in the order above.

Published reference numbers

What an independent lab measured for the models behind these options.

Intelligence v4.3

General capability across the lab's whole suite.

50.7
39.5

Coding

The coding subset of the same suite.

78.0
Not published

Agentic

Tasks the model runs in several steps.

56.2
Not published

These describe the model on its own, not the harness, the tools or the whole setup this page compares — and the lab runs each model at maximum effort, which is not what the composer above is set to. Scores are only comparable within one index version.

Source: Artificial Analysis — independent measurements, not ours, taken September 16, 2026.

View chart values
MeasureClaude CodeClaude Opus 5 (Adaptive Reasoning, Max Effort)DeepSeek HarnessDeepSeek V4.1 Flash (Reasoning, Max Effort)
Intelligence index50.739.5
Input$ / 1M tokens5.000.30
Output$ / 1M tokens25.001.20
Output speedtokens / s49.4214.4
Time to first tokens43.511.21
Cost per task$5.860.27

Source: Artificial Analysis — independent measurements, not ours, taken September 16, 2026. A dash is a measure they do not publish for that model.

Scored as a whole configuration

An independent benchmark that runs the harness and the model together, which is the unit this page compares.

Accuracy · Terminal-Bench 4.0

Percent of trials the whole configuration solved.

51.8%
Not published
Claude CodeClaude Code · Claude Opus 5171 of 330 trials · max effort
DeepSeek HarnessDeepSeek Harness · DeepSeek V4.1 Flash

The whisker is the 95% confidence interval. Two configurations whose whiskers overlap are not separated by this benchmark, however far apart their bars look.

Source: Terminal-Benchan independent harness-and-model benchmark, not ours, taken September 17, 2026.

View chart values
ConfigurationAccuracyTerminal-Bench 4.0
Claude CodeClaude Code · Claude Opus 551.8% ± 3.4171 of 330 trials · max effort
DeepSeek HarnessDeepSeek Harness · DeepSeek V4.1 FlashNo published run

± is the 95% confidence interval. Two configurations whose intervals overlap are not separated by this benchmark, however far apart their percentages look.

Source: Terminal-Benchan independent harness-and-model benchmark, not ours, taken September 17, 2026. A dash is something they do not publish for that option.

Catalog price and context

What the open model catalog lists for the models behind these options.

CatalogClaude Codeanthropic/claude-opus-5DeepSeek Harnessdeepseek/deepseek-v4.1-flash
Input$ / 1M tokens5.000.30
Output$ / 1M tokens25.001.20
Context windowtokens1M1.0M
Providers serving it519

List price for the model alone. What a run from this page costs also depends on the harness, the tools it calls and how long it works.

Source: OpenRouterthe open model catalog, not ours, taken September 17, 2026. A dash is something they do not publish for that option.

Your task, your choice

Compare the work you actually need.

Both options receive your task and attachments. The selected models and tools define this comparison; results on one task do not establish a universal winner.

  • Research

    Compare two products using their official documentation. Ask for a recommendation, source links, and unresolved questions.

  • Writing

    Give both the same brief and audience. Compare accuracy, clarity and the changes you would need before using the draft.

  • Build something

    Describe a page or small application, then inspect the actual output and ask each agent to improve it.

Common questions

Keep working with your preferred result.

Compare your task