Claude Code vs Hermes

Claude Code is built around Claude models and file work; Hermes is general-purpose and runs on whichever model you point it at. The choice usually comes down to whether your task is a codebase or everything else.

Sign in to send. Billed by compute time and model usage.

How they have gone here

Published pairing: Claude Code vs Hermes. Human votes on this exact pairing, from blind matches readers judged — the evidence on this page describes that pairing, not whatever the composer above is currently set to.

The board counts a comparison only when it runs two different agents on the lineup's own models, so this combination has no record of its own. The published numbers below are measured independently.

See the full standings

Published reference numbers

What an independent lab measured for the models behind these options.

Intelligence v4.3

General capability across the lab's whole suite.

50.7
36.3

Coding

The coding subset of the same suite.

78.0
68.8

Agentic

Tasks the model runs in several steps.

56.2
42.3

These describe the model on its own, not the harness, the tools or the whole setup this page compares — and the lab runs each model at maximum effort, which is not what the composer above is set to. Scores are only comparable within one index version.

Source: Artificial Analysis — independent measurements, not ours, taken September 16, 2026.

View chart values
MeasureClaude CodeClaude Opus 5 (Adaptive Reasoning, Max Effort)HermesDeepSeek V4 Pro 0813 (Reasoning, Max Effort)
Intelligence index50.736.3
Input$ / 1M tokens5.001.32
Output$ / 1M tokens25.003.96
Output speedtokens / s49.494.4
Time to first tokens43.511.74
Cost per task$5.860.67

Source: Artificial Analysis — independent measurements, not ours, taken September 16, 2026. A dash is a measure they do not publish for that model.

Scored as a whole configuration

An independent benchmark that runs the harness and the model together, which is the unit this page compares.

Accuracy · Terminal-Bench 4.0

Percent of trials the whole configuration solved.

51.8%
Not published
Claude CodeClaude Code · Claude Opus 5171 of 330 trials · max effort
HermesHermes · DeepSeek V4 Pro

The whisker is the 95% confidence interval. Two configurations whose whiskers overlap are not separated by this benchmark, however far apart their bars look.

Source: Terminal-Benchan independent harness-and-model benchmark, not ours, taken September 17, 2026.

View chart values
ConfigurationAccuracyTerminal-Bench 4.0
Claude CodeClaude Code · Claude Opus 551.8% ± 3.4171 of 330 trials · max effort
HermesHermes · DeepSeek V4 ProNo published run

± is the 95% confidence interval. Two configurations whose intervals overlap are not separated by this benchmark, however far apart their percentages look.

Source: Terminal-Benchan independent harness-and-model benchmark, not ours, taken September 17, 2026. A dash is something they do not publish for that option.

Catalog price and context

What the open model catalog lists for the models behind these options.

CatalogClaude Codeanthropic/claude-opus-5Hermesdeepseek/deepseek-v4-pro
Input$ / 1M tokens5.001.60
Output$ / 1M tokens25.003.20
Context windowtokens1M1.0M
Providers serving it515

List price for the model alone. What a run from this page costs also depends on the harness, the tools it calls and how long it works.

Source: OpenRouterthe open model catalog, not ours, taken September 17, 2026. A dash is something they do not publish for that option.

Your task, your choice

Compare the work you actually need.

Both options receive your task and attachments. The selected models and tools define this comparison; results on one task do not establish a universal winner.

  • Research

    Compare two products using their official documentation. Ask for a recommendation, source links, and unresolved questions.

  • Writing

    Give both the same brief and audience. Compare accuracy, clarity and the changes you would need before using the draft.

  • Build something

    Describe a page or small application, then inspect the actual output and ask each agent to improve it.

Common questions

Keep working with your preferred result.

Compare your task