Hermes vs Kimi Code

Both run a wide range of models rather than one vendor's, which makes the harness itself the thing under test: how each one plans, uses tools and recovers when a step fails. Send a task that takes more than one step and the difference is legible.

Sign in to send. Billed by compute time and model usage.

How they have gone here

Published pairing: Hermes vs Kimi Code. Human votes on this exact pairing, from blind matches readers judged — the evidence on this page describes that pairing, not whatever the composer above is currently set to.

The board counts a comparison only when it runs two different agents on the lineup's own models, so this combination has no record of its own. The published numbers below are measured independently.

See the full standings

Published reference numbers

What an independent lab measured for the models behind these options.

Intelligence v4.3

General capability across the lab's whole suite.

36.3
43.8

Coding

The coding subset of the same suite.

68.8
76.2

Agentic

Tasks the model runs in several steps.

42.3
50.6

These describe the model on its own, not the harness, the tools or the whole setup this page compares — and the lab runs each model at maximum effort, which is not what the composer above is set to. Scores are only comparable within one index version.

Source: Artificial Analysis — independent measurements, not ours, taken September 16, 2026.

View chart values
MeasureHermesDeepSeek V4 Pro 0813 (Reasoning, Max Effort)Kimi CodeKimi K3 (max)
Intelligence index36.343.8
Input$ / 1M tokens1.323.00
Output$ / 1M tokens3.9615.00
Output speedtokens / s94.434.7
Time to first tokens1.744.54
Cost per task$0.672.00

Source: Artificial Analysis — independent measurements, not ours, taken September 16, 2026. A dash is a measure they do not publish for that model.

Catalog price and context

What the open model catalog lists for the models behind these options.

CatalogHermesdeepseek/deepseek-v4-proKimi Codemoonshotai/kimi-k3
Input$ / 1M tokens1.603.00
Output$ / 1M tokens3.2015.00
Context windowtokens1.0M1.0M
Providers serving it1517

List price for the model alone. What a run from this page costs also depends on the harness, the tools it calls and how long it works.

Source: OpenRouterthe open model catalog, not ours, taken September 17, 2026. A dash is something they do not publish for that option.

Your task, your choice

Compare the work you actually need.

Both options receive your task and attachments. The selected models and tools define this comparison; results on one task do not establish a universal winner.

  • Research

    Compare two products using their official documentation. Ask for a recommendation, source links, and unresolved questions.

  • Writing

    Give both the same brief and audience. Compare accuracy, clarity and the changes you would need before using the draft.

  • Build something

    Describe a page or small application, then inspect the actual output and ask each agent to improve it.

Common questions

Keep working with your preferred result.

Compare your task