Codex vs pi

Codex is OpenAI's full coding harness; pi is deliberately small and lets you steer it mid-run. Whether the extra machinery earns its keep depends on how much of the work you want to direct yourself.

Sign in to send. Billed by compute time and model usage.

How they have gone here

Published pairing: Codex vs pi. Human votes on this exact pairing, from blind matches readers judged — the evidence on this page describes that pairing, not whatever the composer above is currently set to.

The board counts a comparison only when it runs two different agents on the lineup's own models, so this combination has no record of its own. The published numbers below are measured independently.

See the full standings

Published reference numbers

What an independent lab measured for the models behind these options.

Intelligence v4.3

General capability across the lab's whole suite.

47.1
47.1

Coding

The coding subset of the same suite.

77.4
77.4

Agentic

Tasks the model runs in several steps.

50.5
50.5

These describe the model on its own, not the harness, the tools or the whole setup this page compares — and the lab runs each model at maximum effort, which is not what the composer above is set to. Scores are only comparable within one index version.

Source: Artificial Analysis — independent measurements, not ours, taken September 16, 2026.

View chart values
MeasureCodexGPT-5.6 Sol (max)piGPT-5.6 Sol (max)
Intelligence index47.147.1
Input$ / 1M tokens4.004.00
Output$ / 1M tokens20.0020.00
Output speedtokens / s64.764.7
Time to first tokens94.9694.96
Cost per task$1.991.99

Source: Artificial Analysis — independent measurements, not ours, taken September 16, 2026. A dash is a measure they do not publish for that model.

Scored as a whole configuration

An independent benchmark that runs the harness and the model together, which is the unit this page compares.

Accuracy · Terminal-Bench 4.0

Percent of trials the whole configuration solved.

37.3%
Not published
CodexCodex · GPT-5.6 Sol123 of 330 trials · max effort
pipi · GPT-5.6 Sol

The whisker is the 95% confidence interval. Two configurations whose whiskers overlap are not separated by this benchmark, however far apart their bars look.

Source: Terminal-Benchan independent harness-and-model benchmark, not ours, taken September 17, 2026.

View chart values
ConfigurationAccuracyTerminal-Bench 4.0
CodexCodex · GPT-5.6 Sol37.3% ± 3.8123 of 330 trials · max effort
pipi · GPT-5.6 SolNo published run

± is the 95% confidence interval. Two configurations whose intervals overlap are not separated by this benchmark, however far apart their percentages look.

Source: Terminal-Benchan independent harness-and-model benchmark, not ours, taken September 17, 2026. A dash is something they do not publish for that option.

Catalog price and context

What the open model catalog lists for the models behind these options.

CatalogCodexopenai/gpt-5.6-solpiopenai/gpt-5.6-sol
Input$ / 1M tokens2.002.00
Output$ / 1M tokens10.0010.00
Context windowtokens1.1M1.1M
Providers serving it33

List price for the model alone. What a run from this page costs also depends on the harness, the tools it calls and how long it works.

Source: OpenRouterthe open model catalog, not ours, taken September 17, 2026. A dash is something they do not publish for that option.

Your task, your choice

Compare the work you actually need.

Both options receive your task and attachments. The selected models and tools define this comparison; results on one task do not establish a universal winner.

  • Research

    Compare two products using their official documentation. Ask for a recommendation, source links, and unresolved questions.

  • Writing

    Give both the same brief and audience. Compare accuracy, clarity and the changes you would need before using the draft.

  • Build something

    Describe a page or small application, then inspect the actual output and ask each agent to improve it.

Common questions

Keep working with your preferred result.

Compare your task