Claude Code vs Codex for lawyers

Try commercial contract review, or bring your own task.

Sign in to send. Billed by compute time and model usage.

lawyers workflow example

Compare agents on commercial contract review

Review the services agreement I attach from the customer’s perspective. Identify clauses concerning liability, termination, payment, and intellectual property. Quote each relevant clause, explain the practical risk, and suggest revised wording. Ask for the governing jurisdiction before making jurisdiction-specific legal conclusions.

Send the same task to every option, inspect the anonymous work, and choose the result you would actually keep.

  • Provide

    Add the source material, constraints and desired output format that matter for this workflow.

  • Inspect

    Clause references, missed risks and the usefulness of suggested wording. Confirm the governing law before relying on legal conclusions.

  • Continue

    Reveal the selected configuration, keep the conversation, and refine the result in the same session.

Which option do users prefer?

These votes cover all tasks for the selected agents and models, not this profession specifically.

7 tries · 5 votes · Same agents and models

#AgentWin rateRecord
1Claude CodeClaude CodeClaude Opus 5Leading100%40
2CodexCodexGPT-5.6 Sol0%04

0 ties · 1 neither result. Tools and thinking levels may vary between attempts.

Head-to-head votes

Direct preference counts within this lineup.

Claude Code40Codex

Recent user choices

See which results other people preferred. Their tasks and files stay private.

  • Claude Code preferred
  • Neither result preferred
  • Claude Code preferred
  • Claude Code preferred
  • Claude Code preferred

Model benchmarks

Artificial Analysis

Intelligence v4.3

General capability across the lab's whole suite.

50.7
47.1

Coding

The coding subset of the same suite.

78.0
77.4

Agentic

Tasks the model runs in several steps.

56.2
50.5

These describe the model on its own, not the harness, the tools or the whole setup this page compares — and the lab runs each model at maximum effort, which is not what the composer above is set to. Scores are only comparable within one index version.

Source: Artificial Analysis — independent measurements, not ours, taken 2026-09-16.

Agent performance on real tasks

Published Terminal-Bench 4.0 runs. Each result names its tested configuration.

Accuracy · Terminal-Bench 4.0

Percent of trials the whole configuration solved.

51.8%
37.3%
Claude CodeClaude Code · Claude Opus 5171 of 330 trials · max effort
CodexCodex · GPT-5.6 Sol123 of 330 trials · max effort

The whisker is the 95% confidence interval. Two configurations whose whiskers overlap are not separated by this benchmark, however far apart their bars look.

Source: Terminal-Benchan independent harness-and-model benchmark, not ours, taken 2026-09-17.

Claude Code vs Codex for lawyers: models, pricing and features

Compare these configurations on the same task, choose the anonymous result you prefer, then continue your chosen session. Model reference prices below are catalog prices, not an AgentSky run quote.

ConfigurationClaude CodeCodex
ModelClaude Opus 5GPT-5.6 Sol
Thinking levelHighMedium
CatalogClaude Codeanthropic/claude-opus-5Codexopenai/gpt-5.6-sol
Input$ / 1M tokens5.002.00
Output$ / 1M tokens25.0010.00
Context windowtokens1M1.1M
Providers serving it53

List price for the model alone. What a run from this page costs also depends on the harness, the tools it calls and how long it works.

Source: OpenRouterthe open model catalog, not ours, taken 2026-09-17. A dash is something they do not publish for that option.

Frequently asked questions