Benchmarks

Jev & decision models

The right decision.
At the right cost.

Compare how well Jev and its alternatives choose, estimate confidence and respond. Explore the trade-offs before choosing a model.

Source: Benchmark Heaven / JevBench, v1.2.1. These are their published measurements; AgentSky has not independently reproduced this run.

Intelligence

Accuracy across easy, standard, judge and hard decisions.

Calibration

How well confidence matches correctness and known probabilities.

Speed

Median and tail latency for one decision at a time.

Cost

The price of a thousand decisions, including pricing assumptions.

Benchmark Heaven / JevBench · v1.2.1

534-decision full suite · 16 systems · measured 2026-09-19

What matters to you?

Equal weights by default. Adjust the balance to explore your priorities.

Weights are normalized to 100%. The geometric mean rewards balance: one strong axis cannot cancel out a weak one. Custom weights change your view, not the published result.

Decision model ranking

Intelligence, calibration, speed and cost. Equal weight, one score.

Score out of 100
  1. 0175.3
  2. 0274.6
  3. 03
    djev (Maisa, diffusion-gemma)

    Maisa (David Villalón)

    74.3
  4. 0469.8
  5. 0568.7
  6. 0667.6
  7. 0766.2
  8. 08
    GPT-5.6 Luna (low reasoning effort)

    OpenAI

    66.0

Partial runs remain in the results table and are excluded from this ranking.

Look beyond the overall score

Compare the same four axes directly. Further right is better.

Jev 1.13.0 (TypeSafe AI)SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)

Intelligence

90.4
85.9

Calibration

82.7
72.6

Speed

83.3
83.7

Cost

51.7
59.2

Every measurement

Sort a column to inspect the trade-offs. Missing measurements stay last within complete and partial runs.

Model
Jev 1.13.0 (TypeSafe AI)

TypeSafe AI

534 / 534 attempted

75.390.482.783.351.7$0.041100.0%99.0%94.5%74.1%0.65s0.72s
SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)

Theodore Lee (TheoLeeCJ)

534 / 534 attempted

74.685.972.683.759.2~$0.023estimated100.0%97.9%95.2%59.5%0.20s0.32s
djev (Maisa, diffusion-gemma)

Maisa (David Villalón)

534 / 534 attempted

74.388.465.491.457.6$0.026announced100.0%97.9%93.2%69.5%0.24s0.31s
69.875.663.283.559.6~$0.022estimated100.0%84.4%74.7%56.8%0.21s0.32s
system-one-open (Gemma 4 E2B LoRA on an L4)

mithalouni

534 / 534 attempted

68.779.556.777.064.1~$0.016estimated100.0%93.8%87.7%49.1%0.65s0.77s
OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16)

razorback16 / Codiv

534 / 534 attempted

67.686.064.883.245.2~$0.067estimated100.0%95.8%91.1%65.5%0.24s0.31s
66.288.977.477.136.1~$0.135estimated100.0%95.8%95.2%71.4%0.68s0.73s
GPT-5.6 Luna (low reasoning effort)

OpenAI

534 / 534 attempted

66.096.889.877.528.2$0.247100.0%97.9%96.6%94.5%0.97s1.82s
open-jev-deberta-v3-large (local CPU)

Kotoba Labs

534 / 534 attempted

64.453.666.466.073.3~$0.0077estimated100.0%49.0%53.4%36.4%1.77s3.35s
Bespoke Nimble 9B (Bespoke Labs)

Bespoke Labs

534 / 534 attempted

63.578.664.582.538.9~$0.109estimated100.0%94.8%89.0%43.6%0.19s0.46s
Gemini 3.1 Flash-Lite

Google

534 / 534 attempted

60.890.368.181.827.1$0.268100.0%99.0%93.2%75.0%0.76s0.88s
DeepSeek V4.1 Flash (thinking default)

DeepSeek

534 / 534 attempted

58.196.196.771.617.1$0.57998.6%99.0%93.2%95.0%1.42s4.89s
system-one (Qwen3-8B, Sean Goedecke)

Sean Goedecke

534 / 534 attempted

56.580.136.884.441.2~$0.092estimated100.0%90.6%91.8%50.0%0.17s0.30s
Qwen3.8 27B (Chutes TEE)

Qwen / Chutes · Partial run

346 / 534 attempted

25.574.692.161.30.0~$2.711estimated98.6%99.0%95.3%21.4%5.75s12.97s
Needle 3, options as tools (post-hoc adapter mode)

Cactus Compute · Partial run

314 / 534 attempted

19.139.5Not measured (scored 0)52.863.7~$0.016estimated66.7%31.3%34.2%3.78s33.64s
Needle 3 (Cactus, 2-bit, local CPU)

Cactus Compute · Partial run

358 / 534 attempted

16.722.4Not measured (scored 0)59.958.1~$0.025estimated47.2%16.7%31.5%7.7%1.69s14.36s

Latency shows raw request time. “Estimated” and “announced” costs are not paid bills. A dash means no measurement; a zero is a measured or explicitly scored zero.

Configuration and coverage notes

SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ). Speed score uses x2 + 0.15 s (assumption, not measured). Latencies shown are raw measurements.

djev (Maisa, diffusion-gemma). Hosted API in free preview: the cost uses djev's announced price ($0.035 per million input tokens, output free); nothing is charged yet. Open-sourcing is planned, not yet released. Probabilities are djev's own (its docs call them experimental and uncalibrated).

open-alternative-jev (Qwen3.5-4B, IkerMoel). Speed score uses x2 + 0.15 s (assumption, not measured). Latencies shown are raw measurements. With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.

system-one-open (Gemma 4 E2B LoRA on an L4). Speed score uses x2 (assumption, not measured). Latencies shown are raw measurements.

OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16). Speed score uses x2 + 0.15 s (assumption, not measured). Latencies shown are raw measurements.

openjev-sglang (Qwen3.6-35B-A3B on SGLang). Speed score uses x2 (assumption, not measured). Latencies shown are raw measurements.

open-jev-deberta-v3-large (local CPU). Speed score uses x2 + 0.15 s (assumption, not measured). Latencies shown are raw measurements.

Bespoke Nimble 9B (Bespoke Labs). Speed score uses x2 + 0.15 s (assumption, not measured). Latencies shown are raw measurements.

system-one (Qwen3-8B, Sean Goedecke). Speed score uses x2 + 0.15 s (assumption, not measured). Latencies shown are raw measurements.

Qwen3.8 27B (Chutes TEE). Partial coverage: excluded from the ranking. The published judge accuracy excludes unattempted items from its denominator. Preserved as published; request errors use failed / attempted requests.

Needle 3, options as tools (post-hoc adapter mode). Label-only output: no probability calibration; the published score treats this axis as zero. Speed score uses x2 + 0.15 s (assumption, not measured). Latencies shown are raw measurements. Partial coverage: excluded from the ranking.

Needle 3 (Cactus, 2-bit, local CPU). Label-only output: no probability calibration; the published score treats this axis as zero. Speed score uses x2 + 0.15 s (assumption, not measured). Latencies shown are raw measurements. Partial coverage: excluded from the ranking.

Source methodology and limitations

External reference measurements, not AgentSky API measurements.

The source score is a geometric mean of Intelligence, Calibration, Speed and Cost with equal weights.

Intelligence uses easy 14%, standard 28%, judge 28% and hard 30%. Coverage denominator differences in the published artifacts are disclosed per row.

Raw serial standard-and-judge latency is shown. Source speed scores adjust demo/self-hosted latency by x2, plus 150 ms on the source author's own servers; these are assumptions.

Measured cost means published tariffs multiplied by measured tokens, not independently verified invoices. Estimated and announced prices are labeled separately.

Errors are failed requests divided by attempted requests from the source outcome aggregates, excluding unattempted items. Partial runs remain unranked.

The hard tier includes 109 held-out items. A public-subset rerun would not cover the same suite.

How to read these results

JevBench tests typed decisions: a state and a bounded set of answers go in, and a choice comes out. A request about a missing package, for example, could map to track_order. These scores describe the tested model configurations.

Confidence should mean something. A well-calibrated model that says “80% sure” should be right about eight times in ten across comparable decisions. JevBench also tests whether the returned probabilities match known distributions.

The default combines the four 0–100 axes with an equal geometric mean, flooring each at 1. A weak axis pulls the score down. Partial runs are shown without a rank.

Source methodology. Intelligence weights easy, standard, judge and hard accuracy. Calibration uses the hard tier. Speed uses median and 95th-percentile latency; cost uses dollars per 1,000 decisions.

Measured and assumed. The source doubles latency for self-hosted and demo endpoints, adding 0.15 seconds on its own servers to approximate production overhead. These are assumptions. The table shows raw latency and marks estimated or announced prices.

A browser agent also depends on its harness, tools and the task. This decision benchmark does not measure browser-task completion.