Jev & decision models
The right decision.
At the right cost.
Compare how well Jev and its alternatives choose, estimate confidence and respond. Explore the trade-offs before choosing a model.
Intelligence
Accuracy across easy, standard, judge and hard decisions.
Calibration
How well confidence matches correctness and known probabilities.
Speed
Median and tail latency for one decision at a time.
Cost
The price of a thousand decisions, including pricing assumptions.
Benchmark Heaven / JevBench · v1.2.1
What matters to you?
Equal weights by default. Adjust the balance to explore your priorities.
Decision model ranking
Intelligence, calibration, speed and cost. Equal weight, one score.
- Jev 1.13.0 (TypeSafe AI)75.3
TypeSafe AI
- SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)74.6
Theodore Lee (TheoLeeCJ)
- djev (Maisa, diffusion-gemma)74.3
Maisa (David Villalón)
- 69.8
- 68.7
- OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16)67.6
razorback16 / Codiv
- 66.2
- GPT-5.6 Luna (low reasoning effort)66.0
OpenAI
Look beyond the overall score
Compare the same four axes directly. Further right is better.
Intelligence
Calibration
Speed
Cost
Every measurement
Sort a column to inspect the trade-offs. Missing measurements stay last within complete and partial runs.
| Model | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 75.3 | 90.4 | 82.7 | 83.3 | 51.7 | $0.041 | 100.0% | 99.0% | 94.5% | 74.1% | 0.65s | 0.72s | |
| 74.6 | 85.9 | 72.6 | 83.7 | 59.2 | ~$0.023estimated | 100.0% | 97.9% | 95.2% | 59.5% | 0.20s | 0.32s | |
| 74.3 | 88.4 | 65.4 | 91.4 | 57.6 | $0.026announced | 100.0% | 97.9% | 93.2% | 69.5% | 0.24s | 0.31s | |
| 69.8 | 75.6 | 63.2 | 83.5 | 59.6 | ~$0.022estimated | 100.0% | 84.4% | 74.7% | 56.8% | 0.21s | 0.32s | |
| 68.7 | 79.5 | 56.7 | 77.0 | 64.1 | ~$0.016estimated | 100.0% | 93.8% | 87.7% | 49.1% | 0.65s | 0.77s | |
| 67.6 | 86.0 | 64.8 | 83.2 | 45.2 | ~$0.067estimated | 100.0% | 95.8% | 91.1% | 65.5% | 0.24s | 0.31s | |
| 66.2 | 88.9 | 77.4 | 77.1 | 36.1 | ~$0.135estimated | 100.0% | 95.8% | 95.2% | 71.4% | 0.68s | 0.73s | |
GPT-5.6 Luna (low reasoning effort) OpenAI 534 / 534 attempted | 66.0 | 96.8 | 89.8 | 77.5 | 28.2 | $0.247 | 100.0% | 97.9% | 96.6% | 94.5% | 0.97s | 1.82s |
| 64.4 | 53.6 | 66.4 | 66.0 | 73.3 | ~$0.0077estimated | 100.0% | 49.0% | 53.4% | 36.4% | 1.77s | 3.35s | |
| 63.5 | 78.6 | 64.5 | 82.5 | 38.9 | ~$0.109estimated | 100.0% | 94.8% | 89.0% | 43.6% | 0.19s | 0.46s | |
Gemini 3.1 Flash-Lite 534 / 534 attempted | 60.8 | 90.3 | 68.1 | 81.8 | 27.1 | $0.268 | 100.0% | 99.0% | 93.2% | 75.0% | 0.76s | 0.88s |
DeepSeek V4.1 Flash (thinking default) DeepSeek 534 / 534 attempted | 58.1 | 96.1 | 96.7 | 71.6 | 17.1 | $0.579 | 98.6% | 99.0% | 93.2% | 95.0% | 1.42s | 4.89s |
| 56.5 | 80.1 | 36.8 | 84.4 | 41.2 | ~$0.092estimated | 100.0% | 90.6% | 91.8% | 50.0% | 0.17s | 0.30s | |
Qwen3.8 27B (Chutes TEE) Qwen / Chutes · Partial run 346 / 534 attempted | 25.5 | 74.6 | 92.1 | 61.3 | 0.0 | ~$2.711estimated | 98.6% | 99.0% | 95.3% | 21.4% | 5.75s | 12.97s |
| 19.1 | 39.5 | Not measured (scored 0) | 52.8 | 63.7 | ~$0.016estimated | 66.7% | 31.3% | 34.2% | — | 3.78s | 33.64s | |
| 16.7 | 22.4 | Not measured (scored 0) | 59.9 | 58.1 | ~$0.025estimated | 47.2% | 16.7% | 31.5% | 7.7% | 1.69s | 14.36s |
Configuration and coverage notes
Source methodology and limitations
How to read these results
JevBench tests typed decisions: a state and a bounded set of answers go in, and a choice comes out. A request about a missing package, for example, could map to track_order. These scores describe the tested model configurations.
Confidence should mean something. A well-calibrated model that says “80% sure” should be right about eight times in ten across comparable decisions. JevBench also tests whether the returned probabilities match known distributions.
The default combines the four 0–100 axes with an equal geometric mean, flooring each at 1. A weak axis pulls the score down. Partial runs are shown without a rank.
Source methodology. Intelligence weights easy, standard, judge and hard accuracy. Calibration uses the hard tier. Speed uses median and 95th-percentile latency; cost uses dollars per 1,000 decisions.
Measured and assumed. The source doubles latency for self-hosted and demo endpoints, adding 0.15 seconds on its own servers to approximate production overhead. These are assumptions. The table shows raw latency and marks estimated or announced prices.
A browser agent also depends on its harness, tools and the task. This decision benchmark does not measure browser-task completion.
