{"version":"v1.2.1","measuredAt":"2026-09-19T17:03:11+00:00","retrievedAt":"2026-09-19T20:02:07.684000+00:00","source":{"name":"Benchmark Heaven / JevBench","url":"https://benchmarkheaven.com/jev-models"},"sourceArtifactUrl":"https://github.com/fstandhartinger/jevbench/blob/d0a11e0511c9c54e34e59611090aee3a20753e2d/results/v1.2/jevbench-v1.2-results.json","sourceSha256":"88ab9abde5cc63a03a47102058c2e3ee6758e693cd015f7cd4109b81d3dc4052","outcomeArtifactUrl":"https://github.com/fstandhartinger/jevbench/blob/d0a11e0511c9c54e34e59611090aee3a20753e2d/results/v1.2/jevbench-v1.2-per-task.json","outcomeSha256":"1ad7d1f9256b45bc93d325b22e11442990862be8f86bd84b3d28d173fcec903d","license":"MIT","licenseUrl":"https://github.com/fstandhartinger/jevbench/blob/d0a11e0511c9c54e34e59611090aee3a20753e2d/LICENSE","decisionCount":534,"hardDecisionCount":220,"rows":[{"id":"jev-1.13.0","label":"Jev 1.13.0 (TypeSafe AI)","provider":"TypeSafe AI","url":"https://docs.typesafe.ai","kind":"jev","status":"complete","axes":{"intelligence":90.4013594852636,"calibration":82.6528888888889,"speed":83.26811926100174,"cost":51.7396879482021},"calibrationKind":"distribution","score":75.32408936881852,"usdPer1k":0.04061412193840432,"costKind":"measured","costNote":"public tariff x measured tokens (https://docs.typesafe.ai/models (output tokens not billed)) | public tariff x measured tokens (hard-tier run)","tiers":{"easy":100,"standard":98.95833333333334,"judge":94.52054794520548,"hard":74.0909090909091},"latencyMs":{"p50":652.4335257709026,"p95":722.1905551850795},"errorRate":0,"sampleCount":534},{"id":"semif-qwen3.5-4b","label":"SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)","provider":"Theodore Lee (TheoLeeCJ)","url":"https://github.com/TheoLeeCJ/openjev","kind":"jev-rebuild","status":"complete","axes":{"intelligence":85.93783727687837,"calibration":72.60139737235289,"speed":83.70422262345133,"cost":59.16231136364247},"calibrationKind":"distribution","score":74.55563524981045,"usdPer1k":0.02297504061919004,"costKind":"estimated","costNote":"ESTIMATE: hosted-provider price, deepinfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.15/M out (same weights (not on OpenRouter), as open-alternative-jev in v1.1.2) x 426 input and 1 output tokens per decision (input tokens measured) | ESTIMATE: deepinfra Qwen/Qwen3.5-4B $0.03/M in, $0.15/M out x 1244 in / 0 out tokens per hard decision","tiers":{"easy":100,"standard":97.91666666666666,"judge":95.2054794520548,"hard":59.54545454545455},"latencyMs":{"p50":197.9611478745937,"p95":315.3164997696876},"errorRate":0,"sampleCount":534,"notes":"Speed score uses x2 + 0.15 s (assumption, not measured). Latencies shown are raw measurements."},{"id":"djev","label":"djev (Maisa, diffusion-gemma)","provider":"Maisa (David Villalón)","url":"https://djev.dev","kind":"jev-rebuild","status":"complete","axes":{"intelligence":88.36249481112495,"calibration":65.41450619010376,"speed":91.35677310946092,"cost":57.57524920597589},"calibrationKind":"distribution","score":74.25567554991633,"usdPer1k":0.025951254681647943,"costKind":"announced","costNote":"ANNOUNCED PRICE (free preview): djev's docs state $0.035 per million input tokens, output tokens free (https://api.djev.dev/docs, 'Usage & credits'; prepaid billing not yet switched on, 19 Sep 2026, so nothing was charged) x measured input tokens (741 per decision on average over all 534 decisions)","tiers":{"easy":100,"standard":97.91666666666666,"judge":93.15068493150685,"hard":69.54545454545455},"latencyMs":{"p50":237.0578795671463,"p95":308.65143015980715},"errorRate":0,"sampleCount":534,"notes":"Hosted API in free preview: the cost uses djev's announced price ($0.035 per million input tokens, output free); nothing is charged yet. Open-sourcing is planned, not yet released. Probabilities are djev's own (its docs call them experimental and uncalibrated)."},{"id":"open-alternative-jev","label":"open-alternative-jev (Qwen3.5-4B, IkerMoel)","provider":"IkerMoel","url":"https://github.com/ikermoel/open-alternative-jev","kind":"jev-rebuild","status":"complete","axes":{"intelligence":75.57456413449563,"calibration":63.16501249259078,"speed":83.47780372641819,"cost":59.62620384142349},"calibrationKind":"distribution","score":69.817630118622,"usdPer1k":0.022171404494382024,"costKind":"estimated","costNote":"ESTIMATE: hosted-provider price, deepinfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.15/M out (as open-alternative-jev) x 383 input and 1 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) | ESTIMATE: deepinfra Qwen/Qwen3.5-4B $0.03/M in, $0.15/M out x 1235 in / 1 out tokens per hard decision","tiers":{"easy":100,"standard":84.375,"judge":74.65753424657534,"hard":56.81818181818182},"latencyMs":{"p50":206.86038956046104,"p95":323.2223108410835},"errorRate":0,"sampleCount":534,"notes":"Speed score uses x2 + 0.15 s (assumption, not measured). Latencies shown are raw measurements. With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order."},{"id":"system-one-open","label":"system-one-open (Gemma 4 E2B LoRA on an L4)","provider":"mithalouni","url":"https://github.com/mithalouni/system-one-open","kind":"jev-rebuild","status":"complete","axes":{"intelligence":79.52521793275218,"calibration":56.699696211408835,"speed":76.96012500732209,"cost":64.1349167978843},"calibrationKind":"distribution","score":68.68493237922856,"usdPer1k":0.015685659145076917,"costKind":"estimated","costNote":"ESTIMATE: hosted-provider price, deepinfra google/gemma-4-E4B-it list price $0.02/M in, $0.1/M out (Gemma 4 E2B is not listed; the nearest larger sibling, Gemma 4 E4B, is listed only on DeepInfra) x 452 input and 2 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) | ESTIMATE: deepinfra google/gemma-4-E4B-it $0.02/M in, $0.1/M out x 1235 in / 2 out tokens per hard decision","tiers":{"easy":100,"standard":93.75,"judge":87.67123287671232,"hard":49.09090909090909},"latencyMs":{"p50":651.7308317124844,"p95":772.4301926791667},"errorRate":0,"sampleCount":534,"notes":"Speed score uses x2 (assumption, not measured). Latencies shown are raw measurements."},{"id":"openjev-razorback16","label":"OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16)","provider":"razorback16 / Codiv","url":"https://github.com/razorback16/openjev","kind":"jev-rebuild","status":"complete","axes":{"intelligence":85.97654628476548,"calibration":64.76011611808524,"speed":83.1778984786352,"cost":45.18071405399585},"calibrationKind":"distribution","score":67.63354664009132,"usdPer1k":0.0671907566081966,"costKind":"estimated","costNote":"ESTIMATE: hosted-provider price, openrouter google/gemma-4-26b-a4b-it list price $0.09/M in, $0.3/M out (DiffusionGemma 26B-A4B is not listed; the same-size Gemma 4 26B-A4B MoE sibling is (size class moe_26B-A4B)) x 410 input and 1 output tokens per decision (input tokens measured) | ESTIMATE: openrouter google/gemma-4-26b-a4b-it $0.09/M in, $0.3/M out x 1222 in / 0 out tokens per hard decision","tiers":{"easy":100,"standard":95.83333333333334,"judge":91.0958904109589,"hard":65.45454545454545},"latencyMs":{"p50":241.27069488167763,"p95":305.2692499011755},"errorRate":0,"sampleCount":534,"notes":"Speed score uses x2 + 0.15 s (assumption, not measured). Latencies shown are raw measurements."},{"id":"openjev-sglang","label":"openjev-sglang (Qwen3.6-35B-A3B on SGLang)","provider":"ekzhang","url":"https://github.com/ekzhang/openjev-sglang","kind":"jev-rebuild","status":"complete","axes":{"intelligence":88.8999584889996,"calibration":77.41176873901938,"speed":77.059799513554,"cost":36.12714266404304},"calibrationKind":"distribution","score":66.15954494186164,"usdPer1k":0.134615554522674,"costKind":"estimated","costNote":"ESTIMATE: hosted-provider price, openrouter qwen/qwen3.6-35b-a3b list price $0.1/M in, $0.9/M out (same base weights) x 667 input and 2 output tokens per decision | ESTIMATE: openrouter qwen/qwen3.6-35b-a3b $0.1/M in, $0.9/M out x 2272 in / 2 out tokens per hard decision","tiers":{"easy":100,"standard":95.83333333333334,"judge":95.2054794520548,"hard":71.36363636363636},"latencyMs":{"p50":677.6718497276306,"p95":726.0066717863083},"errorRate":0,"sampleCount":534,"notes":"Speed score uses x2 (assumption, not measured). Latencies shown are raw measurements."},{"id":"gpt-5.6-luna","label":"GPT-5.6 Luna (low reasoning effort)","provider":"OpenAI","kind":"llm-baseline","status":"complete","axes":{"intelligence":96.821398920714,"calibration":89.79207658196586,"speed":77.5464346646158,"cost":28.204007179931185},"calibrationKind":"distribution","score":66.03444089702963,"usdPer1k":0.24728613154420276,"costKind":"measured","costNote":"public tariff x measured tokens (https://platform.openai.com/docs/pricing (standard tier, read 2026-09-19)) | public tariff x measured tokens (hard-tier run)","tiers":{"easy":100,"standard":97.91666666666666,"judge":96.57534246575342,"hard":94.54545454545456},"latencyMs":{"p50":968.0032916367054,"p95":1817.5220962613816},"errorRate":0,"sampleCount":534},{"id":"open-jev-deberta-v3-large","label":"open-jev-deberta-v3-large (local CPU)","provider":"Kotoba Labs","url":"https://github.com/kotoba-lang/typed-decisions","kind":"jev-rebuild","status":"complete","axes":{"intelligence":53.57632835201329,"calibration":66.3822600310561,"speed":65.97921316871283,"cost":73.3330089966802},"calibrationKind":"distribution","score":64.40697539318792,"usdPer1k":0.007742829572538459,"costKind":"estimated","costNote":"ESTIMATE: hosted-provider price, deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) | ESTIMATE: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) $0.01/M in, $0.0/M out x 1235 in / 0 out tokens per hard decision","tiers":{"easy":100,"standard":48.95833333333333,"judge":53.42465753424658,"hard":36.36363636363637},"latencyMs":{"p50":1767.673410475254,"p95":3349.2883060127488},"errorRate":0.003745318352059925,"sampleCount":534,"notes":"Speed score uses x2 + 0.15 s (assumption, not measured). Latencies shown are raw measurements."},{"id":"nimble-9b","label":"Bespoke Nimble 9B (Bespoke Labs)","provider":"Bespoke Labs","url":"https://github.com/bespokelabsai/nimble","kind":"jev-rebuild","status":"complete","axes":{"intelligence":78.56408260689082,"calibration":64.51395363452104,"speed":82.5436896273356,"cost":38.93708830042782},"calibrationKind":"distribution","score":63.53035216319952,"usdPer1k":0.10850016283992836,"costKind":"estimated","costNote":"ESTIMATE: hosted-provider price, openrouter qwen/qwen3.5-9b list price $0.1/M in, $0.15/M out (a LoRA merge of Qwen3.5-9B; the base weights are listed on OpenRouter (size class dense_9B)) x 990 input and 1 output tokens per decision (input tokens measured) | ESTIMATE: openrouter qwen/qwen3.5-9b $0.1/M in, $0.15/M out x 1215 in / 2 out tokens per hard decision","tiers":{"easy":100,"standard":94.79166666666666,"judge":89.04109589041096,"hard":43.63636363636363},"latencyMs":{"p50":185.20960584282875,"p95":459.8693612962959},"errorRate":0.1348314606741573,"sampleCount":534,"notes":"Speed score uses x2 + 0.15 s (assumption, not measured). Latencies shown are raw measurements."},{"id":"gemini-3.1-flash-lite","label":"Gemini 3.1 Flash-Lite","provider":"Google","kind":"llm-baseline","status":"complete","axes":{"intelligence":90.29052511415526,"calibration":68.08811278368097,"speed":81.78762581772229,"cost":27.14615397142829},"calibrationKind":"distribution","score":60.78233030734062,"usdPer1k":0.26820170492819223,"costKind":"measured","costNote":"public tariff x measured tokens (https://ai.google.dev/gemini-api/docs/pricing (paid tier, read 2026-09-19)) | public tariff x measured tokens (hard-tier run)","tiers":{"easy":100,"standard":98.95833333333334,"judge":93.15068493150685,"hard":75},"latencyMs":{"p50":756.230715662241,"p95":876.1593606323003},"errorRate":0,"sampleCount":534},{"id":"deepseek-flash","label":"DeepSeek V4.1 Flash (thinking default)","provider":"DeepSeek","kind":"llm-baseline","status":"complete","axes":{"intelligence":96.0960806697108,"calibration":96.66517503108454,"speed":71.59823070734842,"cost":17.123750085891416},"calibrationKind":"distribution","score":58.092387016557666,"usdPer1k":0.5788175142273461,"costKind":"measured","costNote":"public tariff x measured tokens (https://api-docs.deepseek.com/quick_start/pricing (cache-miss off-peak; the run is on a Saturday, off-peak all day)) | public tariff x measured tokens (hard-tier run)","tiers":{"easy":98.61111111111111,"standard":98.95833333333334,"judge":93.15068493150685,"hard":95},"latencyMs":{"p50":1416.3860343396664,"p95":4886.470635980367},"errorRate":0.031835205992509365,"sampleCount":534},{"id":"system-one-sg","label":"system-one (Qwen3-8B, Sean Goedecke)","provider":"Sean Goedecke","url":"https://github.com/sgoedecke/system-one","kind":"jev-rebuild","status":"complete","axes":{"intelligence":80.07363013698631,"calibration":36.764309957143894,"speed":84.3621367799104,"cost":41.15248486040298},"calibrationKind":"distribution","score":56.54118326474669,"usdPer1k":0.09153429437798076,"costKind":"estimated","costNote":"ESTIMATE: hosted-provider price, openrouter qwen/qwen3-8b list price $0.117/M in, $0.455/M out (same weights, listed on OpenRouter) x 443 input and 1 output tokens per decision (input tokens measured) | ESTIMATE: openrouter qwen/qwen3-8b $0.117/M in, $0.455/M out x 1258 in / 1 out tokens per hard decision","tiers":{"easy":100,"standard":90.625,"judge":91.78082191780824,"hard":50},"latencyMs":{"p50":166.1309413611889,"p95":304.72867079079145},"errorRate":0,"sampleCount":534,"notes":"Speed score uses x2 + 0.15 s (assumption, not measured). Latencies shown are raw measurements."},{"id":"qwen3.8-27b","label":"Qwen3.8 27B (Chutes TEE)","provider":"Qwen / Chutes","kind":"llm-baseline","status":"partial","axes":{"intelligence":74.60014515231052,"calibration":92.09887121241668,"speed":61.27048999578096,"cost":0},"calibrationKind":"distribution","score":25.47189954757361,"usdPer1k":2.7110372815546095,"costKind":"estimated","costNote":"ESTIMATE: hosted-provider price, openrouter qwen/qwen3.8-27b list price $0.214/M in, $2.55/M out (same weights; our run used a flat-rate Chutes subscription) x 445 input and 393 output tokens per decision | ESTIMATE: openrouter qwen/qwen3.8-27b $0.214/M in, $2.55/M out x 1592 in / 1833 out tokens per hard decision","tiers":{"easy":98.61111111111111,"standard":98.95833333333334,"judge":95.2755905511811,"hard":21.363636363636363},"latencyMs":{"p50":5753.906108438969,"p95":12971.44114784895},"errorRate":0.02601156069364162,"sampleCount":346,"notes":"Partial coverage: excluded from the ranking. The published judge accuracy excludes unattempted items from its denominator. Preserved as published; request errors use failed / attempted requests."},{"id":"needle-3-tools","label":"Needle 3, options as tools (post-hoc adapter mode)","provider":"Cactus Compute","url":"https://github.com/cactus-compute/needle","kind":"small-tool-model","status":"partial","axes":{"intelligence":39.53196347031963,"calibration":0,"speed":52.842525649949984,"cost":63.698846264524306},"calibrationKind":"label-only","score":19.09923092164542,"usdPer1k":0.016219537190082647,"costKind":"estimated","costNote":"ESTIMATE: same per-token price as Needle 3 (openrouter meta-llama/llama-3.2-1b-instruct $0.027/M in, $0.201/M out) x 452 input and 20 output tokens per decision, over the 314 easy/standard/judge decisions it ran (no hard-tier run). The v1.2 score lab had no price for this row and scored it 100; fixed.","tiers":{"easy":66.66666666666666,"standard":31.25,"judge":34.24657534246575,"hard":null},"latencyMs":{"p50":3778.900783509016,"p95":33637.18595951795},"errorRate":0,"sampleCount":314,"notes":"Label-only output: no probability calibration; the published score treats this axis as zero. Speed score uses x2 + 0.15 s (assumption, not measured). Latencies shown are raw measurements. Partial coverage: excluded from the ranking."},{"id":"needle-3","label":"Needle 3 (Cactus, 2-bit, local CPU)","provider":"Cactus Compute","url":"https://github.com/cactus-compute/needle","kind":"small-tool-model","status":"partial","axes":{"intelligence":22.417877404178775,"calibration":0,"speed":59.92368509307548,"cost":58.10061053357013},"calibrationKind":"label-only","score":16.714501353540964,"usdPer1k":0.02492563984585384,"costKind":"estimated","costNote":"ESTIMATE: hosted-provider price, openrouter meta-llama/llama-3.2-1b-instruct list price $0.027/M in, $0.201/M out (no generative model under 1B is listed; the smallest listed one (1B) errs high; about 20 generated tokens for one tool call) x 452 input and 20 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) | ESTIMATE: openrouter meta-llama/llama-3.2-1b-instruct $0.027/M in, $0.201/M out x 1235 in / 20 out tokens per hard decision","tiers":{"easy":47.22222222222222,"standard":16.666666666666664,"judge":31.506849315068493,"hard":7.727272727272727},"latencyMs":{"p50":1687.0602630078793,"p95":14364.453018829225},"errorRate":0,"sampleCount":358,"notes":"Label-only output: no probability calibration; the published score treats this axis as zero. Speed score uses x2 + 0.15 s (assumption, not measured). Latencies shown are raw measurements. Partial coverage: excluded from the ranking."}],"notes":["External reference measurements, not AgentSky API measurements.","The source score is a geometric mean of Intelligence, Calibration, Speed and Cost with equal weights.","Intelligence uses easy 14%, standard 28%, judge 28% and hard 30%. Coverage denominator differences in the published artifacts are disclosed per row.","Raw serial standard-and-judge latency is shown. Source speed scores adjust demo/self-hosted latency by x2, plus 150 ms on the source author's own servers; these are assumptions.","Measured cost means published tariffs multiplied by measured tokens, not independently verified invoices. Estimated and announced prices are labeled separately.","Errors are failed requests divided by attempted requests from the source outcome aggregates, excluding unattempted items. Partial runs remain unranked.","The hard tier includes 109 held-out items. A public-subset rerun would not cover the same suite."]}