← All articles

RESEARCH / MULTIMODAL DECISIONS

CMDB-1500Multimodal Decision Model Benchmark

Comprehensive Multimodal Decision Benchmark 1500

1,200 text questions and 300 image questions, measuring decisions and probability estimates.

Filter
SPX-CDJevGeneral reasoning modelsOther decision models

View exact text results
CMDB-1500 Text results; fixed total 1,200, valid answers counted over the full evaluated scope
RankModel / versionAccuracy · 0–100%Correct / fixed totalValid / evaluated scope
01GPT 6.1 Sol · low80.00%960 / 1,2001,500 / 1,500
02SPX-CD-Pro (2026-10-04)79.33%952 / 1,2001,500 / 1,500
03Claude Sonnet 5.5 · low78.00%936 / 1,2001,500 / 1,500
04StartLux-Decision-27B · BF1677.92%935 / 1,2001,500 / 1,500
05SPX-CD-Flash77.58%931 / 1,2001,500 / 1,500
06GPT 6 Luna · low76.58%919 / 1,2001,500 / 1,500
07GLM 5.3 Flash · low75.75%909 / 1,2001,500 / 1,500
08StartLux-Decision-35B-A3B · BF1675.67%908 / 1,2001,500 / 1,500
09SPX-CD-Omni74.67%896 / 1,2001,500 / 1,500
10Mercury Decide73.92%887 / 1,2001,200 / 1,200
11pplx-decider 27B72.83%874 / 1,2001,500 / 1,500
12Jev72.17%866 / 1,2001,200 / 1,200
13Cloudflare Clef71.50%858 / 1,2001,485 / 1,500
14JPT · 9B71.33%856 / 1,2001,500 / 1,500
15Qwen 3.8 Flash71.08%853 / 1,2001,500 / 1,500
16Kimi K2.671.00%852 / 1,2001,500 / 1,500
17Mapika · Decider 4B v2.169.83%838 / 1,2001,200 / 1,200
18Cloudflare Clef-Flash69.67%836 / 1,2001,485 / 1,500
19Imajev-9B68.83%826 / 1,2001,478 / 1,500
20Kev · 9B68.67%824 / 1,2001,200 / 1,200
21DeepSeek V4.1 Flash68.58%823 / 1,2001,500 / 1,500
22JPT · 4B67.67%812 / 1,2001,500 / 1,500
23Decision Lux · 9B67.42%809 / 1,2001,200 / 1,200
24Jet · v6.2 4B66.92%803 / 1,2001,200 / 1,200
25Open-Jev · Zefan 9B66.67%800 / 1,2001,199 / 1,200
26Nimble · v2 9B66.58%799 / 1,2001,200 / 1,200
27Hopper G · 4B v1.365.92%791 / 1,2001,200 / 1,200
28Malkuth · 4B65.58%787 / 1,2001,200 / 1,200
29Hopper · 4B65.50%786 / 1,2001,200 / 1,200
30Winnow · 12B Q865.33%784 / 1,2001,410 / 1,500
31JevK564.83%778 / 1,2001,200 / 1,200
32Lev · 4B64.58%775 / 1,2001,200 / 1,200
33Mica · 4B BF16 GGUF64.25%771 / 1,2001,200 / 1,200
34Kev · 4B64.17%770 / 1,2001,200 / 1,200
35Jev-Omni · 12B63.58%763 / 1,2001,493 / 1,500
36APUS OpenJev · 9B 3000 High61.42%737 / 1,2001,085 / 1,200
37OpenDecider · Small61.25%735 / 1,2001,200 / 1,200
38APUS OpenJev · 9B 5949 High61.08%733 / 1,2001,085 / 1,200
39OpenDecider · Small TD60.67%728 / 1,2001,200 / 1,200
40APUS OpenJev · 4B 5949 High60.08%721 / 1,2001,085 / 1,200
41Ateve Jev59.75%717 / 1,2001,087 / 1,200
42Bocha Jev58.92%707 / 1,2001,087 / 1,200
43Winnow · E4B Q858.92%707 / 1,2001,410 / 1,500
44Reflex · Stable 4B58.50%702 / 1,2001,387 / 1,500
45Plumb · v5.2 4B58.08%697 / 1,2001,085 / 1,200
46FLock THIS/THAT57.67%692 / 1,2001,156 / 1,200
47Kev · 0.8B51.75%621 / 1,2001,200 / 1,200
48Bosun v3.1 · 1.7B49.75%597 / 1,2001,200 / 1,200
49OpenDecider · Nano46.75%561 / 1,2001,143 / 1,200
50CalDec · Laya46.08%553 / 1,2001,189 / 1,200
51Bosun v3.1 · 0.6B45.58%547 / 1,2001,200 / 1,200
52Laya · Typed Decisions44.58%535 / 1,2001,189 / 1,200
53CalDec · GLiNER43.17%518 / 1,2001,200 / 1,200
54Laya · English41.00%492 / 1,2001,189 / 1,200
55Laya · Multilingual40.25%483 / 1,2001,192 / 1,200
56GLiNER2.5-Decide38.58%463 / 1,2001,200 / 1,200
57Julia-135.75%429 / 1,2001,079 / 1,200
58CLM v0.1 · 8B34.92%419 / 1,2001,200 / 1,200

The multimodal filter shows models evaluated on image questions.

Pooled ECE calibration leaderboard

Taller bars mean lower error ↑
View exact ECE results
Complete ECE results
RankModel / versionECE ↓
01SPX-CD-Pro2.46%
02SPX-CD-Flash2.49%
03StartLux-Decision-27B · BF162.65%
04StartLux-Decision-35B-A3B · BF162.71%
05pplx-decider 27B3.08%
06Reflex · Stable 4B3.33%
07JPT · 9B3.47%
08Plumb · v5.2 4B3.53%
09Imajev-9B3.74%
10GLiNER2.5-Decide3.98%
11Mapika · Decider 4B v2.14.08%
12Kev · 9B4.09%
13OpenDecider · Small TD4.21%
14Decision Lux · 9B4.36%
15Malkuth · 4B4.43%
16Winnow · E4B Q84.56%
17OpenDecider · Small4.58%
18Laya · Typed Decisions4.86%
19Jev4.99%
20OpenDecider · Nano5.00%
21JevK55.05%
22SPX-CD-Omni5.11%
23Kev · 0.8B5.27%
24Nimble · v2 9B5.28%
25Kev · 4B5.63%
26JPT · 4B5.97%
27Ateve Jev6.09%
28Hopper · 4B6.23%
29Bocha Jev6.57%
30Lev · 4B6.78%
31Cloudflare Clef-flash7.90%
32Jet · v6.2 4B7.91%
33Jev-Omni · 12B8.08%
34Open-Jev · Zefan 9B8.62%
35Mica · 4B BF16 GGUF8.66%
36Winnow · 12B Q89.07%
37CalDec · Laya9.24%
38Hopper G · 4B v1.39.29%
39Cloudflare Clef9.40%
40CalDec · GLiNER9.40%
41Bosun v3.1 · 0.6B10.67%
42Mercury Decide10.77%
43APUS OpenJev · 9B 3000 High12.68%
44Bosun v3.1 · 1.7B13.01%
45Laya · English14.42%
46APUS OpenJev · 4B 5949 High14.86%
47APUS OpenJev · 9B 5949 High15.45%
48Laya · Multilingual19.74%
49FLock THIS/THAT30.09%
50CLM v0.1 · 8B31.16%
51Julia-140.55%

ECE uses 10 probability bins and valid choice and yes/no responses with native probabilities. Labels show the actual error. Text and image coverage varies by model.

I/C index leaderboard

0–100; higher is better ↑

View exact I/C results
Complete I/C results
RankModel / versionI ↑C ↑
01SPX-CD-Pro65.4183.58
02Mercury Decide64.5474.53
03StartLux-Decision-27B · BF1664.4883.63
04Winnow · 12B Q860.6068.72
05StartLux-Decision-35B-A3B · BF1660.0582.74
06JPT · 9B59.7083.87
07Cloudflare Clef59.4279.68
08Cloudflare Clef-flash57.4682.20
09pplx-decider 27B57.3680.56
10Plumb · v5.2 4B55.9073.01
11Jev55.4076.52
12Imajev-9B54.7682.45
13Open-Jev · Zefan 9B54.1778.80
14SPX-CD-Flash53.7983.25
15APUS OpenJev · 4B 5949 High53.6458.56
16APUS OpenJev · 9B 3000 High53.0464.91
17Malkuth · 4B52.3081.90
18Jev-Omni · 12B52.3064.26
19JPT · 4B52.2679.91
20APUS OpenJev · 9B 5949 High51.7957.45
21Jet · v6.2 4B51.4080.01
22Ateve Jev48.0775.22
23Hopper G · 4B v1.347.6874.10
24Hopper · 4B47.6179.12
25FLock THIS/THAT47.5943.08
26JevK546.7679.66
27SPX-CD-Omni46.6680.19
28Lev · 4B44.6672.91
29Decision Lux · 9B44.1578.17
30Kev · 4B44.0279.79
31Mapika · Decider 4B v2.143.9479.68
32Bocha Jev42.4673.37
33Kev · 9B41.5980.70
34Nimble · v2 9B39.7683.89
35Mica · 4B BF16 GGUF39.6672.19
36OpenDecider · Small39.4181.38
37Winnow · E4B Q834.4876.40
38OpenDecider · Small TD33.9879.02
39Reflex · Stable 4B31.9777.79
40Bosun v3.1 · 1.7B29.5466.11
41Laya · English25.2567.96
42Bosun v3.1 · 0.6B22.3067.78
43OpenDecider · Nano19.8276.75
44Laya · Typed Decisions18.6878.89
45Laya · Multilingual18.4353.67
46CalDec · Laya15.8578.49
47Kev · 0.8B13.4574.45
48Julia-110.3427.00
49CLM v0.1 · 8B5.0155.44
50CalDec · GLiNER0.0067.28
51GLiNER2.5-Decide0.0074.53

I · Intelligence: measures decision accuracy relative to random guessing.C · Calibration: measures how well predicted probabilities match the correct answers.

I/C uses the same 900 text questions for every model, applying the JevBench v1.5 formulas to CMDB results. ECE and I/C compare models that provide native candidate probabilities.

Add your model

Want your model evaluated on CMDB-1500?

Contact us ↗

CMDB-1500 combines language, image, and action selection tasks in four formats: single choice, yes/no, ordered scoring, and multi-select. The dataset contains 1,500 questions and 322 distinct source images.

What makes up the 1,500 questions?
1,500All questions

Select a category to see its share; select it again to reset.

Questions are sampled across domains, with answer options and source images preserved. Accuracy is calculated over 1,200 text questions, 300 image questions, or all 1,500 questions. Refusals and missing answers count as incorrect. Models evaluated only on text appear in the text leaderboard.

The tasks cover language understanding, business decisions, image understanding, tool use, and action selection.

Language, knowledge & judgment450 questions

Questions from JevBench test knowledge, intent recognition, evidence verification, language understanding, sentiment, and response quality.

ARC-Challenge · MMLU · Banking77 · CLINC150 · MASSIVE · BoolQ · MNLI · ChaosNLI · FEVER · PAWS · SST-5 · STS-B · StrategyQA · Civil Comments · Measuring Hate Speech · SMS Spam · HelpSteer2

Professional & business decisions300 questions

Questions from Atlan Decision Bench test tool selection, task completion checks, citation verification, SQL selection, and business routing.

The collection includes tasks adapted from BFCL, AgentDojo, τ-bench, MT-Bench, and Spider. Models choose an answer from the supplied context.

Visual decisions300 questions

Thirty questions from each of ten visual benchmarks test visual mathematics, image and text understanding, scientific diagrams, and fashion classification. Questions that require multiple images retain all source images.

MathVista · MMMU · MMMU-Pro · AI2D · ScienceQA · Fashion200k · GQA · VQAv2 · TextVQA · VizWiz

GQA, VQAv2, TextVQA, and VizWiz contribute yes/no question subsets.

Safety & adversarial tasks100 questions

Questions from Aegis, PhishNChips, and JevAdvBench test content safety, phishing email and URL detection, and decisions under adversarial perturbations.

Adversarial questions retain their source labels, some of which are not human annotations.

Game actions100 questions

Choose an action from a given game state. OpenJev provides scenarios from racing, Snake, tic-tac-toe, platformers, runner games, and ViZDoom.

Scores measure agreement with reference actions at fixed game states.

Multi-select decisions100 questions

Fifty questions each from MultiRC and GoEmotions ask models to select all correct reading-comprehension answers or identify multiple emotions in a text.

Embodied actions50 questions

ALFRED , ALFWorld, and VirtualHome provide 50 questions about choosing the next action from a task, observations, and action history.

Reference answers come from expert trajectories and evaluate individual action choices.

Chinese transcript selection50 questions

Select the transcript closest to the reference from ChineseHP AISHELL-4 candidates. Answers are determined by character edit distance.

Inputs contain transcript candidates, without audio.

Tool use25 questions

Adapted from BFCL v3, these questions ask whether a user request requires a tool call.

Spatial & complex rules25 questions

Nineteen complex-rule questions and six spatial questions from THIS-THAT test decisions under conflicting rules and spatial relationships. Reference answers come from rules or simulators.

Surd AI

SPX-CD · Native decision models

SPX-CD-Pro, Flash, and Omni return structured choices and candidate probabilities. All three are evaluated with effort=2.

SPX-CD-Omni is open source on Hugging Face, with LoRA weights, inference code, and evaluation results.

TypeSafeStartLuxHugging Face

Other native decision models

Jev, StartLux, Mercury Decide, pplx-decider, JPT, Winnow, and Kev use decision or probability interfaces. The leaderboard distinguishes model sizes, versions, and quantization formats.

GPTClaudeGLM

General reasoning models

GPT, Claude, GLM, Qwen, Kimi, and DeepSeek are included in the accuracy comparison as general reasoning models.

SPX-CD-Pro · Overall #281.20%1,218 / 1,500 correct
SPX-CD-Flash · Overall #579.53%1,193 / 1,500 correct

Text and image accuracy are 79.33% / 88.67% for Pro and 77.58% / 87.33% for Flash.

When a model assigns 80% probability to a set of answers, are about 80% actually correct? ECE measures the gap between predicted confidence and observed accuracy; lower is better. Pooled ECE is 2.46% for Pro and 2.49% for Flash.

Explore CMDB-1500 on Hugging Face ↗