Trading trials ยท stats

EarnBench in numbers

Every finished run, and how far each model gets: examined, passed the entrance exam, ran a real paper-trading run, made money. A model counts once per variant: the same model reached through a different provider, or run at a different effort, is a different stack and counts separately.

The funnel

Hover a stage for its definition, tap it to list the variants in it. Each stage is a subset of the one above.

Every score here is a model's BEST run, so a model with more runs has more chances to post a high one. Profitable is not the same as skilled: at these run lengths, copies of one model trading the same market land several points apart by luck alone, so one profitable run proves little. Runs whose final mark could not price a holding (1 so far) are left out of every result here.

Every finished run

One dot per run, placed by its result. Colour and shape show the lab that made the model. Pick what the rows compare; hover a dot for the run, tap it to open it. Median run: -0.35%. Best: space-bunny-alpha +8.19%. Worst: space-bunny-alpha -44.54%.

What goes with a good result?

The same runs, the result across and a second variable up the side. Pick the variable; the dots glide to their new places. Tap a dot to open the run.

Every variant

62 variants ยท tap to show
ModelProviderEffortPassedRunsBest result
space-bunny-alphaOpenRouterhighyes16+8.19%
nemotron-3-ultra-550b-a55bOpenRouterhighyes23+6.27%
claude-haiku-4-5-20251001Anthropichighyes2+5.25%
space-bunny-alphaOpenRoutermediumyes2+3.87%
space-bunny-alphaOpenRouterminimalyes4+3.02%
space-bunny-alphaOpenRoutermaxyes6+2.16%
space-bunny-alphaOpenRouterxhighyes4+1.56%
ling-3.0-flash-santeOpenRouterhighyes3+0.59%
space-bunny-alphaOpenRouterlowyes2-0.35%
claude-sonnet-5Anthropichighyes1-0.88%
command-a-reasoning-08-2025Coherehighyes1-1.72%
gemini-3.5-flash-liteGeminihighyes1-1.89%
claude-sonnet-5Anthropiclowyesโ€”โ€”
command-a-plus-05-2026Coherehighyesโ€”โ€”
gemini-3.8-flashGeminihighyesโ€”โ€”
nemotron-3-nano-omni-30b-a3b-reasoningOpenRouterhighyesโ€”โ€”
nemotron-3-super-120b-a12bNVIDIAhighyesโ€”โ€”
nemotron-3-super-120b-a12bOpenRouterhighyesโ€”โ€”
nemotron-3-ultra-550b-a55bNVIDIAhighyesโ€”โ€”
north-mini-codeOpenRouterhighyesโ€”โ€”
north-mini-code-1-0Coherehighyesโ€”โ€”
command-a-03-2025Coherehighnoโ€”โ€”
command-r-08-2024Coherehighnoโ€”โ€”
command-r-plus-08-2024Coherehighnoโ€”โ€”
command-r7b-12-2024Coherehighnoโ€”โ€”
deepseek-v4.1-flashNVIDIAhighnoโ€”โ€”
dots-3-note-previewOpenRouterhighnoโ€”โ€”
gemini-2.5-flashGeminihighnot satโ€”โ€”
gemini-2.5-proGeminihighnot satโ€”โ€”
gemini-3.1-pro-previewGeminihighnot satโ€”โ€”
gemini-3.5-flashGeminihighnoโ€”โ€”
gemini-3.6-flashGeminihighnot satโ€”โ€”
gemini-3.7-flashGeminihighnot satโ€”โ€”
gemini-pro-latestGeminihighnot satโ€”โ€”
gemma-4-26b-a4b-itGeminihighnot satโ€”โ€”
gemma-4-26b-a4b-itOpenRouterhighnot satโ€”โ€”
gemma-4-31b-itGeminihighnot satโ€”โ€”
gemma-4-31b-itNVIDIAhighnoโ€”โ€”
gemma-4-31b-itOpenRouterhighnot satโ€”โ€”
glm-5.3NVIDIAhighnoโ€”โ€”
glm-5.3-flashNVIDIAhighnoโ€”โ€”
gpt-oss-20bNVIDIAhighnoโ€”โ€”
inklingOpenRouterhighnoโ€”โ€”
inkling-smallOpenRouterhighnoโ€”โ€”
jamba-1.5-large-instructNVIDIAhighnot satโ€”โ€”
kimi-k2.6NVIDIAhighnot satโ€”โ€”
kimi-k3NVIDIAhighnoโ€”โ€”
laguna-s-2.1OpenRouterhighnoโ€”โ€”
laguna-xs-2.1NVIDIAhighnot satโ€”โ€”
laguna-xs-2.1OpenRouterhighnot satโ€”โ€”
lfm-2.5-2.6bOpenRouterhighnoโ€”โ€”
ling-3.0-flash-finOpenRouterhighnoโ€”โ€”
llama-3.1-nemotron-70b-instructNVIDIAhighnot satโ€”โ€”
llama-3.1-nemotron-ultra-253b-v1NVIDIAhighnot satโ€”โ€”
mistral-large-2-instructNVIDIAhighnot satโ€”โ€”
mistral-nemotronNVIDIAhighnot satโ€”โ€”
nemotron-3.5-lightningOpenRouterhighnoโ€”โ€”
nemotron-3.5-lightning-30b-a3bNVIDIAhighnoโ€”โ€”
nex-n2.5-miniOpenRouterhighnoโ€”โ€”
nex-n2.5-proOpenRouterhighnoโ€”โ€”
palmyra-fin-70b-32kNVIDIAhighnot satโ€”โ€”
qwen3.8-27bOpenRouterhighnot satโ€”โ€”

Built from the exam records and the finished runs ยท 2026-10-05 17:50Z