Skip to main content

Crate cerno_bench

Crate cerno_bench 

Source
Expand description

Measures candidate models against a labelled dataset and writes the comparison table.

Four things decide a model here, and the first one is a gate rather than a score:

  1. Label fidelity — was the model’s most likely first token one of the letters it was offered? A model that writes prose instead is unusable for cerno at any accuracy, even when a letter turns up further down the ranking and an answer can still be read off it.
  2. Accuracy — did it pick the right letter.
  3. Latency — the whole point of one forward pass.
  4. Calibration — how far its confidence has to be flattened to stop lying.

Usage: cerno-bench –models a,b,c [–reference m] [–dataset p] [–out p] [–host url] [–host-kind ollama|openai|vllm|llamacpp|lmstudio]

An API key for the OpenAI-compatible hosts comes from CERNO_HOST_API_KEY.

Structs§

Case 🔒
Dataset 🔒
Outcome 🔒
One model’s result on one case.
Report 🔒

Functions§

arg 🔒
brier 🔒
Mean squared error between the predicted yes-probability and the truth. Noul cases only, since Brier is defined on a binary outcome.
calibration_reading 🔒
What a best-fit temperature says about the model, in words that stay true for any value.
fit_temperature 🔒
Grid-search the temperature that best explains the labelled outcomes.
main 🔒
markdown_table 🔒
Render a markdown table with padded cells, the first column left-aligned and the rest right.
pct 🔒
percentile 🔒
pick_winner 🔒
The best candidate: highest accuracy among models whose top token was an offered label in every case, ties broken by median latency. Fidelity is a gate first — a model cerno cannot read is not a candidate at any accuracy — and the reference is excluded, since it is the yardstick rather than an option.
render 🔒
run_case 🔒
softmax_of 🔒
summarise 🔒
top_token_is_a_label 🔒
Whether the highest-ranked token spells one of offered. tokens is ranked, highest first.