Expand description
Measures candidate models against a labelled dataset and writes the comparison table.
Four things decide a model here, and the first one is a gate rather than a score:
- Label fidelity — was the model’s most likely first token one of the letters it was offered? A model that writes prose instead is unusable for cerno at any accuracy, even when a letter turns up further down the ranking and an answer can still be read off it.
- Accuracy — did it pick the right letter.
- Latency — the whole point of one forward pass.
- Calibration — how far its confidence has to be flattened to stop lying.
Usage: cerno-bench –models a,b,c [–reference m] [–dataset p] [–out p] [–host url] [–host-kind ollama|openai|vllm|llamacpp|lmstudio]
An API key for the OpenAI-compatible hosts comes from CERNO_HOST_API_KEY.
Structs§
Functions§
- arg 🔒
- brier 🔒
- Mean squared error between the predicted yes-probability and the truth. Noul cases only, since Brier is defined on a binary outcome.
- calibration_
reading 🔒 - What a best-fit temperature says about the model, in words that stay true for any value.
- fit_
temperature 🔒 - Grid-search the temperature that best explains the labelled outcomes.
- main 🔒
- markdown_
table 🔒 - Render a markdown table with padded cells, the first column left-aligned and the rest right.
- pct 🔒
- percentile 🔒
- pick_
winner 🔒 - The best candidate: highest accuracy among models whose top token was an offered label in every case, ties broken by median latency. Fidelity is a gate first — a model cerno cannot read is not a candidate at any accuracy — and the reference is excluded, since it is the yardstick rather than an option.
- render 🔒
- run_
case 🔒 - softmax_
of 🔒 - summarise 🔒
- top_
token_ 🔒is_ a_ label - Whether the highest-ranked token spells one of
offered.tokensis ranked, highest first.