Every auto-scoring model ever published, what it claimed when it was trained,
and what it actually did on real boards.
Computed live from the database β nothing here is hand-maintained.
β back to install & setup
S1). When a model
was re-scored on a later exam, segment comes from that restatement. Field is every dart
really thrown on it since. They measure different populations and they routinely disagree; when
they do, the field is the one that counts.
| version | when | segment | label | n exam | train set | darts | bed acc | label acc | accepted | median fix | rigs | sessions | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| loading⦠| |||||||||||||
segment β physical bed on the named exam (inner/outer singles differ; miss wedges count).
label β S/D/T+number; both singles are S1.
n exam β true-gold holdout size that segment/label were measured on.
train set β throws the model was trained on.
bed acc β field: nobody had to tap a bed mix-up (a floor β unnoticed hops count as right).
label acc β field: the score string was right. accepted β darts you never touched.
median fix β how far off the corrected ones were.
A dart counts as accepted only if you left it completely alone. Some corrections just
tidy the label without changing the score, and some fix the bed while the score string was already
right. score accuracy and bed accuracy separate those, so a model is not
punished for cosmetic taps β and not flattered by a bed mix-up that happened to score the same.