The paths a throw takes through detection, the model, and the three gates that
can override it — plus what each record type actually asserts, and which instruments can settle
which questions. Written because these interlock: nearly every wrong turn this month came from
reasoning about one stage while forgetting the one above it.
This page is written for everyone on the project.
The diagrams and the legend carry the story; the mono
annotations and the glossary carry the exact names engineers can search
for. If a word is unfamiliar, it is in the glossary.
Every diagram reads top to bottom. Each box is something the
software does on its own — no person is involved anywhere on these charts until a player
corrects a score afterwards.
A box is a step the system performs.
A diamond is an automatic decision. The labels on its outgoing arrows say which
condition sends a throw down which path.
An arrow means "then". Order matters everywhere here.
Green is a settled outcome — recorded, done.
Amber is caution — recorded but held back, or waiting for its moment.
Red is an override or a must-not.
A dashed group is a list of options in priority order — the first that applies
wins. It is not a sequence.
miss_massMono names are the exact identifiers in code and data. Engineers can search for
them; everyone else can read straight past them.
01
A throw, end to end
The classification happens before the model runs. The
model is only ever asked where, never whether — it cannot see if a dart stayed in
the board, so it gets no say.
In plain terms
When the cameras see movement stop, the software first decides what happened — did a dart
land, bounce out, or did the player just collect their darts — purely by comparing pictures of the
board from before and after. Only when it is sure a new dart is sitting in the board does it ask
the AI model where that dart is.
A dart that sticks in the surround past the scoring ring changes the board, so
it always takes the THROW branch and can never be caught by the bounce logic. That is structural,
not bad luck — see §03.
02
The three gates on the model's answer
Order matters. All three thresholds are measured per model at
export and stamped inside the .onnx — never held on the rig, because none of
these numbers means the same thing from one model to the next. Absent or 0.0 always
means "gate off", so an older model behaves exactly as it always did.
In plain terms
The model's answer is not taken at face value. Three checks stand between its guess and the
scoreboard: is the guess too weak to trust at all, does the model itself believe the dart is off
the board, and — even if the score stands — was it shaky enough to flag for later. The cut-lines
for these checks are measured fresh for every model and travel inside the model file.
Why no_position exists: the point is the heatmap's mean, so when the
belief is diffuse the mean lands near the board centre. Field 2026-07-26 — a dart 216.7mm
out scored 25 at r=7.0mm on confidence 0.043. A dart recorded at the bull that was
never there is a phantom.
Trap.no_position is not expressed as
contact_mm=None. That is the "no proposal" state, and the hands-free /poll
path returns event:False on it — silently dropping a real dart, which is the
missed-dart bug rather than a fix for it. It travels as a flag, with the point kept as
model_contact_mm evidence exactly as a bounce keeps its detection.
Why the angle is kept but the radius replaced. Measured on flagged misses:
the point's angle names the right wedge 87.3% of the time, against 72.5% for an angle read
off the outside-mass centroid. The radius is the part the soft-argmax provably gets wrong — median
12.9mm short. So only the radius is nominal (MISS_NOMINAL_MM = 180, the measured
median radius of a real miss).
03
What the model can actually see
Three different radii get confused constantly. The scoring area
ends at 170mm, the model's input is cropped at 178mm, its output grid reaches
200mm — and the physical board face continues to roughly 225mm.
In plain terms
The model does not see the camera pictures — it sees a flattened, cropped view of the board. The
green disc below is everything it gets. A dart landing outside that disc is still usually assigned
the right direction (its body leans into view), but the distance from centre is guesswork.
discthe shaded area = every pixel the model receives
15.9outer bull
107triple outer
170double outer — scoring ends
178model INPUT extent (data.EXTENT, 448px)
200readout grid (net.EXTENT)
~225physical board face
Between 178mm and the board edge there is a 47mm ring where the model has no pixels at
the tip's location — the cameras capture it, our warp crops it away. It is not blind
there: the dart protrudes toward the cameras so its body leans into view (verified on a cached
sample labelled at 265mm), and the shaft lines are fitted on the UNCROPPED camera frame. That is
why the wedge is usually right and only the radius fails. 58% of human-tapped misses
land beyond 178mm.
Do not "align" 178 to 200. The readout grid is deliberately wider than the
input — that headroom measured better. Widening the input is the open question, and it
trades against mm resolution at a fixed 448px. It also changes the cache: EXTENT is baked into the
warped pixels, so a model trained at one extent and served at another would be silently wrong — the
same failure class as channel order, which costs 18.6 points. Closed 2026-07-27: the warp now
travels with the weights (build_cache records it → the export stamps
input_extent_mm → the rig builds its warp from the stamp; absent = 178/448).
04
What each record type asserts
Several of these produce the identical score. They are kept
distinct because they are different events, and training treats them differently.
In plain terms
Two throws can score the same and still be completely different events — a dart that bounced out
and a dart the software invented both show "MISS", but one is honest data and the other is poison
for training. So the record keeps what happened, not just the number.
flag
the physical claim
score
what training does
bounce
a real dart hit and fell out
MISS · none
kept; position is honestly absent
no_position
a real dart stuck somewhere the model cannot place
MISS · none
kept, excludable from position learning
phantom
no dart was thrown — the opposite of a bounce
must be REMOVED
poison: teaches "empty board → S20"
late_add
the rig missed a dart; a human added it by hand
as tapped
excluded
visit_uncertain
some dart in this visit was hand-added, so which
label belongs to which physical dart is a guess
as tapped
the whole 3-dart round is dropped
Scoring like a bounce is not claiming it bounced. That is why
no_position is a separate flag rather than reusing bounce — the score is
the same, the event is not, and only one of them is evidence about dart retention.
05
Where a throw's data lands
Three sinks, three different jobs. A number missing from one of
them is a question you cannot answer later.
In plain terms
Every throw is written down three times, for three different readers: the teaching material the
next model learns from, the board's own local logbook, and the shared pool that combines all
boards. Each copy answers questions the others cannot.
pools the fleet; how a threshold gets chosen from data
the same, plus rig_id · calibration_id · tags
Record the threshold next to the reading, always. A boolean freezes one cut
into the archive; the raw number lets the whole gate be re-scored at any cut later. The first
session on the miss gate produced one wrong MISS·8 and left no trace of how far over
the cut it had been — which is why miss_mass now rides all three sinks.
05·b
From those sinks to a training run
§05 ends where a throw is written down. This is what happens to
it afterwards — who keeps a copy, which row points at the pixels, and what has to be true before
a retrain can start. Every count here was measured on 2026-08-03, not estimated.
In plain terms
The board keeps its own notes and sends a copy away. The numbers land in a database; the pictures
land in cloud storage; a nightly job copies both somewhere else in case that cloud burns down. To
train a new model you pull all of it back to one machine, rebuild a single file the trainer can
read, freeze it under a version, and point a GPU at that version.
where
what it holds
if it is lost
Pi /dev/shm
today's frames + decision log
gone — RAM-backed by design (SD write budget); survives a restart, not a reboot
restore from the box — drilled 2026-07-20, 15/15 images sha-identical
S3 bucket dart
all pixels — 104,046 objects, 68.2 GiB, presigned-302 reads (2026-09-08)
restore from the box (nightly rclone)
storage box
nightly db + blobs
both primaries still live
Oracle /data
metadata only since 2026-09-07 — sqlite + models + feedback (~296 MB). Local blobs/ retired
db restore from the box; pixels from S3 / the box, not from this volume
R2 snapshots
immutable training sets (data + code digest)
rebuildable — but the reproducibility of past runs is not
workdisk OCD_data
this box's working copies — export, labelled tree, caches, sweep output
nothing — all re-pullable or rebuildable
On the dev box, the bulk is not in the repo./home
runs at 98% (47 GB free) and the dataset alone was 61 GB, so everything large and regenerable
under open-dart-detection/data/ physically lives in
/home/z3n/workdisk/OCD_data/ and is symlinked back — every path in
this section resolves unchanged, and the symlinks are deliberate, not damage. What stays in the
repo is the 2.7 GB git actually tracks: the pinned calibration/,
certified/, tipnet.onnx and the A/B results.
New bulk goes to OCD_data/ too; check
df -h /home before anything that writes tens of GB.
The one column that points at the pixels
capture_image.path holds the S3 object key
verbatim — and it doubles as the relative path inside the export tree, which is why an
rclone copy lands every file exactly where the database already expects
it. That symmetry is also why a half-finished copy is invisible instead of loud.
the pixels. Omit it and nothing errors — see the trap below.
3
COUNT(*) capture_image vs find blobs -type f | wc -l
the guard. These two must be equal before anything is built.
4
model.pull_corrections → model.data
rebuilds the labelling tree, then tipnet_cache.npz.
Every throw from a trusted rig is pulled and cached (test players excluded) —
the shipped recipe is --mode all, i.e. certified base +
human-corrected + accepted. What corrected gates is the
position loss, not membership.
5
dataset_store publish → train
freezes a snapshot whose version digests data and code together, so a pinned
version means exactly one thing forever.
"We only train on corrected throws" is false — and the docstring that says so
is stale. The shipped recipe is
--mode all --seg-loss …: the field (which segment the dart is in)
is supervised for every throw, while the position (mm) loss applies only where a
human actually tapped — base + corrected. That split is the whole point. An accepted throw's
contact_mm is bit-identical to the model's own
model_contact_mm (the server copies the numbers), so training position
on it teaches the model its own answer — the same circularity the 2026-07-24 audit's F10 flags for
seg_acc_game. Which segment it landed in, though, the player
confirmed by accepting the score, so that supervision is real. Sanity check: a nightly run reports
train 5711 / val 476 against ~3,163 corrections in the entire database.
The blob gap — a pull that "succeeds" with a third of the pixels missing.
Step 1 brings the database, step 2 brings the images, and doing only the first fails silently:
pull_corrections runs clean and the captures with no files simply
contribute nothing. Measured 2026-08-03: 47,829 image rows against 29,926 files —
17,903 missing, ~37% of all captures, weighted toward the newest data because that is
exactly what lives on S3 (Oracle's local blobs/ was retired 2026-09-07).
Step 3 is the whole defence.
Three more that are quiet by construction.Stale snapshot — a run pins a snapshot, so a queue line written today will happily train on
weeks-old data and report healthy numbers about the wrong dataset (on 2026-08-03 the newest
published snapshot was 20260722-251ee1d4, n=7,018, while central held
15,943 throws / 3,163 corrections); read the n= in the run's own log
header. Cache rewritten under a live study — build_cache(force=True)
rewrites the npz in place; confirm no trainer is running on either nightly box first.
Stamped throw, missing calibration — the throw is excluded rather than
legacy-warped (correct: silently applying the wrong warp is what poisons a set), so a whole rig
can vanish from a build; the missing_stamped counter is the only trace.
Cloud budget is a wall, not a throttle. Modal stops running apps when the
workspace hits its spend limit — it blocks even zero-cost CPU functions, and a sweep killed
mid-flight writes nothing (this aborted the 2026-07-20 study at $35.62 of $40). The limit is per
billing cycle, so a trivial CPU function is a free probe: if it can create its app, there is room.
One 400-epoch seed on an L4 is about one hour, roughly $1.
06
How a model reaches a board
Publish is not ship. Exactly one model is stable; a rig
either follows it or pins a version by name.
In plain terms
Making a new model available and putting it on people's boards are two separate, deliberate steps —
like printing a book versus shipping it to shops. And even once a board has downloaded a new model,
it waits for a quiet moment to switch: never in the middle of a game.
The swap is request-triggered so it can never change the scorer under a live game. On
an idle board with nobody polling it will sit as "pending" indefinitely — a single
GET /poll applies it and is a no-op when no game is open.
!!
Correction, 2026-07-27 — the far ring is not the main event
In plain terms
A follow-up measurement: the dramatic failures far off the board are real, but rare. Most scoring
mistakes happen well inside the board — so that, not the outer ring, is where the work
went next.
Read this against §02 and §03. Those sections tell the story of
darts the model cannot place, because that is what the field failures looked like. Measured
properly, that is a minority of the problem:
Miss/hit flips are only 30% of score errors (65 of 216 on the true-gold holdout). The
other 70% are sector and ring mistakes.
The "safe" deep interior — more than 40mm inside the double wire — runs a 41.9%
score-error rate. That is simply the mirror of ~57% exact score, and it holds ~2/3 of the errors.
Darts past 210mm are ~0.8% of real play, so widening the model's view was aimed at the
cheapest axis. That experiment was run and is not being resumed.
The readout matters more than the field of view: restricting the soft-argmax to a
window around its peak cuts total score errors 216 → 195 (~+4.4pt exact) on models already
trained, with no retraining.
Full working: docs/research/2026-07-27-where-the-errors-are.md.
07
Which instrument can settle which question
The most expensive lesson here. Every one of these numbers has a
resolution limit, and below it the instrument reads like signal.
In plain terms
Every measuring stick this project owns has a smallest difference it can reliably detect. Read
anything smaller than that off it and you are reading noise — it has cost us wrong conclusions
more than once. Before trusting any comparison, ask: is this instrument even capable of seeing an
effect this small?
instrument
can settle
cannot settle
seg_acc_game paired seeds
effects above ~1.5 points at 6 pairs
anything smaller — per-seed sd is 0.019, and a 3/3 sweep at three pairs went to a
tie at six
the accepted slice (66% of the holdout)
agreement with the incumbent
any geometry change — its gold IS the old system's own answer (F10)
holdout exact %
a collapsed run
ranking two good models — a best-vs-worst live test inverted
(field score_accuracy 0.837 n=123 vs 0.873 n=63; the difference is inside a ±10.5pt CI, so
it neither confirms nor refutes — which is the point)
field score_accuracy
large effects (acceptance 44% → 82%)
a 3-point difference needs ~2,200 throws per arm
model-independent data checks
data quality: sector agreement 0.525 → 0.687, tip residual 23.9 → 11.4mm
nothing about a specific model
a named failure mode + human taps
the sharpest tool we have — 229 tapped misses, 69% mis-scored, AUC 0.948.
⚠ 69% conditions on tapped misses, which selects failures; unconditional ≈34%
needs the failure to be named first
The operating rule: stop ranking recipes, hunt failure modes. Every reliable
answer this project has produced came from a large effect or a model-independent check — never from
a small delta on the holdout.
08
Traps that have already cost a wrong conclusion
In plain terms
Mistakes this project has already made once, kept here so nobody makes them twice.
Two arms scored on two holdouts. A quarantine removes rows from the holdout too
(n=1175 vs 1191). Always re-score every checkpoint on one cache before reading a result.
A single unrepeated measurement. One 138.4ms reading became a reported "20% win"; the
same venv gives 186ms on every repeat. Alternate conditions, ≥3 rounds each.
calibrate/drift cannot detect a camera swap. It folds into ±9° sector-modulo,
so a 120° viewpoint change aliases into a small number. Score every live frame against every
stored calibration instead.
Two swaps cancel. "Nothing moved" and "moved twice" are the same measurement.
A failed nightly job still pops from the queue. It vanishes silently — check the queue
after any failure.
Camera identity is position only. All three report serial SN0001; the
trainer binds by USB path, so index follows the port.
a–z
The words on this page
Plain language first; the mono names are what the
same thing is called in the code and the data.
rig
One physical dartboard installation: the board, three small cameras around it, and a
Raspberry Pi computer running our software.
trainer
Our software on the rig. It watches the cameras, decides what happened, scores the throw,
and collects labelled examples for teaching the next model.
model · TipNet
The neural network. It looks at the flattened camera views and answers one question: where
is the dart's tip? It is never asked whether a dart landed — see §01.
calibration
The measured geometry that maps each camera's picture onto the flat board face. Every throw
records which calibration it was captured under.
warp · extent
Using the calibration, the camera pictures are flattened onto the board plane — the warp.
The extent is how far out from the bullseye (in mm) those flattened pixels reach: 178mm today.
heatmap · belief map
The model's raw answer: not a single point but a grid over the board scoring "how strongly
do I believe the tip is here".
soft-argmax
Turning the belief map into one point by taking its weighted average. Sharp belief gives a
precise point; spread-out belief gives a point that drifts toward the board centre.
confidence · conf
How concentrated the model's belief is (the mass at its peak). Low confidence means the
model is guessing.
miss_mass
The share of the model's belief that lies outside the scoring area. High means the model
itself thinks the dart is off the board.
gate · threshold
An automatic check that can overrule the model's point (§02). Each threshold is measured for
that specific model at export and stamped inside the model file; absent means the gate is off.
bounce · bounce-out
A real dart hit the board and fell out. Recorded as a MISS with — honestly — no position.
phantom
A dart in the record that was never thrown: the opposite of a bounce. It must be removed —
left in, it teaches the model that an empty board contains a dart.
takeout · visit
A visit is one turn: up to three darts. The takeout is the player pulling their darts from
the board at the end of it.
shadow
A check that runs and records its verdict but is not yet allowed to change anything. It is
how a new gate earns trust before it may act.
holdout
A fixed set of past throws with trusted answers, kept aside to test models on. Small by the
cost of trusted answers: it catches a collapsed model, but cannot rank two good ones (§07).
field acceptance · score_accuracy · bed_accuracy
The share of live throws that players let stand (plus label touch-ups). The strongest
real-world measure this project has. Three meters, not one (2026-07-31): acceptance is the
raw not-corrected rate; score_accuracy asks whether the score label was right;
bed_accuracy asks whether the physical field was right too —
single_inner and single_outer share a score but are two beds, up to
90 mm apart with the treble between them. score_accuracy − bed_accuracy is the
field band-confusion rate; 30% of every correction ever filed as "cosmetic" turned out to be a
wrong bed. ⚠ bed_accuracy is a FLOOR — the field only sees a bed error a human corrected.
The holdout's bed_exact grades every throw and is the honest detector metric; never
compare their absolute levels.
publish · promote · stable
Publish puts a candidate model on the shelf — nobody runs it. Promote is the deliberate
one-command step that makes it the single stable model rigs follow (§06).
hot-swap
A rig switching to a newly downloaded model. It waits for an idle moment between requests —
never in the middle of a game.