open-dart-detection · how the system works

How a dart becomes a score

The paths a throw takes through detection, the model, and the three gates that can override it — plus what each record type actually asserts, and which instruments can settle which questions. Written because these interlock: nearly every wrong turn this month came from reasoning about one stage while forgetting the one above it.

This page is written for everyone on the project. The diagrams and the legend carry the story; the mono annotations and the glossary carry the exact names engineers can search for. If a word is unfamiliar, it is in the glossary.

2026-07-27 trainer 0.7.46 model 20260726.uni2-s7-nopos
00

How to read this page

Every diagram reads top to bottom. Each box is something the software does on its own — no person is involved anywhere on these charts until a player corrects a score afterwards.

A box is a step the system performs.
A diamond is an automatic decision. The labels on its outgoing arrows say which condition sends a throw down which path.
An arrow means "then". Order matters everywhere here.
Green is a settled outcome — recorded, done.
Amber is caution — recorded but held back, or waiting for its moment.
Red is an override or a must-not.
A dashed group is a list of options in priority order — the first that applies wins. It is not a sequence.
miss_mass Mono names are the exact identifiers in code and data. Engineers can search for them; everyone else can read straight past them.
01

A throw, end to end

The classification happens before the model runs. The model is only ever asked where, never whether — it cannot see if a dart stayed in the board, so it gets no say.

In plain terms When the cameras see movement stop, the software first decides what happened — did a dart land, bounce out, or did the player just collect their darts — purely by comparing pictures of the board from before and after. Only when it is sure a new dart is sitting in the board does it ask the AI model where that dart is.
motion settles the monitor (MAX over all 3 cameras) sees the movement stop what changed on the board? decided from BOARD CHANGE — the model gets no say board unchanged + hard impact · peak ≥ 120 BOUNCE a dart hit and fell out MISS · sector None auto-recorded into review (AUTO_BOUNCE) board unchanged + faint impact IGNORE no event; logged as a near-miss board emptied (darts pulled) TAKEOUT disputed → prompt; visit kept a NEW dart in the board THROW a dart is in the board ask the model its proposed point now faces the three gates — §02
A dart that sticks in the surround past the scoring ring changes the board, so it always takes the THROW branch and can never be caught by the bounce logic. That is structural, not bad luck — see §03.
02

The three gates on the model's answer

Order matters. All three thresholds are measured per model at export and stamped inside the .onnx — never held on the rig, because none of these numbers means the same thing from one model to the next. Absent or 0.0 always means "gate off", so an older model behaves exactly as it always did.

In plain terms The model's answer is not taken at face value. Three checks stand between its guess and the scoreboard: is the guess too weak to trust at all, does the model itself believe the dart is off the board, and — even if the score stands — was it shaky enough to flag for later. The cut-lines for these checks are measured fresh for every model and travel inside the model file.
the model proposes a point propose() → point, conf, miss_mass the point = soft-argmax = the MEAN of the heatmap belief too weak to place it? yes · conf < 0.1134 no_position MISS · sector None · null contact "the model has no idea where it is" no belief mostly off the board? yes · miss_mass ≥ 0.58 (and point inside 170mm) MISS · N — wedge kept the radius becomes nominal 180mm; the measured angle names the wedge no score the point score_parts(x, y) still not confident? yes · conf < 0.2078 uncertain = true · SHADOW recorded, but changes nothing: the score still stands no the score stands committed to the record
Why no_position exists: the point is the heatmap's mean, so when the belief is diffuse the mean lands near the board centre. Field 2026-07-26 — a dart 216.7mm out scored 25 at r=7.0mm on confidence 0.043. A dart recorded at the bull that was never there is a phantom.
Trap. no_position is not expressed as contact_mm=None. That is the "no proposal" state, and the hands-free /poll path returns event:False on it — silently dropping a real dart, which is the missed-dart bug rather than a fix for it. It travels as a flag, with the point kept as model_contact_mm evidence exactly as a bounce keeps its detection.
Why the angle is kept but the radius replaced. Measured on flagged misses: the point's angle names the right wedge 87.3% of the time, against 72.5% for an angle read off the outside-mass centroid. The radius is the part the soft-argmax provably gets wrong — median 12.9mm short. So only the radius is nominal (MISS_NOMINAL_MM = 180, the measured median radius of a real miss).
03

What the model can actually see

Three different radii get confused constantly. The scoring area ends at 170mm, the model's input is cropped at 178mm, its output grid reaches 200mm — and the physical board face continues to roughly 225mm.

In plain terms The model does not see the camera pictures — it sees a flattened, cropped view of the board. The green disc below is everything it gets. A dart landing outside that disc is still usually assigned the right direction (its body leans into view), but the distance from centre is guesswork.
47mm cropped
discthe shaded area = every pixel the model receives
15.9outer bull
107triple outer
170double outer — scoring ends
178model INPUT extent (data.EXTENT, 448px)
200readout grid (net.EXTENT)
~225physical board face
Between 178mm and the board edge there is a 47mm ring where the model has no pixels at the tip's location — the cameras capture it, our warp crops it away. It is not blind there: the dart protrudes toward the cameras so its body leans into view (verified on a cached sample labelled at 265mm), and the shaft lines are fitted on the UNCROPPED camera frame. That is why the wedge is usually right and only the radius fails. 58% of human-tapped misses land beyond 178mm.
Do not "align" 178 to 200. The readout grid is deliberately wider than the input — that headroom measured better. Widening the input is the open question, and it trades against mm resolution at a fixed 448px. It also changes the cache: EXTENT is baked into the warped pixels, so a model trained at one extent and served at another would be silently wrong — the same failure class as channel order, which costs 18.6 points. Closed 2026-07-27: the warp now travels with the weights (build_cache records it → the export stamps input_extent_mm → the rig builds its warp from the stamp; absent = 178/448).
04

What each record type asserts

Several of these produce the identical score. They are kept distinct because they are different events, and training treats them differently.

In plain terms Two throws can score the same and still be completely different events — a dart that bounced out and a dart the software invented both show "MISS", but one is honest data and the other is poison for training. So the record keeps what happened, not just the number.
flagthe physical claimscorewhat training does
bouncea real dart hit and fell out MISS · nonekept; position is honestly absent
no_positiona real dart stuck somewhere the model cannot place MISS · nonekept, excludable from position learning
phantomno dart was thrown — the opposite of a bounce must be REMOVEDpoison: teaches "empty board → S20"
late_addthe rig missed a dart; a human added it by hand as tappedexcluded
visit_uncertainsome dart in this visit was hand-added, so which label belongs to which physical dart is a guess as tappedthe whole 3-dart round is dropped
Scoring like a bounce is not claiming it bounced. That is why no_position is a separate flag rather than reusing bounce — the score is the same, the event is not, and only one of them is evidence about dart retention.
05

Where a throw's data lands

Three sinks, three different jobs. A number missing from one of them is a question you cannot answer later.

In plain terms Every throw is written down three times, for three different readers: the teaching material the next model learns from, the board's own local logbook, and the shared pool that combines all boards. Each copy answers questions the others cannot.
sinkjobcarries
label.jsontraining gold — what a retrain reads contact_mm · model_contact_mm · confidence · conf_threshold · miss_mass · miss_mass_threshold · corrected · bounce · no_position · visit_uncertain
stats.dbthis rig's own record; answers questions without a network x_mm · proposed_x_mm · score · corrected · confidence · miss_mass · model_version
central metapools the fleet; how a threshold gets chosen from data the same, plus rig_id · calibration_id · tags
Record the threshold next to the reading, always. A boolean freezes one cut into the archive; the raw number lets the whole gate be re-scored at any cut later. The first session on the miss gate produced one wrong MISS·8 and left no trace of how far over the cut it had been — which is why miss_mass now rides all three sinks.
05·b

From those sinks to a training run

§05 ends where a throw is written down. This is what happens to it afterwards — who keeps a copy, which row points at the pixels, and what has to be true before a retrain can start. Every count here was measured on 2026-08-03, not estimated.

In plain terms The board keeps its own notes and sends a copy away. The numbers land in a database; the pictures land in cloud storage; a nightly job copies both somewhere else in case that cloud burns down. To train a new model you pull all of it back to one machine, rebuild a single file the trainer can read, freeze it under a version, and point a GPU at that version.
wherewhat it holdsif it is lost
Pi /dev/shmtoday's frames + decision log gone — RAM-backed by design (SD write budget); survives a restart, not a reboot
Oracle dataset.sqliteall metadata — game_throw 34,680 · capture 34,680 · capture_image 104,040 (2026-09-08) restore from the box — drilled 2026-07-20, 15/15 images sha-identical
S3 bucket dartall pixels — 104,046 objects, 68.2 GiB, presigned-302 reads (2026-09-08) restore from the box (nightly rclone)
storage boxnightly db + blobsboth primaries still live
Oracle /datametadata only since 2026-09-07 — sqlite + models + feedback (~296 MB). Local blobs/ retired db restore from the box; pixels from S3 / the box, not from this volume
R2 snapshotsimmutable training sets (data + code digest) rebuildable — but the reproducibility of past runs is not
workdisk OCD_datathis box's working copies — export, labelled tree, caches, sweep output nothing — all re-pullable or rebuildable
On the dev box, the bulk is not in the repo. /home runs at 98% (47 GB free) and the dataset alone was 61 GB, so everything large and regenerable under open-dart-detection/data/ physically lives in /home/z3n/workdisk/OCD_data/ and is symlinked back — every path in this section resolves unchanged, and the symlinks are deliberate, not damage. What stays in the repo is the 2.7 GB git actually tracks: the pinned calibration/, certified/, tipnet.onnx and the A/B results. New bulk goes to OCD_data/ too; check df -h /home before anything that writes tens of GB.

The one column that points at the pixels

capture_image.path holds the S3 object key verbatim — and it doubles as the relative path inside the export tree, which is why an rclone copy lands every file exactly where the database already expects it. That symmetry is also why a half-finished copy is invisible instead of loud.

hopcolumnreal value
throw → capturegame_throw.capture_idsession-1782269343-r0001-ref
capture → imagecapture_image.capture_idcam0 · median
image → S3capture_image.path blobs/session-1782269343/session-1782269343-r0001-ref/cam0.median.webp
integritycapture_image.sha256cea3c448…4ab9a98e (checked on restore)

Preparing a run — five steps, and the one everybody skips

#stepwhat it does / how it fails
1rsync oracle:/var/lib/docker/volumes/dart-room_dataset-data/_data/ → data/central-export/ metadata. It is a named volume, so it is also a host path — no 16 GB docker cp through /tmp needed, and only the delta moves.
2rclone copy dart-s3:dart/blobs → data/central-export/blobs the pixels. Omit it and nothing errors — see the trap below.
3COUNT(*) capture_image vs find blobs -type f | wc -l the guard. These two must be equal before anything is built.
4model.pull_corrections → model.data rebuilds the labelling tree, then tipnet_cache.npz. Every throw from a trusted rig is pulled and cached (test players excluded) — the shipped recipe is --mode all, i.e. certified base + human-corrected + accepted. What corrected gates is the position loss, not membership.
5dataset_store publish → train freezes a snapshot whose version digests data and code together, so a pinned version means exactly one thing forever.
"We only train on corrected throws" is false — and the docstring that says so is stale. The shipped recipe is --mode all --seg-loss …: the field (which segment the dart is in) is supervised for every throw, while the position (mm) loss applies only where a human actually tapped — base + corrected. That split is the whole point. An accepted throw's contact_mm is bit-identical to the model's own model_contact_mm (the server copies the numbers), so training position on it teaches the model its own answer — the same circularity the 2026-07-24 audit's F10 flags for seg_acc_game. Which segment it landed in, though, the player confirmed by accepting the score, so that supervision is real. Sanity check: a nightly run reports train 5711 / val 476 against ~3,163 corrections in the entire database.
The blob gap — a pull that "succeeds" with a third of the pixels missing. Step 1 brings the database, step 2 brings the images, and doing only the first fails silently: pull_corrections runs clean and the captures with no files simply contribute nothing. Measured 2026-08-03: 47,829 image rows against 29,926 files — 17,903 missing, ~37% of all captures, weighted toward the newest data because that is exactly what lives on S3 (Oracle's local blobs/ was retired 2026-09-07). Step 3 is the whole defence.
Three more that are quiet by construction. Stale snapshot — a run pins a snapshot, so a queue line written today will happily train on weeks-old data and report healthy numbers about the wrong dataset (on 2026-08-03 the newest published snapshot was 20260722-251ee1d4, n=7,018, while central held 15,943 throws / 3,163 corrections); read the n= in the run's own log header. Cache rewritten under a live study — build_cache(force=True) rewrites the npz in place; confirm no trainer is running on either nightly box first. Stamped throw, missing calibration — the throw is excluded rather than legacy-warped (correct: silently applying the wrong warp is what poisons a set), so a whole rig can vanish from a build; the missing_stamped counter is the only trace.
Cloud budget is a wall, not a throttle. Modal stops running apps when the workspace hits its spend limit — it blocks even zero-cost CPU functions, and a sweep killed mid-flight writes nothing (this aborted the 2026-07-20 study at $35.62 of $40). The limit is per billing cycle, so a trivial CPU function is a free probe: if it can create its app, there is room. One 400-epoch seed on an L4 is about one hour, roughly $1.
06

How a model reaches a board

Publish is not ship. Exactly one model is stable; a rig either follows it or pins a version by name.

In plain terms Making a new model available and putting it on people's boards are two separate, deliberate steps — like printing a book versus shipping it to shops. And even once a board has downloaded a new model, it waits for a quiet moment to switch: never in the middle of a game.
train → export stamps thresholds, channel order + warp into the .onnx publish a CANDIDATE on the shelf — running on nobody promote one deliberate command (a human decision) ★ stable exactly one at a time = /v1/models/latest which model does this rig run? · first match wins 1 · DT_MODEL_PIN hard lock set on the box — beats everything › 2 · manual pin a version chosen in this board's settings › 3 · follow stable the default — moves only when promote is run download → verify sha256 → load the file must prove it is intact before it loads hot-swap waits for idle applied on the next /poll, /detect or /game/new — never mid-game
The swap is request-triggered so it can never change the scorer under a live game. On an idle board with nobody polling it will sit as "pending" indefinitely — a single GET /poll applies it and is a no-op when no game is open.
!!

Correction, 2026-07-27 — the far ring is not the main event

In plain terms A follow-up measurement: the dramatic failures far off the board are real, but rare. Most scoring mistakes happen well inside the board — so that, not the outer ring, is where the work went next.
Read this against §02 and §03. Those sections tell the story of darts the model cannot place, because that is what the field failures looked like. Measured properly, that is a minority of the problem:
  • Miss/hit flips are only 30% of score errors (65 of 216 on the true-gold holdout). The other 70% are sector and ring mistakes.
  • The "safe" deep interior — more than 40mm inside the double wire — runs a 41.9% score-error rate. That is simply the mirror of ~57% exact score, and it holds ~2/3 of the errors.
  • Darts past 210mm are ~0.8% of real play, so widening the model's view was aimed at the cheapest axis. That experiment was run and is not being resumed.
  • The readout matters more than the field of view: restricting the soft-argmax to a window around its peak cuts total score errors 216 → 195 (~+4.4pt exact) on models already trained, with no retraining.
Full working: docs/research/2026-07-27-where-the-errors-are.md.
07

Which instrument can settle which question

The most expensive lesson here. Every one of these numbers has a resolution limit, and below it the instrument reads like signal.

In plain terms Every measuring stick this project owns has a smallest difference it can reliably detect. Read anything smaller than that off it and you are reading noise — it has cost us wrong conclusions more than once. Before trusting any comparison, ask: is this instrument even capable of seeing an effect this small?
instrumentcan settlecannot settle
seg_acc_game
paired seeds
effects above ~1.5 points at 6 pairs anything smaller — per-seed sd is 0.019, and a 3/3 sweep at three pairs went to a tie at six
the accepted slice
(66% of the holdout)
agreement with the incumbent any geometry change — its gold IS the old system's own answer (F10)
holdout exact %a collapsed run ranking two good models — a best-vs-worst live test inverted (field score_accuracy 0.837 n=123 vs 0.873 n=63; the difference is inside a ±10.5pt CI, so it neither confirms nor refutes — which is the point)
field score_accuracylarge effects (acceptance 44% → 82%) a 3-point difference needs ~2,200 throws per arm
model-independent
data checks
data quality: sector agreement 0.525 → 0.687, tip residual 23.9 → 11.4mm nothing about a specific model
a named failure mode
+ human taps
the sharpest tool we have — 229 tapped misses, 69% mis-scored, AUC 0.948. ⚠ 69% conditions on tapped misses, which selects failures; unconditional ≈34% needs the failure to be named first

The operating rule: stop ranking recipes, hunt failure modes. Every reliable answer this project has produced came from a large effect or a model-independent check — never from a small delta on the holdout.

08

Traps that have already cost a wrong conclusion

In plain terms Mistakes this project has already made once, kept here so nobody makes them twice.
a–z

The words on this page

Plain language first; the mono names are what the same thing is called in the code and the data.

rig
One physical dartboard installation: the board, three small cameras around it, and a Raspberry Pi computer running our software.
trainer
Our software on the rig. It watches the cameras, decides what happened, scores the throw, and collects labelled examples for teaching the next model.
model · TipNet
The neural network. It looks at the flattened camera views and answers one question: where is the dart's tip? It is never asked whether a dart landed — see §01.
calibration
The measured geometry that maps each camera's picture onto the flat board face. Every throw records which calibration it was captured under.
warp · extent
Using the calibration, the camera pictures are flattened onto the board plane — the warp. The extent is how far out from the bullseye (in mm) those flattened pixels reach: 178mm today.
heatmap · belief map
The model's raw answer: not a single point but a grid over the board scoring "how strongly do I believe the tip is here".
soft-argmax
Turning the belief map into one point by taking its weighted average. Sharp belief gives a precise point; spread-out belief gives a point that drifts toward the board centre.
confidence · conf
How concentrated the model's belief is (the mass at its peak). Low confidence means the model is guessing.
miss_mass
The share of the model's belief that lies outside the scoring area. High means the model itself thinks the dart is off the board.
gate · threshold
An automatic check that can overrule the model's point (§02). Each threshold is measured for that specific model at export and stamped inside the model file; absent means the gate is off.
bounce · bounce-out
A real dart hit the board and fell out. Recorded as a MISS with — honestly — no position.
phantom
A dart in the record that was never thrown: the opposite of a bounce. It must be removed — left in, it teaches the model that an empty board contains a dart.
takeout · visit
A visit is one turn: up to three darts. The takeout is the player pulling their darts from the board at the end of it.
shadow
A check that runs and records its verdict but is not yet allowed to change anything. It is how a new gate earns trust before it may act.
holdout
A fixed set of past throws with trusted answers, kept aside to test models on. Small by the cost of trusted answers: it catches a collapsed model, but cannot rank two good ones (§07).
field acceptance · score_accuracy · bed_accuracy
The share of live throws that players let stand (plus label touch-ups). The strongest real-world measure this project has. Three meters, not one (2026-07-31): acceptance is the raw not-corrected rate; score_accuracy asks whether the score label was right; bed_accuracy asks whether the physical field was right too — single_inner and single_outer share a score but are two beds, up to 90 mm apart with the treble between them. score_accuracy − bed_accuracy is the field band-confusion rate; 30% of every correction ever filed as "cosmetic" turned out to be a wrong bed. ⚠ bed_accuracy is a FLOOR — the field only sees a bed error a human corrected. The holdout's bed_exact grades every throw and is the honest detector metric; never compare their absolute levels.
publish · promote · stable
Publish puts a candidate model on the shelf — nobody runs it. Promote is the deliberate one-command step that makes it the single stable model rigs follow (§06).
hot-swap
A rig switching to a newly downloaded model. It waits for an idle moment between requests — never in the middle of a game.