- Smart Bin Recognition β validator and identifier
Smart Bin Recognition β validator and identifier
Two ONNX graphs behind a web app that answers "what is this bin, and what goes in it?" from a phone camera. They run server-side; nothing here is downloaded to a device.
One of them may ship. The other may not, and the service refuses to start on it. Both are published because an artefact that failed its gate is evidence, not waste, and hiding it would make the next person repeat the measurement.
| validator | identifier | |
|---|---|---|
| question | is there a bin? | which bin? |
| file | v1/validator-v1.onnx |
v1/identifier-v1.onnx |
| architecture | YOLO11n detection | YOLO11s-cls |
| input | 448 Γ 448, NCHW, static batch 1 | 320 Γ 320, NCHW, dynamic batch |
| weights | int8 (U8S8, per-channel) | int8 (U8S8, per-channel) |
| size | 3.15 MB | 5.81 MB |
may_ship |
false |
true |
Every number below carries the split it was measured on and the hardware it was measured on. A figure without both is not quotable, here or in the project that produced it.
The class list, and its order
Read the sidecar, not this table. classes in each *.json is the
authoritative order and the service reads it from there. The order is the
ONNX output index; reordering it silently invalidates every deployed model, and
the project pins it with a test.
validator β 1 class:
| index | id |
|---|---|
| 0 | bin |
identifier β 3 classes, alphabetical, which is the order the training framework emitted and therefore the order the head produces. It is not the order the taxonomy lists them in:
| index | id | what it is |
|---|---|---|
| 0 | igloo |
a glass/bottle bank |
| 1 | wheelie_large |
a 660β1100 L four-wheeled container |
| 2 | wheelie_small |
a 120β240 L two-wheeled household bin |
These are form factors β shapes, not meanings. The model never predicts a waste stream. Meaning comes from a per-jurisdiction region pack (a JSON file) that maps (form factor, colour, region) to a stream, so adding a country is a data change rather than a retrain.
The identifier β passes every gate
| gate | budget | measured | split / hardware |
|---|---|---|---|
| int8 accuracy cost | β€ 0.02 top-1 | 0.0000 | test, Kaggle T4 training run |
| median latency | β€ 25 ms / crop | 9.921 ms | GCE n2-standard-4, 2 of 4 vCPU pinned, Intel Cascade Lake, europe-west3-a, onnxruntime 1.29.0, representative: true |
| p95 latency | β | 11.409 ms | same |
| hardware named | required | yes | same |
gate_result: {"failures": [], "unmeasured": [], "may_ship": true}.
Its evidence is thin, and the card says so
test top-1 is 1.0000 on 47 crops.
| class | crops in test |
capture clusters in test |
|---|---|---|
wheelie_small |
25 | 9 |
wheelie_large |
19 | 8 |
igloo |
3 | 2 |
- 47 correct out of 47 is not certainty. By the rule of three the 95 % lower bound is 0.936.
igloo's contribution is three crops from two scenes. Its per-class number is noisy and must never be quoted clean.- The better estimate is 0.9834 β out-of-fold accuracy over all 403 crops
under
GroupKFoldon capture cluster, against a 0.6823 majority-class baseline. A random split of these crops measures memorisation: 403 crops come from 100 capture clusters, and the largest holds 18 photographs of one bin.
The task is close to saturated, so this measurement can distinguish "int8 and fp32 differ by at most one crop" from "they differ by two or more" and not much finer. The gate passed honestly; it did not pass by a margin the data can resolve.
The validator β trained, real, and it cannot ship
It is not a failed training run. It is a working detector that int8 quantisation destroys.
| gate | budget | measured | split / hardware |
|---|---|---|---|
| int8 accuracy cost | β€ 0.02 mAP@0.5 | 0.727 β | test |
| median latency | β€ 50 ms | 18.252 ms β | GCE n2-standard-4, Cascade Lake, 2 vCPU pinned, representative: true |
| p95 latency | β | 21.345 ms | same |
| value | split | |
|---|---|---|
| mAP@0.5, PyTorch fp32 reference | 0.7524 | test |
| mAP@0.5, the int8 graph published here | 0.025 | test |
| specificity on background frames | 0.9793 | 2 662 hard negatives |
gate_result: {"failures": ["int8 quantisation cost 0.727 map50 (max 0.02)"], "may_ship": false}.
The service loads the validator unconditionally and refuses to start on an
artefact whose sidecar says may_ship: false. So nothing is deployed, and that
is correct rather than pending.
What is known about why
- Quantising the detection head is what collapses it. Leaving
/model.23/in fp32 takes the graph from 0.015 to 0.7481 onvalβ a fifty-fold recovery β for about +5.7 ms and +1.2 MB on an x86 proxy. - It does not follow that nothing else matters. That graph is still quantised everywhere else and still loses 0.0252 against a 0.02 budget. It misses by 0.0052, and the residual is unattributed.
- The three remedies onnxruntime names for this failure mode β S8S8,
reduce_range, U8U8 β all stay at collapse, as does per-tensor. This is not the x86 saturation case it resembles. - Post-training int8 over the whole graph is not viable for this
architecture. Any future YOLO11 detection export here starts from
exclude_head=True.
fp32 is the open route, and it was measured
Measured 2026-08-22, both formats on one Cascade Lake instance with the arms alternated:
| int8 | fp32 | |
|---|---|---|
| validator p50 | 17.921 ms | 24.605 ms |
| against a 50 ms budget | passes | passes |
| concurrent scanners @ 1 bin | 5 | 4 |
| accuracy cost | β0.727 mAP | 0.0 |
An fp32 validator meets the latency budget with 25 ms to spare and costs one concurrent scanner. The ship gate currently refuses unquantised artefacts outright, and its stated reason β "it will not meet the latency budget" β is false on this architecture at this input size. Splitting that gate into per-format profiles is staged for the maintainer's decision and is not merged. No fp32 artefact is published in this repository.
Serving contract
Both graphs, from the sidecars:
- Normalisation: scale
1/255, mean[0,0,0], std[1,1,1]. No ImageNet statistics. - Input name:
images. LayoutNCHW. - NMS is not in the graph. Apply it yourself:
iou 0.45,score 0.35. - The identifier accepts a batch > 1 and the validator does not. A frame can hold several bins β a bank of six is a normal input, not an edge case β so the crops go through one call. The validator sees exactly one frame, every time.
Read imgsz, classes, normalisation and nms from the sidecar rather
than from this card. The service does, which is why a model swap needs no code
change.
Coverage gap β six of ten form factors have no data at all
The taxonomy defines ten form factors. The identifier was trained on three.
| form factor | training crops | status |
|---|---|---|
wheelie_small |
247 | trained |
wheelie_large |
115 | trained |
igloo |
40 | trained |
street_basket |
1 | dropped β one crop in one capture cluster cannot be split across train/val/test, so it can be neither trained nor evaluated |
underground, textile_bank, sack, crate, wall_unit, container_bank |
0 | no data |
Those six keep their ids. An id with no training data is a coverage gap, not
a deletion. Everything the identifier has never seen resolves to unknown, which
is a designed state with a real interface, and the honest answer.
Open Images cannot close this gap. A frozen survey of 384 boxes found it to
be a street-litter corpus β 35 % street_basket, and zero underground,
textile_bank or wall_unit.
There is no geographic holdout
Every training frame comes from one city in one week (Deggendorf, Bavaria),
plus Open Images frames whose region_id is unknown because the source does
not record where a photograph was taken. No split in this data answers "does it
work in another city." The targets min_recall_heldout_city and
min_formfactor_acc_heldout_city are reported as unmeasurable, not as
passed.
Do not quote the predecessor project's 95.2 % mAP as a baseline for these models. It was measured on a random split of one week's photographs in one city.
What this must not be used for
- Do not use the validator as published. It scores 0.025 mAP@0.5. It is here
as the evidence behind a
may_ship: false, not as a detector. - Do not derive disposal advice from these models alone. They output a shape. Being confidently wrong about what goes in which bin is this product's worst failure mode, and the mapping from shape to rule lives in a per-region pack that is versioned, sourced and separately reviewed.
- Do not deploy this outside Deggendorf and expect the numbers above. They are in-distribution numbers on one city, and the card says so twice because it is the most likely way to be wrong with them.
- Do not treat a colour attribute as measured. As of 2026-08-22 colour
agreement against hand-labelled ground truth is 0.5625 on bin bodies, and
lid colour is not measured at all β an upper-band sampler scored 0.1966 and
was deliberately left unwired. Any rule matching on
lid_coloris unreachable. - Do not fine-tune on user-contributed images without human label review. High-agreement machine labels may auto-accept over a public corpus; never over a frame a user contributed, and never for a form factor that has no data yet.
Provenance
Trained on arudaev/smart-bin-detect
at revision 8666aa23 β 18 954 frames: 370 legacy, 1 110 Open Images bins,
17 474 background. The identifier additionally uses a private crop pool whose
403 crops were all adjudicated by a human, blind, in a single pass; against
that pass the pool's shipped streamβshape proposals were wrong on 116 of
403 (28.8 %).
Splits are group-aware on capture cluster, so two photographs of the same bin can never straddle a split.
Model tree for arudaev/smart-bin-detect
Base model
Ultralytics/YOLO11