--- license: mit pipeline_tag: image-segmentation tags: - pytorch - segformer - medical-imaging - semantic-segmentation - wavelet - frequency-domain - boundary-detection --- # WaveSeg **WaveSeg** is a SegFormer-B0 backbone with a lightweight (770-parameter) **Frequency-Boundary Adapter (FBA)** gating its decoder features via a Haar wavelet high-frequency branch. It is evaluated on two 2D medical image segmentation domains: brain MRI lesion segmentation and polyp segmentation. **Locked configuration: WaveSeg = SegFormer-B0 + FBA frequency gating (ablation config A1), with NO boundary-frequency auxiliary loss.** A boundary-frequency loss was also designed and tested at several weights - it is reported as a **negative ablation** (see below), not part of the shipped model. ## Task and datasets - **Task:** binary semantic segmentation, 256x256 input. - **Brain MRI (LGG):** FLAIR-abnormality segmentation on the `mateuszbuda/lgg-mri-segmentation` dataset (Buda et al., *Computers in Biology and Medicine*, 2019) - 110 patients, 3,929 slices, patient-level 70/15/15 split. - **Polyp (Kvasir-SEG):** colonoscopy polyp segmentation on Kvasir-SEG (Jha et al., MMM 2020) - 1,000 images, 800/100/100 split. ## Results All numbers below are read directly from `results/ablation_table_clean.csv` (grouped/averaged by method+dataset across all completed runs - `n` is how many). Region metrics (Dice, mIoU, Boundary-IoU) are higher-is-better; HD95/ASSD (pixel distances at 256x256) are lower-is-better. ### Brain MRI (LGG) | Method | n | Dice | mIoU | Boundary-IoU | HD95 (px) | ASSD (px) | Params | |---|---|---|---|---|---|---|---| | SegFormer-B0 baseline (A0) | 3 | 0.8677 +/- 0.0103 | 0.9159 | 0.7898 | 29.70 | 27.07 | 3,714,401 | | **WaveSeg (A1, ours)** | 2 | **0.8908 +/- 0.0055** | **0.9273** | **0.8112** | **22.83** | **20.56** | 3,715,171 | ### Kvasir-SEG (Polyp) | Method | n | Dice | mIoU | Boundary-IoU | HD95 (px) | ASSD (px) | Params | |---|---|---|---|---|---|---|---| | SegFormer-B0 baseline (A0) | 1 | 0.9049 | 0.9038 | 0.5809 | 18.49 | 4.57 | 3,714,401 | | **WaveSeg (A1, ours)** | 2 | **0.9103 +/- 0.0049** | **0.9084** | **0.5958** | **17.41** | **4.34** | 3,715,171 | WaveSeg (A1) beats the SegFormer-B0 baseline on **every metric, on both datasets**, adding only 770 trainable parameters (3,715,171 vs 3,714,401). ## Honest negative ablation: the boundary-frequency loss does not help The FBA module was originally paired with an auxiliary boundary-frequency loss (BCE+Dice supervising the attention map toward the ground-truth mask's morphological-gradient edge). Tested on LGG at `lambda_boundary` in {0.1, 0.25, 0.5, 1.0}, after seed-averaging (n>=2 where re-run): | lambda_boundary | Dice (LGG) | Beats A1 (gating alone)? | |---|---|---| | 0.0 (= A1, shipped) | 0.8908 +/- 0.0055 | reference | | 0.1 | 0.8826 +/- 0.0174 | no | | 0.25 | 0.8727 | no | | 0.5 | 0.8783 | no | | 1.0 | 0.8753 +/- 0.0069 | no | No tested weight improved on FBA gating alone. This is reported as a deliberate, honest negative ablation rather than dropped: the frequency-domain gating mechanism is doing the work; explicitly supervising its attention map toward GT edges does not add value at any weight tested. See the training repository's `results/tables.md` and `REPORT.md` for the full ablation grid (A0-A5) and the seed-confirmation methodology. ## Training details - **Backbone:** `nvidia/mit-b0` (ImageNet-pretrained SegFormer-B0 encoder), fresh decode head. - **Hardware:** local development/inference on an RTX 4060 (8GB); full training runs on Kaggle T4/P100 and locally, all under an 8GB VRAM budget. - **Precision:** AMP (mixed precision), channels-last memory format, batch size 16, 256x256 input. - **Optimizer:** AdamW, lr 6e-5, weight decay 0.01, 100 epochs. - **Reproducibility:** base seed 42 everywhere (data splits fixed across all seed variants); multi-seed confirmation runs (seeds 42/123/456 depending on the row) give `n>=2` for every number reported above except the two Kvasir/LGG A0 rows still at `n=1`. - **Metrics:** Dice, mIoU (mean of foreground+background IoU), Boundary-IoU (Cheng et al., CVPR 2021), HD95 and ASSD (via `scipy.ndimage`; a fixed image-diagonal penalty is used for the empty-mask edge case so every method aggregates over the identical image set). ## Intended use Research and educational demonstration of frequency-domain gating for medical image segmentation boundary quality. **Not validated for clinical use** - trained on two small public research datasets (3,929 and 1,000 images respectively), with no independent multi-site validation. ## Limitations - Both datasets are 2D, single-institution(-family) research collections; generalization to other scanners, contrast protocols, or endoscopy hardware is untested. - The `n=1` rows above (Kvasir/LGG A0 baseline on Kvasir) have no seed-variance estimate. - The boundary-frequency loss's negative result was tested only on LGG and only up to `lambda_boundary=1.0`; it does not rule out a different loss formulation working. - Small lesions/polyps (a few hundred pixels) are the most common failure mode for both the baseline and WaveSeg. ## How to use See the project README for full reproduction instructions (training, evaluation, and the Gradio demo). The demo Space loads these same weights: . ## Citation If you use this model, please cite the datasets and metrics it builds on: ```bibtex @inproceedings{xie2021segformer, title={SegFormer: Simple and efficient design for semantic segmentation with transformers}, author={Xie, Enze and Wang, Wenhai and Yu, Zhiding and Anandkumar, Anima and Alvarez, Jose M and Luo, Ping}, booktitle={Advances in Neural Information Processing Systems}, volume={34}, pages={12077--12090}, year={2021} } @article{buda2019association, title={Association of genomic subtypes of lower-grade gliomas with shape features automatically extracted by a deep learning algorithm}, author={Buda, Mateusz and Saha, Ashirbani and Mazurowski, Maciej A}, journal={Computers in Biology and Medicine}, volume={109}, pages={218--225}, year={2019}, publisher={Elsevier} } @inproceedings{jha2020kvasir, title={Kvasir-SEG: A segmented polyp dataset}, author={Jha, Debesh and Smedsrud, Pia H and Riegler, Michael A and Halvorsen, P{\aa}l and Lange, Thomas de and Johansen, Dag and Johansen, H{\aa}vard D}, booktitle={International Conference on Multimedia Modeling}, pages={451--462}, year={2020}, organization={Springer} } @inproceedings{cheng2021boundary, title={Boundary IoU: Improving object-centric image segmentation evaluation}, author={Cheng, Bowen and Girshick, Ross and Doll{\'a}r, Piotr and Berg, Alexander C and Kirillov, Alexander}, booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition}, pages={15334--15342}, year={2021} } ``` ## Author Sarvarbek Erniyazov. Licensed under the MIT License (see `LICENSE`).