Ordinal severity grading of coffee leaf disease via knowledge-distilled vision transformers.
A deep-learning study applying knowledge-distilled vision transformers (ViT) to ordinal severity grading of Hemileia vastatrix (Roya) and Mycena citricolor (Ojo de Gallo) on field-captured coffee leaves.
Knowledge distillation transfers learned representations from a large pretrained ViT teacher into a smaller student network suitable for smartphone-class deployment. Ordinal regression heads preserve the natural ranking of severity classes, a structure that standard cross-entropy discards. Coffee leaf rust alone has cost Latin American producers an estimated USD 3.2 billion since the 2012 outbreak (Avelino et al., 2015); this work targets smallholder-deployable diagnostics. Manuscript in preparation, first author. Target venues: Computers and Electronics in Agriculture or Agriculture.
-
01
Field Capture
1,400 leaves · iPhone 11 + 3D-printed rig · ColorChecker · 4,200 images at 0°/45°/abaxial.
-
02
Annotation
EMSAM → CVAT → STAPLE consensus · 3 phytopathologists · 10% triplicate subset · Krippendorff α ≥ 0.80.
-
03
Model Training
Method A: SegFormer-B3 (45M) · Method B: FastViT-T8 (4M) with DINOv2-LoRA teacher and CORN ordinal head.
-
04
Evaluation
3-fold grouped CV · BCa paired bootstrap · TOST equivalence · ±0.05 margin · power 0.97 at Δκ = 0.
-
05
Deployment
FastViT to LiteRT (Android) / Core ML (iOS) · on-device inference, fully offline · 113× fewer FLOPs.
5 architectures plotted on the params–κ plane. The 4 M-parameter FastViT-T8 is hypothesised to match the 45 M SegFormer-B3 to within ±0.05 κ.
Five preregistered floors; all reported regardless of outcome. Bars show target value on the 0–1 scale.
Two One-Sided Tests (TOST) with equivalence margin ±0.05. Expected: 95% CI of Δκ falls fully within the shaded zone.
Probability of correctly declaring equivalence at N = 1,400. Power = 0.97 at Δ = 0, dropping outside ±0.05.
A first working system for the thesis' second target disease, Mycena citricolor (ojo de gallo). It segments necrotic lesions pixel by pixel, measures severity as lesion area ÷ leaf area, and assigns a G0–G3 grade, running fully offline on a phone.
Seven segmentation models plus a transformer classifier, all on the same held-out test set, ranked by lesion Dice. The deployed 17 MB model sits third, behind only a 106 MB and a 324 MB network.
| Model | Grade | QWK | Dice | IoU | Size | On-device |
|---|---|---|---|---|---|---|
| SegFormer-B5 (1024px) | 100% | 1.00 | 0.940 | 0.916 | 324 MB | · |
| DeepLabV3+ ResNet-50 (1024px) | 100% | 1.00 | 0.933 | 0.906 | 106 MB | · |
| DeepLabV3+ MobileNetV2 (1024px) | 100% | 1.00 | 0.924 | 0.895 | 17 MB | Yes |
| SegFormer-B5 (512px) | 100% | 1.00 | 0.911 | 0.878 | 324 MB | · |
| DeepLabV3+ EfficientNet-B7 (512px) | 96.4% | 0.944 | 0.893 | 0.851 | 242 MB | · |
| DeepLabV3+ MobileNetV2 (512px) | 98.2% | 0.972 | 0.892 | 0.850 | 17 MB | Yes |
| DeepLabV3+ ResNet-50 (512px) | 98.2% | 0.972 | 0.891 | 0.856 | ~100 MB | · |
| PLA-ViT (ViT-B/16 classifier) | 98.2% | 0.972 | n/a | n/a | 328 MB | · |
Per-leaf metrics on the real held-out test set (worst-face aggregation). Grade accuracy saturates at 100% for several models, so lesion Dice and model size are the true differentiators. PLA-ViT is a whole-leaf classifier and produces no lesion map.
- Small beats large. A 17 MB MobileNet matched or beat ResNet-50 (106 MB), EfficientNet-B7 (242 MB), SegFormer-B5 (324 MB) and a ViT classifier (328 MB) on grade, and ranks third on lesion Dice. On this dataset, efficiency won, not capacity.
- Two paradigms converge. Direct ViT classification and segment-then-measure reached identical grade accuracy (98.2%, QWK 0.972); but only segmentation yields an interpretable severity map and a defensible percentage a researcher can verify.
- Field-ready, not lab-bound. The whole system runs offline on a phone (ONNX-Runtime WASM, no server, no signal) with GPS geotagging, survey-session incidence tracking, three languages and CSV/PDF export.
A decision-support research tool for coffee-leaf disease assessment, not a validated clinical diagnosis. Reported metrics are on flatbed scans; phone photos yield estimates.
Test the deployed model on a coffee leaf
Pick one of my sample leaves, or upload your own coffee-leaf photo, and the segmentation model runs entirely on your device via ONNX Runtime Web. It localizes the ojo de gallo (Mycena citricolor) lesions, measures the affected leaf area, and returns an ordinal severity grade (G0–G3). No image ever leaves your browser.