Life sciences · Preprint
arXiv · September 4, 2026
Early or partial results. Treat as a signal, not a conclusion.
This is a technical preprint demonstrating that Vision Transformer compression via combined pruning, quantization, and knowledge distillation can reduce model size 54.5-fold while maintaining ~95% accuracy on a proprietary chilli disease classification task. The work is unreviewed, single-dataset, and does not establish clinical, agronomic, or real-world deployment value; it presents an engineering proof-of-concept that a smaller model matches a baseline, but also shows that a directly-trained smaller model performs comparably, questioning the added benefit of the proposed pipeline.
Ablation and technical comparison study (preprint, unreviewed). Chilli (Capsicum annuum) plant disease image dataset; 3-class classification task; cross-village and cross-device out-of-distribution test split used. Total image count, class labels, and split proportions not specified.. Intervention: Vision Transformer compression pipeline: Hessian-Balanced Adaptive Block Pruning (H-BAC), quantization, and attention-based knowledge distillation, applied sequentially.. Compared with: Uncompressed FP32 baseline Vision Transformer; directly-trained INT8 student model of same final size (6.01 MB).. India (implied by chilli crop focus and village-split dataset design); specific field locations not named..
Baseline FP32 Vision Transformer achieved 95.13% accuracy on 3-class chilli disease dataset Fully integrated compression pipeline achieved 54.5x model size reduction (327.42 MB to 6.01 MB) at 95.13 ± 2.32% accuracy Individual compression techniques achieved 74–98% model size reduction while matching baseline accuracy
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
This work is not clinical. For agricultural technologists, it suggests model compression is feasible for on-device plant disease detection, but does not demonstrate field efficacy, cost-benefit, or comparative advantage over simpler approaches in real-world use.
A single-center, uncontrolled technical validation study of model compression methods on a proprietary dataset with no peer review, clinical validation, or comparison to deployed agricultural systems.
As stated by the source record.
Quoted from the source exactly as published.
This work is not clinical. For agricultural technologists, it suggests model compression is feasible for on-device plant disease detection, but does not demonstrate field efficacy, cost-benefit, or comparative advantage over simpler approaches in real-world use.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Chilli (Capsicum annuum) is one of India's most economically significant crops, yet its productivity is persistently threatened by diseases that are difficult to identify without expert intervention. While Vision Transformers (ViTs) have achieved high classification accuracy, their large computational footprint makes deployment on resource constrained devices challenging. Existing compression approaches typically address pruning, quantization, and knowledge distillation in isolation, leaving the potential benefits and interactions of their combined application insufficiently explored. We propose a unified Vision Transformer compression framework that combines Hessian-Balanced Adaptive Block Pruning (H-BAC), guided by second-order sensitivity estimation, with quantization and attention-based knowledge distillation. To systematically identify the most effective configuration within each compression family, each technique is first evaluated independently through controlled ablation studies, after which the best-performing components are integrated into a sequential deployment pipeline tailored to real-world agricultural constraints. On a chilli 3-class village-split dataset with a genuine cross-village, cross-device out-of-distribution test split, the resulting compressed models match or exceed the 95.13% FP32 baseline's accuracy, alongside 74-98% model size reduction, and the fully integrated compression pipeline achieves a 54.5x size reduction (327.42 MB to 6.01 MB) at 95.13 +/- 2.32% accuracy across four tested configurations. A direct comparison further reveals that, on this dataset, a directly-trained student of the same final size, without pruning or distillation, reaches comparable accuracy of 94.87%, at the same 6.01 MB INT8 size, indicating where H-BAC and knowledge distillation are, and are not yet shown to be, worth their computational cost.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.