STEMM Institute Press
Science, Technology, Engineering, Management and Medicine
A Controlled Evaluation with a Last-Stage Multiscale Variant
DOI: https://doi.org/10.62517/jbdc.202601314
Author(s)
Zhaoyang Chen*, Tuchao Li
Affiliation(s)
College of Artificial Intelligence and Big Data, Guangzhou Vocational University of Science and Technology, Guangzhou, Guangdong, China *Corresponding Author
Abstract
Efficiency claims for compact convolutional networks are commonly supported by parameter counts and floating-point operation estimates, although neither metric directly measures execution time. This study examines how accuracy, static complexity, and observed GPU runtime align for three established lightweight networks and one controlled architectural variant. MobileNetV2, MobileNetV3-Small, ShuffleNetV2, and a MobileNetV3 variant with a parallel multi-scale module in its final inverted-residual block were trained from scratch on CIFAR-10 and CIFAR-100. A fixed data split and training recipe were used throughout, with three independent random seeds per model and dataset, yielding 24 formal runs. ShuffleNetV2 produced the highest mean test accuracy on both datasets, reaching 93.76% on CIFAR-10 and 74.24% on CIFAR-100. MobileNetV2 was fastest in the hardware benchmark despite requiring approximately four times the FLOPs of MobileNetV3. The multi-scale intervention increased CIFAR-10 accuracy by only 0.39 percentage points while adding 45.09% more parameters and 43.89% more FLOPs. On CIFAR-100, it reduced accuracy by 4.18 percentage points and also lowered throughput. The findings show that low FLOPs should not be treated as a sufficient proxy for runtime and that a multi-scale operator can become counterproductive when its placement and fusion cost disturb an already compact representation.
Keywords
Compact Neural Networks; Hardware-aware Evaluation; Inference Latency; Depthwise Convolution; Model Benchmarking; Small-Image Classification
References
[1] D. Qin, C. Leichner, M. Delakis, et al., "MobileNetV4: Universal models for the mobile ecosystem," in Computer Vision - ECCV 2024. Cham, Switzerland: Springer, 2024, pp. 78-96. [2] M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, "MobileNetV2: Inverted residuals and linear bottlenecks," in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, 2018, pp. 4510-4520. [3] A. Howard, M. Sandler, B. Chen, et al., "Searching for MobileNetV3," in Proc. IEEE/CVF Int. Conf. Computer Vision, 2019, pp. 1314-1324. [4] X. Zhang, X. Zhou, M. Lin, and J. Sun, "ShuffleNet: An extremely efficient convolutional neural network for mobile devices," in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, 2018, pp. 6848-6856. [5] N. Ma, X. Zhang, H.-T. Zheng, and J. Sun, "ShuffleNet V2: Practical guidelines for efficient CNN architecture design," in Computer Vision - ECCV 2018. Cham, Switzerland: Springer, 2018, pp. 122-138. [6] W. Yu, P. Zhou, S. Yan, and X. Wang, "InceptionNeXt: When Inception meets ConvNeXt," in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, 2024, pp. 5672-5683. [7] J. Hu, L. Shen, and G. Sun, "Squeeze-and-Excitation networks," in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, 2018, pp. 7132-7141. [8] P. K. A. Vasu, J. Gabriel, J. Zhu, O. Tuzel, and A. Ranjan, "MobileOne: An improved one millisecond mobile backbone," in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, 2023, pp. 7907-7917. [9] J. Chen, S.-H. Kao, H. He, et al., "Run, don't walk: Chasing higher FLOPS for faster neural networks," in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, 2023, pp. 12021-12031. [10] S. Zagoruyko and N. Komodakis, "Wide residual networks," in Proc. British Machine Vision Conf., 2016, pp. 87.1-87.12. [11] I. Loshchilov and F. Hutter, "SGDR: Stochastic gradient descent with warm restarts," in Int. Conf. Learning Representations, 2017. [12] C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, "Rethinking the Inception architecture for computer vision," in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2016, pp. 2818-2826. [13] H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz, "mixup: Beyond empirical risk minimization," in Int. Conf. Learning Representations, 2018. [14] J. Ansel, E. Yang, H. He, et al., "PyTorch 2: Faster machine learning through dynamic Python bytecode transformation and graph compilation," in Proc. 29th ACM Int. Conf. Architectural Support for Programming Languages and Operating Systems, vol. 2, 2024, pp. 929-947.
Copyright @ 2020-2035 STEMM Institute Press All Rights Reserved