STEMM Institute Press
Science, Technology, Engineering, Management and Medicine
Multimodal Agent Development and Educational Application Based on GLM-4
DOI: https://doi.org/10.62517/jbdc.202601331
Author(s)
Shoutong Huang*, Yu Ma*
Affiliation(s)
School of Electronic and Electrical Engineering, Ningxia University, Yinchuan, Ningxia, China *Corresponding Author
Abstract
Recent progress in large language models has moved artificial intelligence from isolated text reasoning toward integrated multimodal perception, generation, and action. Built on the GLM-4 ecosystem released by Zhipu AI, this work develops a multimodal agent that couples image-text understanding with content generation, and deploys it as an image-text question-answering service for junior secondary education. We describe the conceptual evolution of the GLM family, analyze the engineering trade-offs of cross-modal alignment, long-context memory, latency, and safety, and implement a front-end and back-end separated agent using GLM-4V and CogView together with FastAPI and Streamlit. Evaluated on fifty biology textbook illustrations, the agent reaches 84% image-text understanding accuracy with a 6% factual-error rate. We further design a multimodal education assistant and discuss its social value and the practical challenges of deployment.
Keywords
GLM-4; Multimodal Agent; CogView; Educational Application; FastAPI; Image-Text Question Answering
References
[1] Du, Z., Qian, Y., Liu, X., Ding, M., Qiu, J., Yang, Z., et al. GLM: General Language Model Pretraining with Autoregressive Blank Infilling. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL 2022), pp. 320-335. [2] Yao, S., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., Narasimhan, K. ReAct: Synergizing Reasoning and Acting in Language Models. In Proceedings of the 11th International Conference on Learning Representations (ICLR 2023), Kigali, Rwanda, May 1-5, 2023. [3] Xi, Z., Chen, W., Guo, X., He, W., Ding, Y., Hong, B., et al. The Rise and Potential of Large Language Model Based Agents: A Survey. Science China Information Sciences, 2025, 68(2): 121101. [4] Kasneci, E., Sessler, K., Kuchemann, S., Bannert, M., Dementieva, D., Fischer, F., et al. ChatGPT for Good? On Opportunities and Challenges of Large Language Models for Education. Learning and Individual Differences, 2023, 103: 102274. [5] Li, M., Zhang, H., Wang, L. Applications and Challenges of Multimodal Large Models in Educational Scenarios. China Distance Education, 2025, 41(3): 45-54. [6] Hong, W., Ding, M., Zheng, W., Liu, X., Tang, J. CogVideo: Large-scale Pretraining for Text-to-Video Generation. In Proceedings of the 11th International Conference on Learning Representations (ICLR 2023), Kigali, Rwanda, May 1-5, 2023, pp. 1-24. [7] Liu, H., Li, C., Wu, Q., Lee, Y. J. Visual Instruction Tuning. In Advances in Neural Information Processing Systems 36 (NeurIPS 2023), pp. 34892-34916. [8] Hong, W., Wang, W., Lv, Q., Liu, J., Wang, W., Cao, J., et al. CogAgent: A Visual Language Model for GUI Agents. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2024), pp. 14221-14230. [9] Schick, T., Dwivedi-Yu, J., Dessi, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., et al. Toolformer: Language Models Can Teach Themselves to Use Tools. In Advances in Neural Information Processing Systems 36 (NeurIPS 2023), pp. 68539-68551. [10] Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., et al. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In Advances in Neural Information Processing Systems 35 (NeurIPS 2022), pp. 24824-24837. [11] Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., et al. Training Language Models to Follow Instructions with Human Feedback. In Advances in Neural Information Processing Systems 35 (NeurIPS 2022), pp. 27730-27744. [12] Liu, X., Yu, H., Zhang, H., Xu, Y., Lei, X., Lai, H., et al. AgentBench: Evaluating LLMs as Agents. In Proceedings of the 12th International Conference on Learning Representations (ICLR 2024), Vienna, Austria, May 7-11, 2024.
Copyright @ 2020-2035 STEMM Institute Press All Rights Reserved