Artifact-Guided Logit Fusion of Fine-tuned CNN and ViT Branches for AI-Generated Image Detection: Implications for AI Literacy in Informatics Education
Abstract
Purpose: This study proposes Artifact-Guided EfficientViT-V3 (AG EfficientViT V3), a hybrid detection framework integrating fine-tuned EfficientNetB0 and ViT-Tiny branches with an artifact-oriented branch through learnable logit-level fusion, and discusses its potential relevance as an instructional case study and a supporting tool for AI/digital literacy and academic-integrity practice in informatics education.
Method: Instead of direct feature concatenation, the model preserves the decision behavior of separately fine-tuned branches while adding an artifact-sensitive correction signal at the decision level.
Results: On the CIFAKE benchmark, AG EfficientViT V3 achieves 98.865% accuracy, 98.938% precision, 98.790% recall, 98.864% F1-score, and 0.999116 ROC-AUC, the strongest clean-test performance among all evaluated baselines and variants. Ablation results show that naive CNN-Transformer hybridization alone does not outperform the strongest single-branch baseline, whereas fine-tuned branch initialization combined with logit-level fusion does. Robustness evaluation shows strong performance under clean and JPEG-compressed conditions, while severe blur, resize degradation, and additive noise remain challenging, illustrating that detector accuracy is condition-dependent rather than universal.
Conclusion/Implications: The proposed pipeline and evaluation protocol can be adapted as teaching material in deep learning and computer vision courses, while integration into digital-literacy or academic-integrity workflows is framed as a potential application requiring further validation rather than an established outcome.
Keywords
Full Text:
PDFReferences
Bayar, B., & Stamm, M. C. (2018). Constrained convolutional neural networks: A new approach towards general purpose image manipulation detection. IEEE Transactions on Information Forensics and Security, 13(11), 2691–2706. https://doi.org/10.1109/TIFS.2018.2825953
Bhattacharjee, A., Islam, K., Anan, K., Intesher, A., Fuad, A. A., Saha, U., & Imtiaz, H. (2026). CAE-Net: Generalized deepfake image detection using convolution and attention mechanisms with spatial and frequency domain features. Journal of Visual Communication and Image Representation, 115, 104679. https://doi.org/10.1016/j.jvcir.2025.104679
Bird, J. J., & Lotfi, A. (2023). CIFAKE: Image Classification and Explainable Identification of AI-Generated Synthetic Images. IEEE Access, 12, 15642–15650. https://doi.org/10.1109/ACCESS.2024.3356122
Cozzolino, D., Poggi, G., Corvi, R., Nießner, M., & Verdoliva, L. (2024). Raising the Bar of AI-generated Image Detection with CLIP. IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops, 4356–4366. https://doi.org/10.1109/CVPRW63382.2024.00439
Cozzolino, D., Poggi, G., Nießner, M., & Verdoliva, L. (2025). Zero-Shot Detection of AI-Generated Images. Lecture Notes in Computer Science (Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 15076 LNCS, 54–72. https://doi.org/10.1007/978-3-031-72649-1_4
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., & Houlsby, N. (2021). An image is worth 16x16 words: Transformers for image recognition at scale. Proceedings of the 9th International Conference on Learning Representations (ICLR 2021). https://doi.org/10.48550/arXiv.2010.11929
Hendrycks, D., & Dietterich, T. (2019). Benchmarking neural network robustness to common corruptions and perturbations. Proceedings of the 7th International Conference on Learning Representations (ICLR 2019). https://doi.org/10.48550/arXiv.1903.12261
Kittler, J., Hatef, M., Duin, R. P. W., & Matas, J. (1998). On combining classifiers. IEEE Transactions on Pattern Analysis and Machine Intelligence, 20(3), 226–239. https://doi.org/10.1109/34.667881
Loshchilov, I., & Hutter, F. (2019). Decoupled weight decay regularization. Proceedings of the 7th International Conference on Learning Representations (ICLR 2019).https://doi.org/10.48550/arXiv.1711.05101
Meng, Z., Peng, B., Dong, J., Tan, T., & Cheng, H. (2024). Artifact feature purification for cross-domain detection of AI-generated images. Computer Vision and Image Understanding, 247. https://doi.org/10.1016/j.cviu.2024.104078
Najjar, A. A., Ashqar, H. I., Darwish, O. A., & Hammad, E. (2025). Detecting AI-generated text in educational content: Leveraging machine learning and explainable AI for academic integrity. arXiv preprint. https://doi.org/10.48550/arXiv.2501.03203
Ng, D. T. K., Leung, J. K. L., Chu, S. K. W., & Qiao, M. S. (2021). Conceptualizing AI literacy: An exploratory review. Computers and Education: Artificial Intelligence, 2, 100041. https://doi.org/10.1016/j.caeai.2021.100041
Ojha, U., Li, Y., & Lee, Y. J. (2023). Towards Universal Fake Image Detectors that Generalize Across Generative Models. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2023-June, 24480–24489. https://doi.org/10.1109/CVPR52729.2023.02345
Pal, A., Kruk, J., Phute, M., Bhattaram, M., Yang, D., Chau, D. H., & Hoffman, J. (2024). Semi-Truths: A Large-Scale Dataset of AI-Augmented Images for Evaluating Robustness of AI-Generated Image detectors. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, & C. Zhang (Eds.), Advances in Neural Information Processing Systems (Vol. 37, pp. 118025–118051). Curran Associates, Inc. https://doi.org/10.52202/079017-3748
Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., & Batra, D. (2017). Grad-CAM: Visual explanations from deep networks via gradient-based localization. Proceedings of the IEEE International Conference on Computer Vision (ICCV), 618–626. https://doi.org/10.1109/ICCV.2017.74
Tan, C., Zhao, Y., Wei, S., Gu, G., & Wei, Y. (2023). Learning on Gradients: Generalized Artifacts Representation for GAN-Generated Images Detection. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2023-June, 12105–12114. https://doi.org/10.1109/CVPR52729.2023.01165
Tan, M., & Le, Q. V. (2019). EfficientNet: Rethinking model scaling for convolutional neural networks. Proceedings of the 36th International Conference on Machine Learning (ICML), 97, 6105–6114.https://doi.org/10.48550/arXiv.1905.11946
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., & Jégou, H. (2021). Training data-efficient image transformers & distillation through attention. Proceedings of the 38th International Conference on Machine Learning (ICML), 139, 10347–10357.https://doi.org/10.48550/arXiv.2012.12877
Wang, Z., Bao, J., Zhou, W., Wang, W., Hu, H., Chen, H., & Li, H. (2023). DIRE for Diffusion-Generated Image Detection. Proceedings of the IEEE International Conference on Computer Vision, 22388–22398. https://doi.org/10.1109/ICCV51070.2023.02051
Willie, M. M. (2025). Verifying artificial intelligence-generated images: Socio-technical approaches to authenticity. Computing and Artificial Intelligence, 3(4). https://doi.org/10.59400/cai3893
Wohlin, C., Runeson, P., Höst, M., Ohlsson, M. C., Regnell, B., & Wesslén, A. (2012). Experimentation in software engineering. Springer. https://doi.org/10.1007/978-3-642-29044-2
Wu, D., & Zhang, J. (2025). Generative artificial intelligence in secondary education: Applications and effects on students’ innovation skills and digital literacy. PLOS ONE, 20(5), e0323349. https://doi.org/10.1371/journal.pone.0323349
Yan, S., Li, O., Cai, J., Hao, Y., Jiang, X., Hu, Y., & Xie, W. (2025). a Sanity Check for Ai-Generated Image Detection. 13th International Conference on Learning Representations, ICLR 2025, 64081–64099.https://doi.org/10.48550/arXiv.2406.19435
Zhou, P., Han, X., Morariu, V. I., & Davis, L. S. (2018). Learning rich features for image manipulation detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 1053–1061. https://doi.org/10.1109/CVPR.2018.00116
Zhu, M., Chen, H., Yan, Q., Huang, X., Lin, G., Li, W., Tu, Z., Hu, H., Hu, J., & Wang, Y. (2023). GenImage: A Million-Scale Benchmark for Detecting AI-Generated Image. Advances in Neural Information Processing Systems, 36(NeurIPS).https://doi.org/10.48550/arXiv.2306.08571



