Artifact-Guided Logit Fusion of Fine-tuned CNN and ViT Branches for AI-Generated Image Detection: Implications for AI Literacy in Informatics Education

Mochamad Rizal Fauzan, Yusuf Athallah Adriyansyah, Feri Adriyanto

Abstract

Background: The rapid advancement of generative artificial intelligence has made it increasingly difficult to distinguish authentic images from synthetic visual content, raising concerns for digital and AI literacy and academic integrity in higher education. Informatics curricula increasingly need concrete, evaluated examples of how AI-generated image detectors are designed, tested, and limited.
Purpose: This study proposes Artifact-Guided EfficientViT-V3 (AG EfficientViT V3), a hybrid detection framework integrating fine-tuned EfficientNetB0 and ViT-Tiny branches with an artifact-oriented branch through learnable logit-level fusion, and discusses its potential relevance as an instructional case study and a supporting tool for AI/digital literacy and academic-integrity practice in informatics education.
Method: Instead of direct feature concatenation, the model preserves the decision behavior of separately fine-tuned branches while adding an artifact-sensitive correction signal at the decision level.
Results: On the CIFAKE benchmark, AG EfficientViT V3 achieves 98.865% accuracy, 98.938% precision, 98.790% recall, 98.864% F1-score, and 0.999116 ROC-AUC, the strongest clean-test performance among all evaluated baselines and variants. Ablation results show that naive CNN-Transformer hybridization alone does not outperform the strongest single-branch baseline, whereas fine-tuned branch initialization combined with logit-level fusion does. Robustness evaluation shows strong performance under clean and JPEG-compressed conditions, while severe blur, resize degradation, and additive noise remain challenging, illustrating that detector accuracy is condition-dependent rather than universal.
Conclusion/Implications: The proposed pipeline and evaluation protocol can be adapted as teaching material in deep learning and computer vision courses, while integration into digital-literacy or academic-integrity workflows is framed as a potential application requiring further validation rather than an established outcome.

Keywords

AI-generated image detection; AI literacy; artifact-guided logit fusion; CIFAKE; informatics education.

Full Text:

PDF

References

Bayar, B., & Stamm, M. C. (2018). Constrained convolutional neural networks: A new approach towards general purpose image manipulation detection. IEEE Transactions on Information Forensics and Security, 13(11), 2691–2706. https://doi.org/10.1109/TIFS.2018.2825953

Bhattacharjee, A., Islam, K., Anan, K., Intesher, A., Fuad, A. A., Saha, U., & Imtiaz, H. (2026). CAE-Net: Generalized deepfake image detection using convolution and attention mechanisms with spatial and frequency domain features. Journal of Visual Communication and Image Representation, 115, 104679. https://doi.org/10.1016/j.jvcir.2025.104679

Bird, J. J., & Lotfi, A. (2023). CIFAKE: Image Classification and Explainable Identification of AI-Generated Synthetic Images. IEEE Access, 12, 15642–15650. https://doi.org/10.1109/ACCESS.2024.3356122

Cozzolino, D., Poggi, G., Corvi, R., Nießner, M., & Verdoliva, L. (2024). Raising the Bar of AI-generated Image Detection with CLIP. IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops, 4356–4366. https://doi.org/10.1109/CVPRW63382.2024.00439

Cozzolino, D., Poggi, G., Nießner, M., & Verdoliva, L. (2025). Zero-Shot Detection of AI-Generated Images. Lecture Notes in Computer Science (Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 15076 LNCS, 54–72. https://doi.org/10.1007/978-3-031-72649-1_4

Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., & Houlsby, N. (2021). An image is worth 16x16 words: Transformers for image recognition at scale. Proceedings of the 9th International Conference on Learning Representations (ICLR 2021). https://doi.org/10.48550/arXiv.2010.11929

Hendrycks, D., & Dietterich, T. (2019). Benchmarking neural network robustness to common corruptions and perturbations. Proceedings of the 7th International Conference on Learning Representations (ICLR 2019). https://doi.org/10.48550/arXiv.1903.12261

Kittler, J., Hatef, M., Duin, R. P. W., & Matas, J. (1998). On combining classifiers. IEEE Transactions on Pattern Analysis and Machine Intelligence, 20(3), 226–239. https://doi.org/10.1109/34.667881

Loshchilov, I., & Hutter, F. (2019). Decoupled weight decay regularization. Proceedings of the 7th International Conference on Learning Representations (ICLR 2019).https://doi.org/10.48550/arXiv.1711.05101

Meng, Z., Peng, B., Dong, J., Tan, T., & Cheng, H. (2024). Artifact feature purification for cross-domain detection of AI-generated images. Computer Vision and Image Understanding, 247. https://doi.org/10.1016/j.cviu.2024.104078

Najjar, A. A., Ashqar, H. I., Darwish, O. A., & Hammad, E. (2025). Detecting AI-generated text in educational content: Leveraging machine learning and explainable AI for academic integrity. arXiv preprint. https://doi.org/10.48550/arXiv.2501.03203

Ng, D. T. K., Leung, J. K. L., Chu, S. K. W., & Qiao, M. S. (2021). Conceptualizing AI literacy: An exploratory review. Computers and Education: Artificial Intelligence, 2, 100041. https://doi.org/10.1016/j.caeai.2021.100041

Ojha, U., Li, Y., & Lee, Y. J. (2023). Towards Universal Fake Image Detectors that Generalize Across Generative Models. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2023-June, 24480–24489. https://doi.org/10.1109/CVPR52729.2023.02345

Pal, A., Kruk, J., Phute, M., Bhattaram, M., Yang, D., Chau, D. H., & Hoffman, J. (2024). Semi-Truths: A Large-Scale Dataset of AI-Augmented Images for Evaluating Robustness of AI-Generated Image detectors. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, & C. Zhang (Eds.), Advances in Neural Information Processing Systems (Vol. 37, pp. 118025–118051). Curran Associates, Inc. https://doi.org/10.52202/079017-3748

Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., & Batra, D. (2017). Grad-CAM: Visual explanations from deep networks via gradient-based localization. Proceedings of the IEEE International Conference on Computer Vision (ICCV), 618–626. https://doi.org/10.1109/ICCV.2017.74

Tan, C., Zhao, Y., Wei, S., Gu, G., & Wei, Y. (2023). Learning on Gradients: Generalized Artifacts Representation for GAN-Generated Images Detection. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2023-June, 12105–12114. https://doi.org/10.1109/CVPR52729.2023.01165

Tan, M., & Le, Q. V. (2019). EfficientNet: Rethinking model scaling for convolutional neural networks. Proceedings of the 36th International Conference on Machine Learning (ICML), 97, 6105–6114.https://doi.org/10.48550/arXiv.1905.11946

Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., & Jégou, H. (2021). Training data-efficient image transformers & distillation through attention. Proceedings of the 38th International Conference on Machine Learning (ICML), 139, 10347–10357.https://doi.org/10.48550/arXiv.2012.12877

Wang, Z., Bao, J., Zhou, W., Wang, W., Hu, H., Chen, H., & Li, H. (2023). DIRE for Diffusion-Generated Image Detection. Proceedings of the IEEE International Conference on Computer Vision, 22388–22398. https://doi.org/10.1109/ICCV51070.2023.02051

Willie, M. M. (2025). Verifying artificial intelligence-generated images: Socio-technical approaches to authenticity. Computing and Artificial Intelligence, 3(4). https://doi.org/10.59400/cai3893

Wohlin, C., Runeson, P., Höst, M., Ohlsson, M. C., Regnell, B., & Wesslén, A. (2012). Experimentation in software engineering. Springer. https://doi.org/10.1007/978-3-642-29044-2

Wu, D., & Zhang, J. (2025). Generative artificial intelligence in secondary education: Applications and effects on students’ innovation skills and digital literacy. PLOS ONE, 20(5), e0323349. https://doi.org/10.1371/journal.pone.0323349

Yan, S., Li, O., Cai, J., Hao, Y., Jiang, X., Hu, Y., & Xie, W. (2025). a Sanity Check for Ai-Generated Image Detection. 13th International Conference on Learning Representations, ICLR 2025, 64081–64099.https://doi.org/10.48550/arXiv.2406.19435

Zhou, P., Han, X., Morariu, V. I., & Davis, L. S. (2018). Learning rich features for image manipulation detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 1053–1061. https://doi.org/10.1109/CVPR.2018.00116

Zhu, M., Chen, H., Yan, Q., Huang, X., Lin, G., Li, W., Tu, Z., Hu, H., Hu, J., & Wang, Y. (2023). GenImage: A Million-Scale Benchmark for Detecting AI-Generated Image. Advances in Neural Information Processing Systems, 36(NeurIPS).https://doi.org/10.48550/arXiv.2306.08571