Minimal Perturbation Adversarial Attacks on Malware Classification Systems

Abstract

This paper proposes a new approach to generate adversarial examples targeting malware classification models that treat binary executables as images. The method introduces carefully targeted obfuscation, altering specific bytes in the binary that translate into adversarial pixel perturbations in its image representation. By employing syntactic obfuscation techniques, the malware can be modified while fully retaining its functionality. Experimental evaluations show that this approach reliably misleads CNN-based detection systems, resulting in incorrect classifications. Furthermore, we examine the challenges related to selective obfuscation, especially when altered bytes overlap with non-instructional regions or structural components of the executable. Collectively, these findings shed light on the vulnerabilities of image-based malware detectors and contribute to the development of more robust defenses against adversarial threats.

Publication
2025 IEEE/WIC International Conference on Web Intelligence and Intelligent Agent Technology (WI-IAT)