Intelligent Cybersecurity Framework for Detecting and Explaining Malicious Prompt Injection Attacks in Large Language Models
DOI:
https://doi.org/10.1234/mwyfz002Keywords:
Large Language Models, BERT, classification, MPDD datasetAbstract
Large Language Models (LLMs) are increasingly vulnerable to malicious prompt injection attacks that manipulate intended instructions and compromise system security. This paper proposes an intelligent cybersecurity framework combining Modern BERT for contextual prompt classification with Integrated Gradients for explainable detection. The framework preprocesses prompts, generates contextual representations, performs malicious-prompt classification, applies threshold-based security decisions, and identifies influential tokens responsible for predictions. Using the MPDD dataset, the proposed Modern BERT model achieves 97.78% accuracy, demonstrating effective malicious prompt detection. The integration of explainable analysis improves transparency and supports informed security decisions, providing a practical approach for strengthening the protection of LLM-based applications against prompt injection attacks.