Intelligent Cybersecurity Framework for Detecting and Explaining Malicious Prompt Injection Attacks in Large Language Models

Authors

  • Wilson Bakasa Author
  • Ranjan Arulanandham Author

DOI:

https://doi.org/10.1234/mwyfz002

Keywords:

Large Language Models, BERT, classification, MPDD dataset

Abstract

Large Language Models (LLMs) are increasingly vulnerable to malicious prompt injection attacks that manipulate intended instructions and compromise system security. This paper proposes an intelligent cybersecurity framework combining Modern BERT for contextual prompt classification with Integrated Gradients for explainable detection. The framework preprocesses prompts, generates contextual representations, performs malicious-prompt classification, applies threshold-based security decisions, and identifies influential tokens responsible for predictions. Using the MPDD dataset, the proposed Modern BERT model achieves 97.78% accuracy, demonstrating effective malicious prompt detection. The integration of explainable analysis improves transparency and supports informed security decisions, providing a practical approach for strengthening the protection of LLM-based applications against prompt injection attacks.

Downloads

Published

2026-07-25