Evaluating AI Risk and Governance in Generative AI Systems: A Prompt-Level Analysis
Evaluating AI Risk and Governance in Generative AI Systems: A Prompt-Level Analysis
Ziya Anjum
JAMIA HAMDARD (Deemed to be University)
New Delhi 110062, India Centre for Distance and Online Education
Under the guidance of: Dr. Abdul Majid Farooqi
Assistant Professor at
Centre for Distance and Online Education, Dept. of CSE, Jamia Hamdard, New Delhi, India
---------------------------------------------------------------------***---------------------------------------------------------------------
Abstract - Generative Artificial Intelligence systems—particularly those built on Large Language Models (LLMs)—have become central to modern enterprise computing, yet they carry with them a class of vulnerabilities that traditional cybersecurity models were never designed to address. Decoder-only transformer architectures process system instructions and untrusted user inputs as a single undifferentiated sequence of tokens, which makes them susceptible to direct and indirect prompt injections, jailbreaking attacks, and retrieval poisoning. Existing governance frameworks—including the NIST AI Risk Management Framework, the EU AI Act, and ISO/IEC 42001—offer valuable compliance guidance at the organizational level, but they stop short of providing the kind of execution-level security blueprints that engineering teams actually need.
This paper introduces the Prompt-Level AI Risk Governance Framework (PLAIRGF), a four-phase architectural model covering Prevention, Detection, Response, and Continuous Improvement. Rather than depending on simulated metrics to evaluate the framework, we validate PLAIRGF through mathematical formalization of key risk indicators—including Attack Success Rate (ASR), Risk Severity Score (R), the Governance Readiness Index (GRI), and a newly proposed Human-in-the-Loop Alignment Coefficient (η)—combined with rigorous defensive capability mapping and step-by-step operational trace walkthroughs of representative adversarial attack scenarios. The paper concludes with a structured alignment between PLAIRGF controls and international compliance standards, offering organizations a practical, audit-ready foundation for deploying LLM-based systems securely.
Keywords: Generative AI Security, Large Language Models, Prompt Injection, AI Governance, Retrieval-Augmented Generation, Human-in-the-Loop, PLAIRGF