By Cyber Defense Technologies February 24, 2026 5 min read
Organizations are deploying AI quickly: chat assistants for employees, customer-facing agents, document analysis, code generation, and machine learning models inside products and mission systems. Many of these deployments were built fast, and security has not always kept pace.
AI systems inherit the familiar risks of any software: vulnerable infrastructure, weak access control, exposed data. But they also introduce new ones, because their behavior is shaped by data and natural language rather than only by code.
Where AI systems can be attacked
-
Users and prompts
Prompt injection and jailbreaks that override instructions or extract data.
-
Application and agents
Tools and actions an AI agent can take on a user’s behalf, and the permissions behind them.
-
Model
Evasion, model extraction and theft of the model weights themselves.
-
Data and retrieval
Poisoned training data or documents that change what the model learns or retrieves.
-
Pipeline and supply chain
Compromised models, libraries and datasets pulled from outside sources.
-
Infrastructure
The cloud, compute and storage that host models and data.
Prompt injection
Applications built on large language models (LLMs) combine instructions from developers with input from users and other sources. Prompt injection occurs when an attacker supplies input that the model treats as instructions, overriding its intended behavior.
There are two forms:
- Direct prompt injection: a user types instructions designed to make the model ignore its rules, reveal its system prompt or produce prohibited output. Attempts to bypass safety rules are often called jailbreaks.
- Indirect prompt injection: the malicious instructions are hidden in content the model reads, such as a web page, an email, a document or a database record. When the model processes that content, it follows the attacker's instructions without the user ever seeing them.
Indirect injection becomes especially dangerous when AI systems act as agents: reading email, browsing the web, calling tools and taking actions. An injected instruction might cause an agent to forward sensitive data, change records or trigger transactions.
Filtering helps but cannot fully solve the problem, because the model processes instructions and data through the same channel. Defense therefore depends on limiting what a compromised model could do.
Data poisoning
Machine learning models learn from data. If an attacker can influence that data, they can influence the model. Poisoning can:
- degrade a model's accuracy overall
- plant a hidden backdoor that triggers specific behavior on specific inputs
- bias results in the attacker's favor
For systems using retrieval-augmented generation (RAG), where a model draws on a document collection when answering, poisoning the documents can change answers without touching the model at all.
Model evasion, extraction and theft
- Evasion: crafting inputs that cause a model to misclassify, such as altering malware so a detection model misses it, or modifying images so a vision model fails.
- Extraction: querying a model repeatedly to reconstruct its behavior, effectively copying it.
- Data extraction: coaxing a model to reveal sensitive information from its training data.
- Model theft: stealing model weights directly from storage or infrastructure. For organizations whose models represent significant investment or sensitive capability, this is a serious risk.
Supply chain risk
Organizations increasingly download pretrained models, datasets and libraries from public repositories. Model files can contain malicious code, and models can be modified to include backdoors. The AI supply chain deserves the same scrutiny as any software supply chain: known sources, integrity verification, scanning and controlled promotion into production.
Frameworks that help
- MITRE ATLAS catalogs adversary tactics and techniques against AI systems, much as MITRE ATT&CK does for traditional networks. It is useful for planning tests and describing findings.
- The OWASP Top 10 for LLM Applications identifies the most critical risks in LLM-based applications, including prompt injection, sensitive information disclosure, supply chain risks and excessive agency.
- The NIST AI Risk Management Framework (AI 100-1) provides a structure for governing, mapping, measuring and managing AI risk across an organization, and NIST's adversarial machine learning taxonomy describes attack types in detail.
Defending AI systems
Practical controls, most of them familiar security principles applied to a new kind of system:
- Least privilege for AI agents. Give an AI system access only to the data and tools it needs. An assistant that summarizes documents does not need permission to send email.
- Human approval for consequential actions. Require confirmation before an agent deletes, sends, pays or changes anything important.
- Treat model output as untrusted. Validate and sanitize output before it reaches other systems, such as databases, browsers or command interpreters.
- Separate trust levels. Keep untrusted content, such as external web pages and inbound email, away from systems with sensitive access, or process it with restricted capabilities.
- Protect the data pipeline. Control who can add or change training and retrieval data, and monitor for unexpected changes.
- Secure the model supply chain. Use vetted sources, verify integrity and scan model files before use.
- Monitor and log. Record prompts, outputs and tool calls, and watch for abuse patterns.
- Test adversarially. Include AI-specific attacks in penetration tests and red team exercises before deployment and after significant changes.
Governance matters too
Many organizations do not know how many AI systems they run, which data those systems can reach, or who is accountable for them. An inventory of AI systems, clear ownership, risk assessments and acceptable-use policies are the foundation. For government and defense systems, AI components must also fit into existing authorization processes such as the Risk Management Framework.
Questions to ask before deploying an AI system
- What data can the system access, and who can see its output?
- What actions can it take, and which require human approval?
- Where does untrusted content enter, such as email, web pages or uploaded documents?
- How are models, datasets and libraries sourced and verified?
- What is logged, and who reviews it?
- Has the system been tested adversarially, including prompt injection?
- Who owns the system, and how will it be updated and retired?
How CDT can help
CDT's AI/ML cybersecurity services include adversarial testing of AI systems mapped to MITRE ATLAS, secure architecture for ML pipelines and LLM applications, and AI governance aligned with the NIST AI Risk Management Framework.