Generative AI Security Guidelines
1. Overview
A. Concept of Generative AI
Generative AI is AI that learns the statistical patterns of training data through LLMs (large language models), diffusion models, and the like, and then generates new content on its own—text, images, audio, code, and more. In 2023 Korea's National Cyber Security Center (NCSC), under the National Intelligence Service, published the "Security Guidelines for Using Generative AI such as ChatGPT," presenting principles for safe use by the public sector and enterprises.
The fundamental difference between generative AI and conventional IT systems lies in the unpredictability of its behavior. Traditional software operates deterministically according to fixed logic, so the boundary between input and output is clear, and perimeter-based controls such as firewalls and access control fit well. Generative AI, however, produces output by probabilistically predicting the next token, so it can return different answers to the same input, and even the very definition of a "wrong output" is ambiguous. In other words, natural-language input is simultaneously a command and data, and the accuracy of output cannot be guaranteed in advance—this dual uncertainty creates a new attack surface.
Moreover, the fact that content a user enters in natural language can flow directly into model training, logs, and caches is critical from an information-security standpoint. In traditional systems "input" meant structured data that had passed form validation, but in generative AI a block of source code a developer pastes in or the full text of a customer complaint an agent relays all become prompts. As a result, new threats arise across the entire input–model–output pipeline, so perimeter-defense-centric legacy controls are insufficient, and dedicated guidelines tailored to generative AI's threat model are required.
B. Background and Necessity
The root reason guidelines are needed is that the pace of AI adoption has outrun the pace of building out security systems. Within just a few months of ChatGPT's release in late 2022, generative AI penetrated office, development, and support work, yet most organizations used it without any policy on "what may be entered" or "how far output may be trusted." Uncontrolled, autonomous adoption (Shadow AI) soon led to incidents such as confidential leaks, copyright infringement, and the incorporation of misinformation.
A second driver is the concretization of regulation and accountability. As the EU AI Act, Korea's AI Framework Act, and others impose transparency and risk-management obligations on high-risk AI, using generative AI is no longer a matter of convenience but one of legal compliance. Organizations must be able to demonstrate "who used what, under which controls," and this requires standardized usage and security guidelines to come first.
C. Examples of Services in Use
Guidelines are needed in practice because generative AI is already used widely in work, which in turn increases the points at which confidential information may be exposed. Each field below has a different threat profile. For example, the key risk in development is source-code leakage, while in customer-facing channels it is personal-data leakage and prompt injection.
| Field | Example | Main risk |
|---|---|---|
| Productivity | Document summarization/drafting, translation, minutes | Entry of trade secrets/undisclosed info |
| Development | Code generation/review, test automation | Leakage of source code/credentials (Secret) |
| Customer-facing | Chatbot support, FAQ auto-response | Personal-data leakage, prompt injection |
| Creative/Marketing | Image/design/copy generation | Copyright infringement, deepfake misuse |
Because the nature of risk differs by field, a uniform "ban everything" or "allow everything" policy is inappropriate in both cases. A blanket ban actually encourages Shadow AI (covert use), while blanket permission creates a control vacuum. Therefore, a risk-based approach that assesses the risk level per work domain and differentiates the scope of permission and strength of control becomes the starting point of the guidelines.
2. The Overall Structure of Generative AI Security Threats
Generative AI threats become systematically understandable when organized around the three points of the data flow (input, model/service, output) and the governance layer that surrounds them. The structure diagram below shows at a glance where each threat arises.
flowchart LR
subgraph GOV["Governance (policy·training·audit)"]
subgraph FLOW["Data flow"]
I["Input (prompt)"] --> P["Model/Service"]
P --> O["Output (generated content)"]
end
end
I -.sensitive input.-> R1["Data leakage"]
I -.malicious instruction.-> R2["Prompt injection"]
P -.training·RAG tampering.-> R3["Data poisoning"]
P -.repeated queries·API exposure.-> R5["Model theft·inversion"]
O -.factuality limits.-> R4["Hallucination"]
O -.abuse of generation.-> R6["Malicious content"]
At the input stage, if a user puts confidential or personal information into a prompt, that content may be stored in training data or logs and later reproduced or leaked (data leakage), and untrusted input may override system instructions to bypass controls (prompt injection). Injection in particular is evolving into indirect prompt injection, where instructions hidden in documents or web pages enter via RAG, so an attack can succeed even without the user directly entering malicious input.
At the model/service stage, tampering with training or RAG data can implant bias or backdoors (data poisoning), and repeatedly querying an API to replicate the model or infer the training data becomes a problem (model theft/membership inference). This simultaneously threatens the organization's intellectual property and the personal data used in training.
At the output stage, the intrinsic limit of statistical generation—hallucination—presents nonexistent facts plausibly and misleads decisions, and the generation capability is abused to create malware, phishing emails, and deepfakes. Hallucination is not a bug but a byproduct of probabilistic generation, so it cannot be completely eliminated and must be managed through evidence provision and verification.
| Threat | Main cause | Possible security threat |
|---|---|---|
| Data leakage | Confidential/personal info entered in prompt → stored in training·logs | Trade-secret/personal-data exposure, reproduced leakage |
| Prompt injection | Untrusted input/documents override system instructions | Bypassing instructions, privilege takeover, data leakage |
| Data poisoning | Tampering with training·RAG data | Bias·backdoors, inducing wrong answers |
| Hallucination | Factuality limits of statistical generation | Misinformation in work, decision errors |
| Misuse | Abuse of generation capability | Malware·phishing·deepfake generation |
| Model theft·inversion | API exposure·repeated queries | Model replication, training-data·membership inference |
For example, in 2023 a global manufacturer had a developer paste internal source code into a chatbot for review, and the incident in which confidential code flowed into an external model service became known, highlighting that input-stage data leakage is the most realistic threat. Afterward the company at one point banned internal use of generative AI entirely, then shifted its policy toward limited re-permission after building a closed, in-house model.
3. Security Considerations in Development/Use and Layered Countermeasures
The core of the response strategy is to build layered controls by input, model, output, and governance (Defense in Depth) matched to where threats occur. The architecture below shows how a user request is processed safely as it passes through multiple control gates.
sequenceDiagram
participant U as User
participant F as Input filter (DLP·masking)
participant G as AI gateway (auth·logging)
participant M as LLM/RAG
participant V as Output review (filter·evidence check)
U->>F: Enter prompt
F->>F: Detect·mask sensitive info
F->>G: Sanitized request
G->>M: Pass authenticated·isolated request
M->>M: Separate system/user prompts
M->>V: Generated result
V->>V: Harmful·hallucination check, watermarking
V->>U: Safe response
A. Input-stage control. This is the highest-priority line of defense. To keep sensitive information from entering the model in the first place, use DLP (data loss prevention) and regex/NER-based detection to mask resident registration numbers, card numbers, API keys, and the like. At the organizational level, establish input-ban policies and approval procedures that specify "what must not be entered," minimizing exceptions that bypass technical controls. Input-stage control is the most cost-effective because it blocks at the source rather than cleaning up after an incident.
B. Model/service-stage control. To prevent prompt injection, clearly isolate the system prompt from user input and process instructions and data separately, on the premise that user input is untrusted. If you use RAG, verify the source and integrity of documents before indexing to block indirect injection and poisoning. Also, to prevent model theft, apply authentication and rate limiting to the API and log every query to detect anomalous query patterns.
C. Output-stage control. Rather than trusting the generated result as is, place a filter/review layer. Inspect for harmful, discriminatory, or personal-data-containing content, and for work where factuality matters, enforce the provision of evidence (sources) so a human can verify hallucinations. Insert watermarking/provenance (e.g., C2PA) into generated content to trace deepfake misuse. The output stage is most effective when combined with a "human-in-the-loop" final review.
D. Governance layer. The organization's policy, training, and audit wrap around all the technical controls above. Usage policies and approval procedures, regular staff training, private models/network separation for sensitive work, and audits of usage history belong here. Since most data leakage stems from carelessness rather than malice, governance is the last safety net that reduces the "human error" technology cannot block.
A point to note when designing layered controls is that each layer must operate independently. Even if the input filter is breached, model-stage isolation filters out injection; and even if that fails, output review blocks sensitive-data leakage—composing defense in depth so that a single control failure does not immediately become an incident. Since generative AI continually spawns new threats, a design that relies on a single layer of control is dangerous.
| Category | Security consideration | Countermeasure |
|---|---|---|
| Input | Block sensitive-info ingress | Input filtering·masking, DLP, sensitive-info input-ban policy |
| Model/Service | Prevent injection·poisoning·theft | Prompt isolation (system/user separation), input validation, access control·logging, RAG data verification, rate limiting |
| Output | Control harmful·inaccurate results | Output filter·review, evidence provision, watermarking, hallucination verification, human-in-the-loop |
| Governance | Organization-wide control | Usage policy·approval, staff training, private model·network separation, audit |
4. Comparison with Conventional Information Security and Cases
Understanding where generative AI security differs from conventional information security makes clear why separate controls are needed. Traditional security protected the interior by drawing a "trust boundary," but in generative AI the natural-language input itself is a potential command, so an attack can succeed even inside the boundary. Also, traditional security dealt with confidentiality, integrity, and availability (CIA), whereas generative AI adds new axes of factuality (accuracy) and bias.
| Aspect | Conventional security | Generative AI security |
|---|---|---|
| Target | Data·systems (structured) | Data + model + output (unstructured) |
| Attack vector | Code vulnerabilities·network | Natural-language prompts·training data |
| Core properties | Confidentiality·integrity·availability (CIA) | CIA + factuality·bias·explainability |
| Control method | Perimeter·signature based | Layer·verification·governance based |
The difference to note especially in this table is the attack vector. In conventional security an attacker targeted code vulnerabilities or network paths, but in generative AI the normal usage path of "natural-language input" itself becomes the attack means. Malicious instructions may be embedded within normal traffic that has passed the firewall, so it is hard to judge traffic as right or wrong at the network layer. For this reason the center of gravity of control shifts from the network perimeter to the application·data·output verification layers.
As a concrete case, a public institution drafting complaint responses with generative AI cited a nonexistent statutory provision (a hallucination), leaving the lesson that results must not be trusted without output verification (evidence provision). Conversely, the financial sector is building model cases in which it builds closed LLMs in-house and logs·audits every query, taking the productivity gains while controlling the risk of data leakage.
Another case: as development-productivity tools spread internally, the problem emerged of developers inadvertently including secrets such as API keys and DB connection details in code-completion requests. In response, some firms introduced input filters that detect and block secret patterns at the IDE-plugin stage—a good illustration of the layered-defense principle of "block as close as possible to where the threat arises (the input stage)." The common lesson of the three cases is that generative AI security is not completed by a control at any single point; it becomes effective only when input blocking, output verification, and organizational policy work together.
5. Deep Dive: International/Domestic Standards and Latest Trends
Generative AI security is rapidly converging beyond individual corporate efforts into standardization and regulatory frameworks. From a professional engineer's perspective, one must be able to design a risk-based management system by linking the frameworks below.
- OWASP Top 10 for LLM Applications: A de facto standard summarizing threats specific to LLM applications, covering prompt injection, sensitive-information disclosure, supply chain, data poisoning, excessive agency, and more. As the era arrives in which AI agents autonomously call tools, the weight of the "excessive delegation of authority" threat is growing.
- NIST AI RMF (AI Risk Management Framework) 1.0: Manages AI risk through the four functions Govern·Map·Measure·Manage, and in 2024 added a separate Generative AI Profile (NIST AI 600-1) presenting control items for risks unique to generative AI.
- ISO/IEC 42001: An international standard for an AI management system (AIMS) that, like ISMS (27001), requires organizations to have a certifiable system for continuously managing AI risk.
- Regulation: The EU AI Act imposes transparency and documentation obligations through a risk-based (prohibited·high·limited·minimal) approach and is applied in phases. Korea's AI Framework Act (Act on the Development of AI and Establishment of a Foundation of Trust) stipulates transparency (labeling of generated content) obligations for high-impact and generative AI. However, the detailed enforcement decrees and criteria are still being refined, so the latest notices must be checked.
Thus standards are layered from "what to block (OWASP)" to "how the organization continuously manages (NIST·ISO)" to "what the state mandates (regulation)," and practice moves toward mapping these three and implementing them as the company's own controls.
A notable recent trend is that the center of gravity of threats is expanding from single-prompt attacks to autonomous agents and multimodality. An AI agent that calls tools and autonomously performs multiple steps has a large blast radius, since a single injection can lead to actual system manipulation (file deletion, payment, etc.). Multimodal models that handle images and audio also create a new attack surface, attempting injection via instructions hidden inside images. Therefore the defense system too must move beyond a text-centric approach to expand into agent action approval and multimodal input verification.
6. Considerations and Implications
- Parallel operation of the technical·policy·human triad: Technical controls such as DLP and prompt isolation alone are insufficient. Usage policy and staff training must go together for genuine defense. Since most data leakage stems from carelessness rather than malice, investing in people is itself a security investment.
- Trade-off between data sovereignty and deployment-model choice: For sensitive work, use closed/on-premises LLMs and RAG instead of external commercial APIs so data does not leave the organization. However, since in-house builds carry large GPU/operational costs, a hybrid strategy that differentially applies commercial API, private instance, and on-premises by data sensitivity is realistic.
- Regulatory·governance linkage: Link the EU AI Act, Korea's AI Framework Act, ISO/IEC 42001, and others to establish a risk-based management system, and reflect trustworthiness requirements such as content labeling and explainability in design up front (Security/Compliance by Design).
- Privilege minimization in the AI-agent era: For agents that autonomously call tools, "excessive agency" becomes a new threat. It is necessary to minimize the authority granted to agents and place human approval for risky actions.
- Making offense-defense continuous (ongoing red teaming): Adversarial-prompt and jailbreak techniques keep evolving, so red teaming and monitoring must run continuously—not as a one-off check—to keep updating defenses.
- Securing transparency·explainability: The content labeling (watermarking) and decision-evidence provision required by regulation must be reflected at service design time, not as an afterthought; this contributes both to user trust and to accountability for legal responsibility.
References
- Korea NIS National Cyber Security Center, "Security Guidelines for Using Generative AI such as ChatGPT" — https://www.ncsc.go.kr
- OWASP Top 10 for LLM Applications — https://owasp.org/www-project-top-10-for-large-language-model-applications/
- NIST, AI Risk Management Framework — https://www.nist.gov/itl/ai-risk-management-framework
- EU Artificial Intelligence Act — https://artificialintelligenceact.eu/
In one line: Generative AI security controls the threats of data leakage, prompt injection, data poisoning, hallucination, misuse, and model theft layer by layer across input (DLP), model (isolation·verification), output (filter·evidence·watermarking), and governance (policy·training·audit), centered on the parallel technical·policy·training triad, securing data sovereignty, and linking standards·regulations such as OWASP, NIST, and the EU AI Act.