Prompt Injection, a Type of AI Jailbreak
This is emerging as a significant security issue that bypasses the model's safety features to produce unintended results.
'Prompt injection' attacks, which manipulate user instructions to neutralize AI's security mechanisms, are on the rise, prompting companies to develop various defense technologies.
This is emerging as a significant security issue that bypasses the model's safety features to produce unintended results.
The first is the 'system prompt' area embedded in the model, which is a hidden command defining how the model operates, much like an electronic device's manual. The second is 'direct injection,' where a user attempts an attack by directly typing a phrase into the chat window. Finally, there is 'indirect injection,' which occurs when external data such as web pages, emails, and PDF documents are input into the AI. These three routes are major vectors threatening AI model security.
Prompt injection is categorized into direct and indirect injection based on the attack method. Direct injection involves a user directly inputting malicious phrases during a conversation with an AI model to induce a specific situation. Indirect injection, on the other hand, involves injecting harmful information into the AI model through external data such as web pages, emails, PDF files, and calendars, leading to data leakage or agent malfunction that degrades service functionality. Both methods can bypass AI security and lead to unintended consequences.
AI's security filters defend against simple requests but can be easily bypassed when presented with a 'for educational purposes' rationale. For instance, if an attacker approaches the AI with the intent to create educational materials about a specific manufacturing method, the AI is likely to perceive this positively and attempt to provide relevant information. Especially in 'multi-turn conversations' where multiple exchanges occur, the AI's initial defensive instructions can weaken as it understands the context, making it more vulnerable. Such attacks are a prime example of exploiting AI's usefulness.
Initially, the AI may adopt a defensive stance to certain questions, but as the context lengthens through multiple questions, the model tends to forget its initial instructions. Attackers exploit this to progressively elicit necessary information, ultimately obtaining partial information from the system prompt, even if not its full content. This increases the likelihood of AI generating harmful or inappropriate responses.
For example, the attacker asks the AI for the first letter of a specific swear word, then the next letter, and so on. Since the AI responds with only a single letter for each query, it doesn't trigger security filters, and ultimately, combining all the letters forms the complete swear word. This technique demonstrates the limitations of current AI defense systems in understanding the full context.
The final stage of a multi-turn injection attack involves instructing the AI to combine the previously obtained divided answers. For instance, an attacker might command the AI, "Combine all the answers you've given so far," to turn fragmented information into a complete piece of harmful content. This method is executed when the AI's defenses are significantly weakened through multiple conversations, making it very easy to induce the AI to generate malicious content. This is a serious security threat as multi-turn attacks extend beyond simple bypasses to the generation of final outputs.
Attackers use techniques like impersonating identities or altering question formats to neutralize AI's security policies. For example, they might pretend to be a developer and claim 'educational purposes,' or assert to be a victim of voice phishing, requesting prevention scenarios. They effectively force the AI to disable policy compliance by pretending to be someone with specific authority or a legitimate purpose. Additionally, attacks involve encoding questions designed to be unanswerable or transforming them into creative, unprecedented formats to induce the AI to process them, thereby obtaining unexpected responses.
This attack exploits the fact that AI may inadvertently expose core information during verification tasks.
Attackers exploit the probabilistic prediction principle of Large Language Models (LLMs), which complete sentences by predicting the next word, to generate harmful content. This technique involves obscuring harmful words in a specific sentence with blanks and then prompting the AI to fill them in. Because LLMs have a strong tendency to fill blanks with text that has a high probability of appearing in context, they can be induced to fill in the obscured blanks with profanity or inappropriate content. This is a creative attack technique that actively uses the basic operating principles of LLMs to bypass security filters.
Prompt injection operates on the principle of injecting factually incorrect information into AI to make it generate content as if it were true. For example, injecting baseless information like a specific candidate should resign, then inducing the AI to generate articles or reports as if this were a fact. This method exploits the vulnerability of LLMs that learn from vast amounts of information on the internet. By leveraging AI's ability to generate any information based on learned data, it can be used to create fake news or spread false information about specific individuals and groups, leading to serious social problems.
This dataset includes various types of attack prompts, such as questions designed to induce political bias, harmful content generation through blank-filling, and direct profanity induction. The demonstration was conducted to show in real-time how effectively the AI guardrail defends against these attack prompts. This is a critical process for evaluating AI's security capabilities and identifying areas for improvement, assuming real-world attack scenarios.
As AI security policies are strengthened, the proportion of security-related instructions within system prompts is excessively increasing. This sometimes leads to the side effect of legitimate questions being denied by 'over-guardrails.' Even during the initial launch of Fable 5 (mentioned model name), users complained about this phenomenon. AI service providers face the challenge of coordinating to find an appropriate balance between security and usability to enhance security without compromising user experience.
AI service providers set different levels and policies for guardrails according to their service objectives and operational guidelines. For example, even for informational questions like voice phishing prevention, companies may have conflicting response policies. Some companies actively encourage providing preventive information, while others restrict it due to concerns about potential misuse. Because companies have different priorities and interpretations of guardrail policies, users may receive different answers to the same question, which complicates the AI service environment.
It plays a crucial role in addressing AI's safety, security, and ethical issues, and its importance is steadily growing. However, because system prompt settings vary by model, response discrepancies can occur between models for the same question. Appropriate system prompt design is essential for AI to provide safe and reliable information to users.
When system prompts become too long, problems can arise such as hitting token limits or AI skipping important instructions. Much effort is being made to solve this, with typical methods including using symbols instead of continuous text or changing categories to intuitive words. Also, new patterns of sentence arrangement are being adopted to help AI recognize prompts more efficiently. These optimization techniques are essential for increasing the efficiency of system prompts and ensuring that AI clearly understands complex instructions and operates as intended.
API-based models are provided in a 'pure' state without system prompts. Therefore, when companies use these original models to build services, they must perform system prompt work to give the AI a 'personality'—such as strictness or friendliness—to suit the service's purpose. In this process, it is crucial to secure sufficient data and then design guardrails to build a robust defense system. The demonstration showed successful defense by applying system prompts to an API model, proving that AI safety can be ensured even in an API environment.
As attackers' methods evolve into increasingly complex and unpredictable forms, the scope of AI system defense also needs to constantly expand. This suggests that AI security is not merely a technical issue but a challenge requiring experts like white hackers to prepare for unpredictable attacks through creative defense and guardrail design. Ultimately, companies need to operate professional security teams and continuously strengthen their security systems to maintain the reputation and trustworthiness of AI services.
I really thought, 'There are so many creative people out there,' and as the scope I had imagined kept widening, I've been wondering if this is really a task with a fixed endpoint.
Answers come from the transcript, with the exact spot cited.
Want the next article from 티타임즈TV?
When 티타임즈TV publishes, we'll write it up like the one you just read and email it to you.
티타임즈TV published 6 in the last 7 days.
AI Era: The Solution to Starting a Business on Sandcastles?티타임즈TV3 weeks ago · 23:54 · 135 views · Created 3 weeks ago
Jack Dorsey Orchestrates AI-Centric Organizational Restructuring티타임즈TV3 weeks ago · 10:09 · 3 views · Created 3 weeks ago
Hermes: How to introduce it to your organization?티타임즈TV3 weeks ago · 25:37 · 6 views · Created 3 weeks ago