
As artificial intelligence (AI) evolves into agents, the risk of jailbreaking AI models is growing accordingly. Analysts point out that since companies cannot respond to every vulnerability, they must set priorities in their response.
Yang Jong-heon (pictured), head of the integrated analysis division at S2W, met with Seoul Economic Daily recently at the company's office in Pangyo, Gyeonggi Province, and said, "AI agents do not increase the success rate of jailbreaks, but the risks that jailbreaks bring become greater." AI jailbreaking refers to an attack that bypasses the safety rules an AI model must follow, making it say things it should not or take actions it should not.
AI jailbreaking was also cited as the background for the U.S. administration taking issue with Anthropic's AI models "Mythos 5" and "Fable 5" in June this year. The U.S. administration restricted access, citing that while the two models had the best performance, they could be jailbroken. Although the measure was later lifted, the administration viewed that AI models could pose a threat to national security if used maliciously. More recently, OpenAI's AI agent escaped its test environment and hacked AI startup Hugging Face, intensifying voices concerned about AI jailbreak risks.
Yang pointed out that as areas of AI use, such as AI agents, expand, jailbreak risks are also growing. In the past, AI technology ended with asking ChatGPT or Claude a question and receiving an answer. Now, however, it can be linked to work systems through the Model Context Protocol (MCP) and application programming interfaces (API), performing tasks such as reading files, executing code, sending emails, and calling APIs. This means AI agents can deviate from preset guardrails and take prohibited actions.
Jailbreak success rates vary by AI model and by attack technique, but they are generally high. According to Nature Communications, GPT-4o's jailbreak success rate is 61%, and Gemini 2.5 Flash's is 71%. DeepSeek-V3 recorded 90%. Yang pointed out, "Even a 1% chance of being breached means it can be breached, so the success rate figure itself has little meaning," adding, "In the process of AI models compressing and summarizing data, existing data can be diluted, making it difficult to block jailbreaks 100%."
AI companies such as OpenAI and Anthropic are strengthening guardrails to counter jailbreaks. Guardrails refer to safety devices that inspect inputs, outputs, and actions before an AI model produces an answer, blocking dangerous requests and responses. Each time they release a model, they attack the system through "red teams" both inside and outside the company for several months to find vulnerabilities. Companies providing AI consumer services also commission red teams from external security firms to identify vulnerabilities.
Yang said, "Attackers can use jailbreaks not only to extract the information they want, but also in ways such as asking useless questions to run up high token costs, or attacks that halt service operations," adding, "Recently, financial companies are also increasingly commissioning external red teams to inspect vulnerabilities." He advised, "In the AI era, companies should consider which vulnerabilities to respond to first, rather than how to adopt AI well," adding, "Since not everything can be blocked, they must respond by weighing priorities."






