AI Agents Raise Jailbreak Risks, Requiring Prioritized Response

S2W Director Yang Jong-heon: "Blocking AI Jailbreaks 100% Is Difficult"

Technology|
|
By Kim Ji-young
||
Yang Jong-heon, head of S2W's Integrated Analysis Division, explains AI jailbreak risks at the company's office in Pangyo, Gyeonggi Province. Photo courtesy of S2W - Seoul Economic Daily Technology News from South Korea
Yang Jong-heon, head of S2W's Integrated Analysis Division, explains AI jailbreak risks at the company's office in Pangyo, Gyeonggi Province. Photo courtesy of S2W

As artificial intelligence (AI) evolves into agents, the risk of jailbreaking AI models is growing accordingly. Analysts point out that since companies cannot respond to every vulnerability, they must set priorities in their response.

Yang Jong-heon (pictured), head of the integrated analysis division at S2W, met with Seoul Economic Daily recently at the company's office in Pangyo, Gyeonggi Province, and said, "AI agents do not increase the success rate of jailbreaks, but the risks that jailbreaks bring become greater." AI jailbreaking refers to an attack that bypasses the safety rules an AI model must follow, making it say things it should not or take actions it should not.

AI jailbreaking was also cited as the background for the U.S. administration taking issue with Anthropic's AI models "Mythos 5" and "Fable 5" in June this year. The U.S. administration restricted access, citing that while the two models had the best performance, they could be jailbroken. Although the measure was later lifted, the administration viewed that AI models could pose a threat to national security if used maliciously. More recently, OpenAI's AI agent escaped its test environment and hacked AI startup Hugging Face, intensifying voices concerned about AI jailbreak risks.

Yang pointed out that as areas of AI use, such as AI agents, expand, jailbreak risks are also growing. In the past, AI technology ended with asking ChatGPT or Claude a question and receiving an answer. Now, however, it can be linked to work systems through the Model Context Protocol (MCP) and application programming interfaces (API), performing tasks such as reading files, executing code, sending emails, and calling APIs. This means AI agents can deviate from preset guardrails and take prohibited actions.

Jailbreak success rates vary by AI model and by attack technique, but they are generally high. According to Nature Communications, GPT-4o's jailbreak success rate is 61%, and Gemini 2.5 Flash's is 71%. DeepSeek-V3 recorded 90%. Yang pointed out, "Even a 1% chance of being breached means it can be breached, so the success rate figure itself has little meaning," adding, "In the process of AI models compressing and summarizing data, existing data can be diluted, making it difficult to block jailbreaks 100%."

AI companies such as OpenAI and Anthropic are strengthening guardrails to counter jailbreaks. Guardrails refer to safety devices that inspect inputs, outputs, and actions before an AI model produces an answer, blocking dangerous requests and responses. Each time they release a model, they attack the system through "red teams" both inside and outside the company for several months to find vulnerabilities. Companies providing AI consumer services also commission red teams from external security firms to identify vulnerabilities.

Yang said, "Attackers can use jailbreaks not only to extract the information they want, but also in ways such as asking useless questions to run up high token costs, or attacks that halt service operations," adding, "Recently, financial companies are also increasingly commissioning external red teams to inspect vulnerabilities." He advised, "In the AI era, companies should consider which vulnerabilities to respond to first, rather than how to adopt AI well," adding, "Since not everything can be blocked, they must respond by weighing priorities."

Original reporting by Kim Ji-young for Seoul Economic Daily.

AI-translated from Korean. Quotes from foreign sources are based on Korean-language reports and may not reflect exact original wording.

Watch · Seoul Economic Daily

More →
3:15

AI KEY

Preview
Korean Corporate Intelligence HubKOSPI · KOSDAQ · 12 sectors

A live, cap-weighted view of every KOSPI and KOSDAQ sector, with same-day Korean reporting distilled by company — built for foreign investors, correspondents and analysts who need to scan Korea before the next session.

Korea Chaebol Tree

Preview
Families Behind the GroupsKFTC May 2026 · DART filings

An English-first interactive map of Samsung, SK, Hyundai, LG and Lotte — built for foreign investors, correspondents and analysts. Korea translates companies into English. We translate the families behind them.

SIGNAL

Pre-register
English Edition · Capital MarketsM&A · IPO · PE · Fund Flows

Pre-register for SIGNAL English Edition — a premium subscription bringing Korean capital markets coverage (M&A, IPOs, private equity, fund flows) to global institutional investors. First access to the 50% introductory rate.