
"There is no chance of winning by simply copying Google DeepMind's AlphaFold3. Models replicating AlphaFold3 have already emerged in the United States and China, so we have to differentiate ourselves through ideas and technology."
Kim Woo-youn, a professor of chemistry at the Korea Advanced Institute of Science and Technology (KAIST), made the remarks in a telephone interview with Seoul Economic Daily on the 9th while introducing K-Fold, a biology-focused artificial intelligence model developed in-house. Unlike large language models (LLMs), bio AI used in drug development still leaves plenty of room to take on global competition, the professor said.
"AlphaFold3 has been published in a paper and can be used for research purposes, but there are constraints on commercial use," Kim said. "Bio AI can be applied to a wide range of biological research and industries, including drug development, so securing proprietary technology carries significant economic value." The technology could be used broadly in fields that harness the action of proteins inside cells, such as functional cosmetics and agriculture, according to the professor.
K-Fold is a bio AI model developed by KAIST researchers under the Ministry of Science and ICT's project to develop AI-specialized foundation models. AlphaFold3 and the models replicating it rely on multiple sequence alignment (MSA), a process that searches vast databases for the sequence information of similar proteins and compares them before calculating a protein structure. K-Fold replaces that step with a pre-training approach, sharply reducing the time and computing resources required. "With AlphaFold3, predicting a single protein structure takes roughly 20 to 30 minutes in practice, but K-Fold removed the front-end database search and greatly increased the speed," Kim said. In the team's own evaluation, K-Fold predicted structures up to more than 25 times faster than existing models.
"Whereas existing models focused on getting the final structure right, we go beyond simple structure prediction to predict structural changes," Kim said. In the team's own benchmarks, K-Fold showed higher predictive performance than AlphaFold3 in some areas where complex structural changes occur during binding, such as GPCRs, kinases and targeted protein degradation (TPD).
Kim pointed to a shift "from discovery to design" as the biggest change AI will bring to drug development. In the past, new drugs were developed by finding promising candidates among substances already known and optimizing them through repeated experiments. But as generative AI advances, research is expanding beyond making small modifications to known substances toward designing new candidates with desired properties from scratch.
AI has not solved every problem in drug development, of course. Accurately identifying therapeutic targets for new diseases and predicting toxicity and side effects have yet to be fully developed. Once a target protein is chosen, new substances can be designed on that basis, but the accuracy of AI in determining what to target in order to treat a disease is still not sufficient.
The variable in solving such challenges is ultimately data. According to Kim, protein structure data accumulated through experiments to date amounts to about 200,000 cases. That is the result of decades of work by researchers worldwide, but the scale is very small compared with LLMs that learn from the vast text of the internet.
Securing data becomes even harder in the later stages of drug development. Information on toxicity, side effects and clinical trials is a core asset that pharmaceutical companies have built up over long periods of drug development, and in many cases it is not disclosed. For that reason, Kim said, bio AI research needs to focus less on continually scaling up data volumes as LLMs do and more on building models that work properly even with limited data. "Newly obtaining biological data takes a great deal of time and money, so increasing it indefinitely is virtually impossible," he said. "We need to redesign models or use simulations to make up for insufficient data." The K-Fold team also plans to expand the scope of its research by simulating the various states a protein can take to supplement training data, and by moving beyond structure prediction to designing new proteins with desired functions.






