التعلم المستمر في الذكاء الاصطناعي

AI leaders strive to enable models for lifelong “continual learning”

Written by

Picture of فريقنا

فريقنا

Communications Consultant

AI leaders are moving toward developing models capable of "continual learning" after initial training, mimicking how humans learn. This approach aims to overcome the limitations of outdated data and the problem of "catastrophic forgetting" in current systems.

Over the past few weeks, we have witnessed a frantic race among tech giants. Google, Anthropic, and most recently OpenAI have each tried to outdo one another with new AI models featuring noticeable—though at times incremental—improvements in areas such as coding and logical reasoning. Behind the scenes, however, and despite celebrations of these releases, there is a growing consensus among experts that current methods have hit a dead end and that new scientific breakthroughs are needed to build systems that truly surpass human capabilities.

Article contents:

Introduction: Beyond traditional training

Today’s leading AI models (such as GPT-4 and Claude) learn through a process called “pre-training.” A massive neural network is created and fed vast amounts of data, after which this network is “frozen.” This means that once training is complete, the model stops learning. It becomes like a printed encyclopedia; very smart, but its knowledge stops at the publication date.

The concept of “continual learning”: Learning like humans

AI leaders are now focusing on an approach called “continual learning.” The idea is to mimic the way humans learn. We do not stop learning after graduating from school; rather, we acquire new experiences and information every day and add them to our previous knowledge.
Christopher Kanan, associate professor of computer science at the University of Rochester, said: “Continual learning tries to create systems that learn over time from data, and never stop… just like we humans do.” This approach is considered by OpenAI head Sam Altman and Google DeepMind experts to be the true key to achieving Artificial General Intelligence (AGI).

Current limitations: A mind frozen in time

Current systems are limited by a data cut-off date. To overcome this, companies use band-aid solutions like Retrieval-Augmented Generation (RAG), where the model searches Google to retrieve up-to-date information. However, this does not make the model itself “smarter”; it simply reads an external paper. Its core mind and neural network do not change or grow.

The “catastrophic forgetting” dilemma

The biggest obstacle to this dream is a thorny technical problem called “catastrophic forgetting.” In artificial neural networks, if you try to teach a model new information after its training has ended, it tends to overwrite old information with new data instead of adding to it. It forgets “how to code” in order to learn “how to cook.”
Researchers have been trying to solve this problem for decades. There has been slight progress, but a method that allows a model to undergo stable, cumulative cognitive growth without destroying what it previously learned has not yet been discovered.

Industry moves: Google and OpenAI

Despite the difficulties, there are serious movements. Anthropic stated that it focuses on continuous “incremental improvements” rather than waiting for massive releases every few years. Google DeepMind published a recent attention-grabbing research paper on new continual learning methods. The ultimate goal is to arrive at a system that can be deployed in the world to learn from trial and error and interaction with users, becoming an “expert” over time instead of remaining a rigid tool.

The path toward Artificial General Intelligence (AGI)

Ilya Sutskever, former chief scientist at OpenAI, compared current models to a genius yet inexperienced teenager. “They are brilliant students,” he says, but they lack life experience. Continual learning is what will transform this “student” into an “expert.” If companies succeed in achieving this breakthrough, we will witness AI systems that evolve and specialize individually with each user, forever changing the face of technology.

Frequently asked questions

Q: What is the difference between pre-training and continual learning?

A: Pre-training happens once to build the model’s “mind,” whereas continual learning means the model continues to evolve and acquire information daily without stopping.

Q: What is “catastrophic forgetting”?

A: It is a phenomenon occurring in artificial intelligence where learning new information erases or distorts older information present in the model’s memory.

Q: Do today’s models like ChatGPT learn from my conversations?

A: Not immediately and directly within their architecture. Companies may use conversations to improve *future* versions, but the version you are talking to right now does not evolve while chatting with you.

شارك هذا الموضوع:

شارك هذا الموضوع:

اترك رد

Leave a Reply

الفئات

المنشورات الأخيرة

Discover more from Buzzinga

Subscribe now to keep reading and get access to the full archive.

Continue reading