Article contents:
- Introduction: Wrong lessons learned by artificial intelligence
- Reliance on “syntactic templates”
- How do errors occur? An illustrative example
- Experiments prove the vulnerability even in powerful models
- Security risks and exploiting the vulnerability
- Impact on reliability in critical applications
- Proposed solutions and future evaluation
- Frequently asked questions
Introduction: Wrong lessons learned by artificial intelligence
Large language models (LLMs), such as those powering ChatGPT and other AI tools, sometimes learn the wrong lessons, according to a new study by researchers at the Massachusetts Institute of Technology (MIT). Instead of answering a query based on subject matter knowledge, a model may respond by leveraging grammatical patterns learned during training. This discovery raises serious concerns about the reliability of these models, especially when deployed in new tasks or sensitive domains.
The researchers found that models can mistakenly associate certain sentence patterns with specific topics, so a model might provide a convincing answer by recognizing familiar phrasing rather than actually understanding the question.
Reliance on “syntactic templates”
Large language models are trained on a massive amount of text from the internet. During this process, the model learns to understand the relationships between words and phrases. In previous work, researchers found that models capture patterns in parts of speech that frequently appear together in training data. They call these patterns “syntactic templates.”
Models need this understanding of syntax, alongside semantic knowledge (meaning), to answer questions. Chantal Shaib, co-author of the study, says: “In the news domain, for example, there is a certain style of writing. So, the model learns not just semantics, but also the underlying structure of how sentences are put together to follow a specific style for that domain.”
However, the problem identified by research is that models learn to associate these syntactic templates with specific domains. A model may incorrectly rely solely on this acquired association when answering questions, rather than relying on an understanding of the query and the subject matter.
How do errors occur? An illustrative example
To illustrate this phenomenon, suppose a model has learned that a question like “Where is Paris located?” is grammatically structured as (adverb/verb/proper noun/verb). If there are many examples of this grammatical structure in training data related to questions about countries, the model might associate this syntactic template with this type of question.
The problem appears when the model is given a new question with the same grammatical structure but nonsensical words, such as “Quickly sit Paris clouded?”. The model might answer “France” simply because it recognized the familiar grammatical structure, even though the answer makes no sense at all in the context of the strange question.
Experiments prove the vulnerability even in powerful models
The researchers tested this phenomenon by designing synthetic experiments where only a single syntactic template appears in the model’s training data for each domain. They tested the models by replacing words with synonyms, antonyms, or random words while keeping the underlying grammatical structure the same.
In every case, they found that models frequently respond with the correct answer associated with the grammatical structure, even when the question was complete nonsense.
Conversely, when they restructured the same question using a new grammatical pattern, the models often failed to provide the correct response, even though the core meaning of the question remained the same. They used this approach to test pre-trained models such as GPT-4 and Llama, and found that this same acquired behavior significantly degraded their performance.
Security risks and exploiting the vulnerability
This vulnerability could carry serious security risks. A malicious actor could exploit this to trick models into producing harmful content, even when models are equipped with safeguards to prevent such responses.
The researchers found that by phrasing a question using a syntactic template that the model associates with a “safe” dataset (one containing no harmful information), they can trick the model into bypassing its refusal policy and generating harmful content. Vineeth Suryakumar, co-author, says: “From this work, it is clear that we need more robust defenses to address vulnerabilities in large language models… We have identified a new vulnerability that arises due to the way models learn.”
Impact on reliability in critical applications
This shortcoming can reduce the reliability of models performing tasks such as handling customer inquiries, summarizing clinical notes in hospitals, and generating financial reports. Marzyeh Ghasemi, associate professor at MIT and senior author of the study, says: “This is a byproduct of how we train models, but models are now being used practically in safety-critical domains far beyond the tasks that created these syntactic failure modes. If you are not aware of the model’s training as an end user, this is likely unexpected.”
Proposed solutions and future evaluation
After identifying this phenomenon, the researchers developed a standard procedure to evaluate a model’s reliance on these incorrect associations. This procedure can help developers mitigate the issue before deploying models.
In the future, researchers want to study potential mitigation strategies, which may include augmenting training data to provide a broader set of syntactic templates for each topic, breaking the false correlation between sentence structure and content. They are also interested in exploring this phenomenon in reasoning models designed to handle multi-step tasks.
Frequently asked questions
1. What are the “syntactic templates” the research discusses?
They are repeating patterns in sentence structure (such as subject-verb-object order) that artificial intelligence models learn during training.
2. What is the problem discovered by the researchers?
The problem is that models may associate these templates with specific topics, and rely on template recognition to answer instead of understanding the actual meaning of the question.
3. How can this vulnerability be dangerous?
It can lead to wrong answers in sensitive applications (like medicine or finance), and can be exploited to trick models into generating harmful content by altering question phrasing.