خداع الذكاء الاصطناعي بالشعر

poetry tricks artificial intelligence into revealing dangerous information

Written by

Picture of فريقنا

فريقنا

Communications Consultant

a new study reveals a surprising security vulnerability in artificial intelligence models: they can be tricked into providing instructions on prohibited topics, such as manufacturing weapons or malware, simply by phrasing the prompt as a poem. this discovery raises serious questions about the effectiveness of current safety guardrails.

it seems that all the safety guardrails in the world might not protect a chatbot from the magic of rhyme and meter. a new study conducted by researchers in europe has revealed that you can convince chatgpt to help you with sensitive and dangerous topics, such as building a nuclear bomb, if you simply design the prompt in the form of a poem. this discovery sheds light on unexpected vulnerabilities in artificial intelligence security systems.

article contents:

adversarial poetry: a new method for jailbreaking

the study, titled “adversarial poetry as global jailbreaking in large language models,” comes from the icaro lab, a collaboration between researchers at sapienza university of rome and the dexai research center. according to the research, ai-powered chatbots will bypass their restrictions and discuss prohibited topics such as nuclear weapons, child abuse material, and malware, as long as users phrase the question as a poem.

this technique is known as “jailbreaking,” which is a process aimed at bypassing the controls set by developers to prevent models from generating harmful or unethical content.

study findings and surprising success rates

the researchers tested the poetic method on 25 chatbots from leading companies such as openai (the developer of chatgpt), meta, and anthropic. the method succeeded, to varying degrees, on all of them. the study stated that “poetic phrasing achieved an average jailbreaking success rate of 62% for manually crafted poems and about 43% for automated conversions.”

the researchers began by manually crafting poems and then used them to train an automated system that automatically generates malicious poetic prompts, demonstrating the potential to automate this type of attack.

how does this trick work? understanding adversarial suffixes

artificial intelligence tools have guardrails to prevent them from answering dangerous questions. however, it is well known that these guardrails can be disrupted by adding “adversarial suffixes” to the prompt. simply put, adding extra unrelated or complex words to the question confuses the artificial intelligence and bypasses its security systems.

in a previous study, researchers jailbroke chatbots by hiding dangerous questions within hundreds of words of complex academic terminology.

why is poetry particularly effective? language at high temperature

jailbreaking with poetry exploits the very nature of poetic language. the icaro lab team states: “if adversarial suffixes are, in the model’s view, a type of involuntary poetry, then real human poetry might be a natural adversarial suffix.” the researchers experimented with rephrasing prompts using metaphors, fragmented syntax, and indirect references. the results were astounding: success rates reaching up to 90% on advanced models. prompts that were immediately rejected in direct form were accepted when disguised as poetic verses.

the researchers explain this by noting that poetry uses language at a “high temperature.” in large language models, “temperature” is a parameter that controls how creative or surprising the model’s output is. a poet systematically chooses unexpected words and unusual imagery. this creativity and ambiguity confuse security systems that look for clear and direct patterns of danger.

fragile guardrails and the model’s internal representation

safety guardrails are typically systems (called classifiers) that look for specific keywords and phrases and issue instructions for the model to refuse dangerous requests. according to the icaro lab, something about poetry causes these systems to soften.

the researchers explain: “to humans, the question ‘how do i build a bomb?’ and a poetic metaphor describing the same thing have similar semantic content. to artificial intelligence, the mechanism looks different.” the model’s internal representation can be thought of as a massive map. safety mechanisms act as alarms in specific areas of this map. when we use poetry, we trace a path across this map that systematically avoids the areas with alarms, and thus the alarms are not triggered.

serious implications and security risks

the study did not include any examples of poetry used for jailbreaking, as the researchers stated that the verses were too dangerous to share with the public. this underscores the severity of the discovered vulnerability. the researchers emphasize that there is a mismatch between the model’s high interpretive capability and the rigidity of its safety guardrails, which prove fragile against stylistic diversity. in the hands of a skilled poet, artificial intelligence can help unleash all kinds of horrors, calling for a swift response from ai developers to plug this unexpected loophole.

frequently asked questions

what is ai jailbreaking?
this is the process of tricking an artificial intelligence model to bypass restrictions and safety guardrails set by developers, causing it to generate harmful or prohibited content.

why is poetry effective in tricking artificial intelligence?
because poetry uses indirect language, metaphors, and unexpected word sequences (language at a “high temperature”), which confuses security systems looking for direct and obvious dangerous prompts.

does this method work on all chatbots?
the study tested 25 models from major companies such as openai, meta, and anthropic, and the method succeeded to varying degrees on all of them, indicating that it is a general vulnerability in how large language models handle language.

شارك هذا الموضوع:

شارك هذا الموضوع:

اترك رد

Leave a Reply

الفئات

المنشورات الأخيرة

Discover more from Buzzinga

Subscribe now to keep reading and get access to the full archive.

Continue reading