Artificial intelligence can write polished Arabic. The recent launch of the MBZUAI Arab culture AI benchmark aims to evaluate how well AI understands regional language and cultural nuances. That does not necessarily mean it knows how people in the Arab world actually speak.
Researchers at Abu Dhabi’s Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) are digging into that gap with a new benchmark designed to test AI models on Arab culture, everyday conversations and regional dialects.
The project, called ArabCulture-Dialogue, looks beyond Modern Standard Arabic and asks a more awkward question for today’s large language models: can an AI system understand the cultural setting of a conversation and then reply in a dialect that sounds like it actually belongs there?
So far, the answer is mixed.
AI Understands the Situation Better Than It Speaks the Dialect
MBZUAI’s research found something interesting. Strong AI models can often recognize the culturally appropriate response to an Arabic conversation, even when the dialogue shifts between Modern Standard Arabic and a local dialect.
Producing that dialect is another story.
When researchers asked models to translate into dialectal Arabic or continue a conversation using a particular national dialect, performance dropped significantly. A model might understand what is happening culturally while still replying in language that sounds too formal, too generic or simply from the wrong part of the Arab world.
That distinction matters because Arabic is not one uniform spoken language.
Modern Standard Arabic dominates formal writing, education and much of the media. Everyday conversation happens in regional dialects that can differ sharply in vocabulary, phrasing and pronunciation.
Training and testing AI mainly on formal Arabic can therefore create an illusion of fluency.
ArabCulture-Dialogue Covers 13 Arabic-Speaking Countries
The ArabCulture-Dialogue dataset includes conversations from 13 Arabic-speaking countries and covers both Modern Standard Arabic and national dialects.
According to MBZUAI, the benchmark spans 12 everyday subject areas and 54 more specific subtopics.
Those conversations are not limited to abstract language tests. They involve ordinary cultural situations involving weddings, food, parenting, agriculture, arts, games and other parts of daily life.
MBZUAI worked with 26 native Arabic speakers across the participating countries. The speakers rewrote and translated dialogues into the way people would naturally speak rather than relying on literal word-for-word translations.
That approach is important. Dialects are messy. People code-switch. Spellings vary. Expressions make sense in one city and sound strange somewhere else.
A benchmark built only from textbook Arabic would miss much of that.
Researchers Put AI Models Through Three Tests
The benchmark measures AI performance using three main tasks.
First, models have to select a culturally appropriate response to a conversation. Then they are tested on translation between Modern Standard Arabic and a particular dialect. Finally, the AI is asked to continue a conversation while staying inside the requested dialect.
Recognition was generally easier than generation.
Researchers reported that models were better at understanding customs shared broadly across Arab societies. More local traditions and expressions were tougher.
Emirati and several North African dialect conversations were among the more difficult cases.
The problem becomes especially obvious when models are told to sound local.
MBZUAI reported an example in which GPT-5 produced high-quality conversational text but was identified as using the correct national dialect only 45% of the time under a strict automatic dialect test. For some dialects, including UAE, Libyan, Sudanese and Yemeni Arabic, that strict score fell to zero in the reported experiment.
The sentences could still sound fluent. They just did not necessarily sound local.
Giving AI More Cultural Context Can Help
The research also suggests that AI models may already contain more cultural knowledge than their initial responses reveal.
When models were given clearer information about the country or region associated with a conversation, their accuracy improved.
That raises an interesting possibility. Some cultural failures may not come purely from missing information inside the model. The model may have trouble retrieving the right cultural knowledge at the right moment.
Prompting helps, but only to a point.
Fine-tuning models with dialect-specific data improved some results, particularly for dialects with distinctive vocabulary. Moroccan Arabic showed notable gains in one experiment. Other dialects remained difficult to separate cleanly from neighboring varieties or from Modern Standard Arabic.
Arabic AI Has a Data Problem
There is a larger issue behind the benchmark.
Arabic has hundreds of millions of speakers, but dialectal Arabic remains far less represented in the datasets commonly used to train language models than English or even Modern Standard Arabic.
Researchers working on Arabic NLP repeatedly run into the same problem: dialect data is comparatively scarce, inconsistent and full of regional variation.
A separate 2026 MBZUAI-related study adapting open-source language models for Egyptian, Moroccan, Palestinian, Saudi and Syrian Arabic described dialectal training data as limited, noisy and highly heterogeneous.
Another community-built dataset, Palm, collected cultural and linguistic material from all 22 Arab countries and similarly found uneven performance across countries and dialects. Some regions were represented considerably better than others.
More Arabic text alone will not necessarily solve this. The type of Arabic matters.
Why This Research Matters Beyond Translation
Getting dialects right is not just about making chatbots sound friendlier.
An AI assistant used for education, government services, healthcare information, customer support or local search can misunderstand users if it treats regional Arabic as badly written Modern Standard Arabic.
That becomes more important as AI systems move deeper into UAE government services and decision-making, where language and cultural context can affect how people interact with public-sector technology.
Culture adds another layer.
Expressions around hospitality, family relationships, celebrations, food and social etiquette often carry meaning that disappears in literal translation.
A system can technically translate the words while still missing the conversation.
That is what makes ArabCulture-Dialogue more interesting than another language benchmark. It tests whether AI can function inside the conversation rather than simply decode it.
MBZUAI Is Pushing Toward More Localized Arabic AI
The research also fits into the UAE’s wider effort to build domestic artificial intelligence capabilities.
The country has made AI a strategic priority through initiatives including the UAE National Strategy for Artificial Intelligence 2031, while regional researchers and companies continue developing Arabic-focused models and datasets.
The same push toward more localized AI can be seen in efforts to keep AI inference and data processing inside the UAE and in UAE-built AI platforms moving onto local and on-premise hardware.
For MBZUAI, the immediate lesson from ArabCulture-Dialogue is fairly clear.
Today’s AI systems can sometimes recognize Arab cultural context surprisingly well. Getting those same systems to speak naturally across the Arab world’s enormous range of dialects is still unfinished work.
And that gap may become much more noticeable as AI assistants move away from formal prompts and into normal everyday conversations.

