AI Search Optimization in the Middle East: The Arabic Citation Gap

- #AIOverviews
- #AIsearchoptimization
- #ArabicSEO
- #ChatGPT
- #GenerativeEngineOptimization
- #MENASEO
Arabic has more than 400 million speakers, and the UAE has the highest generative AI adoption rate of any country on earth. Arabic accounts for 0.6% of the pages in Common Crawl. That mismatch between how many people in the Gulf ask AI assistants questions and how little Arabic source material exists for those assistants to cite is the defining opportunity in MENA digital marketing right now.
Key Takeaways
- The UAE ranks first globally for generative AI adoption at 70.1%, against a global average of 17.8% (Microsoft AI Economy Institute, May 2026). Qatar ranks tenth at 41.8%.
- Arabic is 0.6677% of the most recent Common Crawl. English is 40.45%, roughly 61 times larger (Common Crawl language statistics, CC-MAIN-2026-34).
- Multilingual retrieval systems suppress Arabic sources. Arabic queries recover 28.9 to 33.8 on character 3-gram recall with standard rerankers, against 51.4 to 53.6 with oracle evidence, a deficit of 18 to 21 points (Institute of Software, Chinese Academy of Sciences, April 2026).
- Modern Standard Arabic has zero native speakers and 335 million second-language users (Ethnologue 28). People ask in dialect. Models answer in Fusha. Your content has to serve both.
- Translated benchmarks flatter models; native Arabic exposes the real gap. GPT-4 scored 72.5% on native ArabicMMLU against a reported 80% on the translated MMLU (MBZUAI, February 2024), which is why every Arabic-first lab now trains on native rather than machine-translated data.
- None of the large citation-language studies published to date measures the source-language composition of AI search citations for Arabic queries. As of August 2026, no such measurement could be found.
The 70/0.6 Problem
The Gulf has the world’s highest demand for AI answers and one of the world’s thinnest supplies of source material to answer them with. Call it the 70/0.6 problem: 70.1% adoption in the UAE, 0.6% of the crawled web in Arabic. Every Arabic query an assistant receives has to be answered from a corpus that barely represents the language it was asked in.

The demand side is documented and it is extreme. Microsoft’s AI Economy Institute puts the UAE first in the world at 70.1% in its Global AI Diffusion report for Q1 2026, ahead of Singapore at 63.4% and more than double the United States at 31.3%. Qatar sits tenth at 41.8%. Saudi Arabia is at 29.4%, Jordan 29.7%, Oman 26.5%, Kuwait 21.1%. The global average is 17.8%. Microsoft estimates adoption from anonymised telemetry rather than surveys, so treat the absolute numbers as directional, but the ranking is consistent across editions and the UAE has held first place since H1 2025.
Saudi Arabia’s own regulator confirms the trend from a different method entirely. The Communications, Space and Technology Commission reported in its fifth Saudi Internet Report, published 19 July 2026, that 45.2% of Saudi internet users now use AI tools, more than double the previous year. ChatGPT is the most-downloaded AI app in the Kingdom, followed by Google Gemini, DeepSeek, and the national HUMAIN platform. Usage peaks at 55.7% among 20 to 29 year olds, and women use AI tools more than men, 52.8% against 39.4%.
Deloitte’s 2026 Digital Consumer Trends survey of 1,000 Saudi consumers found 66% actively using AI tools, up 17 percentage points in a year. The single most common use case was searching for information, at 51%. That is not a productivity story. That is search substitution, measured directly.
Similarweb’s July 2026 rankings put chatgpt.com as the third most-visited website in the UAE, behind only google.com and youtube.com. It ranks sixth in Qatar, seventh in Saudi Arabia, eighth in both Kuwait and Egypt.

Now the supply side. Common Crawl’s published language statistics for CC-MAIN-2026-34 show 14,286,482 Arabic pages, or 0.6677% of the crawl. English accounts for 865,530,140 pages, or 40.4526%. Arabic sits 21st on the list, or 20th among identified languages once an unlabelled bucket is excluded. Its share has been essentially flat for six years, moving from 0.5307% in early 2020 to 0.6677% now. W3Techs, measuring a different universe of sites, arrives at almost the same answer: Arabic is the content language of 0.6% of websites, against 49.5% for English.
Set that against Arabic’s speaker base. Ethnologue ranks Arabic the fifth most spoken language in the world, and the Qatar Computing Research Institute puts it at over 400 million native speakers. Against a world population near 8.2 billion, that is roughly 5% of humanity holding 0.67% of the indexed web, an underrepresentation of about seven to one.
Qatar Computing Research Institute put the constraint in plain terms in the Fanar 2.0 paper published March 2026, noting that “Arabic content represents only ≈0.5% of web data” as a core design problem it had to engineer around. The Jais 2 team at MBZUAI, Inception, and Cerebras wrote in July 2026 that Arabic “continues to be underrepresented in global training corpora” and that general-purpose models such as Llama 3, Gemma, and Qwen 2.5 “achieve only partial proficiency in Arabic.”
The people building Arabic language models say the corpus is too small. The people publishing in Arabic have not yet understood what that means for them.
What Happens When Someone Asks an AI Assistant a Question in Arabic
An Arabic query does not simply get answered from Arabic sources. Retrieval systems demonstrably prefer English and Latin-script documents, and they suppress Arabic sources even when those sources contain the better answer. This is measured, published behaviour, not speculation.

The clearest evidence comes from the Institute of Software at the Chinese Academy of Sciences, in a study published April 2026 on language bias in multilingual retrieval-augmented generation. Across 13 languages, more than 70% of the top five retrieved documents came from either English or the query language alone. For Arabic specifically, character 3-gram recall landed at 32.7 to 33.8 with a BGE reranker and 28.9 to 31.4 with a Qwen3 reranker. With oracle evidence, meaning the genuinely best available documents regardless of language, the same measure reached 51.4 to 53.6. That is a retrieval deficit of 18 to 21 points. The authors found that useful evidence “is distributed across multiple languages rather than dominated by any single one,” yet rerankers “systematically suppress such ‘answer-critical’ documents.” Worth noting the scope: this is a research retrieval pipeline on an open benchmark, not a measurement of any commercial AI search product.
A separate study from Chung-Ang University, published in Findings of ACL 2025, found the same pattern from a different angle. Retrievers show a strong preference for high-resource languages, English above all. Generators favour Latin-script languages over non-Latin ones, which puts Arabic script on the wrong side of the divide. Arabic scored 40.39 on their monolingual rank-shift measure with BGE-m3, against English figures often above 56. The authors note that these preferences do not consistently improve answer quality, so the bias is not earning its keep.
There is a mechanistic explanation underneath. Research from EPFL published in 2024 found that multilingual transformers decode a semantically correct next token in their middle layers but give “higher probability to its version in English than in the input language” before moving into an input-language-specific region. The abstract concept space, in the authors’ words, “lies closer to English than to other languages.” That work studied Llama-2 and did not isolate Arabic, so read it as a mechanism rather than an Arabic measurement.
Google confirms the practical shape of this asymmetry in its own product. When Gemini added Arabic support in May 2024, Google stated the model understands questions in more than 16 Arabic dialects and provides responses in Modern Standard Arabic. Dialect in, Fusha out. That single design decision shapes everything about how Arabic content gets retrieved and quoted.
One study touches Arabic citation behaviour directly. Profound analysed 3.25 billion citations, collected in March 2026 and published that April, across 14 countries and 10 languages including Arabic. In UAE and Saudi markets, Google AI Overviews showed what the researchers called a near-complete inversion of the English baseline: Instagram citations reached 29%, YouTube fell to 26% and lost its usual top position entirely, and LinkedIn nearly doubled. Profound sells AI visibility software, so this is a vendor study with disclosed methodology rather than peer-reviewed work, but it is the only large-sample dataset that includes Arabic markets at all.
Here is the honest boundary of what anyone knows. Two other large citation-language studies exist, from Temso AI in April 2026 and Weglot in August 2026. Neither includes Arabic. Temso covers Spanish, Dutch, German, Swedish, Italian and French. Weglot covers English, French, Spanish and Japanese. No published measurement of the source-language composition of AI search citations for Arabic queries could be found for this article. If someone quotes you a percentage for it, ask which study it came from.
The Dialect-In, Fusha-Out Rule
Arabic search behaviour splits across three registers that no other major market has to reconcile. People type queries in dialect or in Arabizi, models are strongest in Modern Standard Arabic, and models answer in MSA regardless of how the question arrived. Content that serves only one register loses at either retrieval or extraction.
Start with a fact that surprises most marketers: Modern Standard Arabic has no native speakers. Ethnologue lists MSA with zero L1 speakers and 335 million L2 users, because Arabic speakers first learn a local dialect and acquire MSA through formal education. The register that language models are optimised to produce is nobody’s mother tongue.
The performance cost of dialect is measured. DialectalArabicMMLU, published by IBM Research AI with NYU Abu Dhabi and MBZUAI in October 2025, tested 19 open-weight models on 3,135 questions per dialect across 32 domains. Average accuracy was 51.9% on Modern Standard Arabic and 47.7% across dialects, a gap of 4.2 percentage points. By dialect: Emirati 49.8%, Egyptian 48.9%, Saudi 48.2%, Syrian 46.6%, Moroccan 45.0%. The authors reported that performance “consistently declines across all dialects compared to MSA and English, and this trend holds consistently across all Arabic-enabled LLMs.”

QCRI’s AraDiCE benchmark shows the same slope on reading comprehension. Llama-3-8B-Instruct scored 0.85 accuracy in English, 0.74 in MSA, 0.71 in Levantine, 0.73 in Egyptian. The same work found that dialect-to-English translation consistently beats English-to-dialect, meaning models understand dialects better than they generate them.
Arabizi is worse. Researchers at the University of Geneva tested seven models on Franco-Arabic transliteration in 2025. GPT-4o, the best performer, scored BLEU 17.39 to 20.16 translating Arabizi into English but only 8.40 to 10.15 translating it into Arabic. Averaged across all seven models, the scores fell by dialect: Egyptian 9.65, Lebanese 7.52, Algerian 4.24. A significant share of young Gulf and Levantine users type in Arabizi, and models handle it roughly half as well in Arabic as in English.
The rule that follows is simple to state and rare to see implemented. Capture the question in the register people use. Deliver the answer in the register the model produces.
In practice that means an Arabic page carries dialect and Arabizi phrasings in its question layer, in H2s, FAQ entries, and the natural language of the opening lines, so that semantic retrieval has something to match against a real query. The answer itself is written in clean Modern Standard Arabic, because that is what the model is strongest at parsing and what it will produce when it paraphrases you. Pages built entirely in formal MSA miss the query. Pages built entirely in dialect get retrieved and then poorly extracted. Both halves have to be present.
Why Translating Your English Content Into Arabic Does Not Work
Translated Arabic and native Arabic are not interchangeable inputs to a retrieval system. Every benchmark that separates them shows translated content flattering model performance while native content exposes the real gap, and the labs building Arabic models have responded by excluding machine translation from their training data entirely.
The cleanest evidence is in the ArabicMMLU paper from MBZUAI, February 2024. GPT-4 scored 72.5% on ArabicMMLU, which is built from 14,575 native Arabic questions drawn from real curricula across eight countries. The GPT-4 technical report gives 80% for the English-to-Arabic translated MMLU. The ArabicMMLU authors attribute the eight-point gap to their dataset containing “a higher proportion of Arabic-specific content.”
Read that finding twice, because it is the whole argument. Translated Arabic measures whether a model can handle English concepts rendered in Arabic words. Native Arabic measures whether it can handle Arabic concepts, Arabic references, and Arabic framing. The second is what your customers actually ask about, and it is where models are weakest and where good content is scarcest.
The Technology Innovation Institute in Abu Dhabi reached the same conclusion from the model side. When TII built Falcon Arabic in May 2025, it trained on what it described as “100% native Arabic datasets,” explicitly avoiding machine-translated content, and added 32,000 Arabic-specific tokens to its vocabulary. Jais 2 built a custom 150,272-token vocabulary for the same reason. These are expensive engineering decisions made specifically because translated Arabic was not good enough.
There is a further reason to care about the quality of Arabic sources. Researchers at King Saud University, Prince Sattam Bin Abdulaziz University, and the Institute of Public Administration in Riyadh published a benchmark of Arabic RAG noise robustness in Frontiers in Big Data on 17 August 2026, testing six models across 300 questions and 6,196 documents. At zero noise, Claude-4 Sonnet scored 91.67% and GPT-4o scored 74.33%. At 80% noise, Claude-4 Sonnet degraded only 4.67 percentage points to 87.00%, while GPT-4o fell to 61.00%, a drop of 13 points.
Read that as a spread between models rather than a comparison with English, since the study has no English arm. The practical point stands either way: which Arabic sources get retrieved alongside yours changes the answer materially, and on some models it changes it by more than 13 points. In a corpus that is 0.67% of the web and heavily padded with machine translation, thin translated Arabic adds noise to a retrieval pool that is already noisy, and that noise degrades the answer your brand appears in.
The position this article takes, and it is a position rather than a consensus: most Gulf brands running an English-first content operation should stop translating and start commissioning. A smaller volume of Arabic content written by Arabic speakers, about Arabic-market realities, with Gulf-specific figures, regulations, and examples, will outperform a full translated mirror of the English site. The translated mirror competes in a category the models already handle. Native Arabic content competes in the category where they are starved.
Getting Eligible Before Getting Cited
Most Gulf sites fail AI search on eligibility, not on content quality. Three technical controls decide whether an assistant can see your pages at all, and two of them changed during 2026 in ways most published guidance has not caught up with.
The first is OpenAI’s crawler split, and it is routinely misunderstood. OAI-SearchBot is the agent that surfaces sites in ChatGPT’s search features. GPTBot crawls content that may be used for model training. They are independent settings. Blocking GPTBot to keep your content out of training does nothing to your ChatGPT Search visibility. Blocking OAI-SearchBot removes you from ChatGPT search answers, though OpenAI notes such pages can still appear as navigational links. Plenty of MENA sites that added a blanket AI-bot block in 2024 are currently invisible in ChatGPT and do not know it.
Two things changed here recently. On 9 December 2025, OpenAI removed the language requiring ChatGPT-User, the agent that fetches pages when a user asks for them, to comply with robots.txt. Any guide stating that all OpenAI bots respect robots.txt is now wrong. OpenAI has also added OAI-AdsBot, which validates the safety of pages submitted as ads on ChatGPT and appears in almost no published crawler list.
The second control is the one nobody in the region is checking. Google’s official generative AI optimization guide, last updated 10 July 2026, states that “in addition to the technical requirements for Search, a site must be included in Search generative AI features in Search Console to be eligible for display in generative AI features on Google Search.” This is a real gate with three settings, Include, Exclude, and Inherit from parent, covering AI Overviews, AI Mode, and generative AI in Discover. Google’s own framing is blunt: sites that opt out “will not receive traffic or impressions from our generative AI features.” The control launched for UK sites in June 2026 under regulatory pressure and was still rolling out to other countries through July. Check whether it is live on your property before assuming the default holds.
The related myth is worth killing directly. Blocking Google-Extended does not remove you from AI Overviews or AI Mode. Google states plainly that “Google-Extended does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.” It controls Gemini Apps and Vertex AI grounding only. AI Overviews and AI Mode are served from the Googlebot-built Search index. This has been misstated across the industry since 2023.
The third control is the set of crawlers most regional robots.txt files have never heard of. Meta-WebIndexer is Meta AI’s search crawler, the direct analogue of OAI-SearchBot, and it is missing from nearly every published AI crawler list. Amazon has split Amazonbot into Amzn-SearchBot and Amzn-User. Anthropic runs ClaudeBot for training, Claude-SearchBot for search indexing, and Claude-User for user-initiated fetches, and disabling Claude-SearchBot “may reduce your site’s visibility and accuracy in user search results.” PerplexityBot surfaces and links sites in Perplexity results and should be allowed if you want to appear there.
Two practical constraints round this out. Analysis of OpenAI’s retrieval stack by the agency RESONEO, published August 2026, found a hard 4 MB page size limit, above which pages are rejected outright rather than truncated. And Bing’s webmaster guidance states that sitemap lastmod values and IndexNow directly influence how quickly updates reach AI-generated answers, while changefreq and priority are ignored.
What the Evidence Says Actually Moves Citations
Google published its own generative AI optimization guide in May 2026 and its central claim is that there is no separate discipline. The guide addresses the industry’s own vocabulary directly, naming both answer engine optimization and generative engine optimization, then says: “From Google Search’s perspective, optimizing for generative AI search is optimizing for the search experience, and thus still SEO.” It goes on to name five popular tactics and tell site owners to stop.
One disclosure before the numbers. Almost every large citation study in this field is published by a company that sells SEO or AI visibility software: Ahrefs, Surfer, Semrush, Seer, Profound, Previsible, RESONEO. Their methodologies are usually disclosed and their sample sizes are real, but none of it is peer-reviewed and each of them has a commercial interest in the conclusion. Where a finding below comes from a vendor with a direct stake in it, that is noted. Google’s own documentation is the only primary source in this section.
The strongest published correlation in this field is query fan-out. Google’s guide defines it officially as “a set of concurrent, related queries generated by the model,” giving the example of a lawn weeds query fanning out into herbicide, chemical-free removal, and prevention sub-queries. Surfer analysed 10,000 keywords, 173,902 URLs and 33,000 fan-out queries in December 2025 and found that URLs ranking across fan-out queries are 161% more likely to be cited, with a Spearman correlation of 0.77 between fan-out rankings and AI citations. That is the highest correlation in any study reviewed for this article, and it matches Google’s own description of the mechanism. You do not rank for the prompt. You rank for the questions the prompt decomposes into.
Brand mentions correlate with AI visibility far more strongly than backlinks. Ahrefs studied 75,000 brands in May 2025 and found branded web mentions at Spearman 0.664 against backlinks at 0.218. This finding is repeatedly and incorrectly credited to Seer Interactive across the SEO industry. It is Ahrefs. Seer’s separate work points the same way on links: its strongest reported correlation is roughly 0.65 for Google page-one rankings, while its study of top news site mentions found roughly 0.07, which Seer itself described as almost no relationship.
Review profiles show the largest jump anyone has measured. Seer Interactive and Trustpilot analysed 804,491 AI responses across 1,926 brands and 15,783 prompts in May 2026, matched by domain rating cohort. Brands with no review profile had a median citation rate of 1%. Brands with a minimal profile, meaning 1 to 13 reviews, hit 53.5%. Nothing to something is associated with 52 percentage points. Good to excellent is worth about six. Two caveats belong with that number: Trustpilot co-authored a study about the value of Trustpilot profiles, and the data is correlational and self-selected, so brands that collect reviews may differ from brands that do not in ways the cohort matching does not capture. For Gulf brands with almost no review presence on international platforms, it is still the cheapest item on the list to test.
Now the things that do not work, each with the evidence against them.
Ranking in Google’s top 10 is a weaker enabler than the industry believes, and the number moved sharply. Ahrefs reported in July 2025 that 76% of AI Overview citations came from top-10 pages. Its own study of 863,000 keywords and 4 million AI Overview URLs, published March 2026, put the figure at 37.9%, with 31.2% from positions 11 to 100 and 31.0% from beyond position 100. For standalone assistants the relationship was never strong: across 15,000 long-tail queries, only 12% of AI assistant citations ranked in Google’s top 10, with ChatGPT in-text citations at 8.0% and Perplexity the outlier at 28.6%. The stale 76% figure is still widely republished, so check the date on any version you are quoted.
llms.txt does nothing measurable. Ahrefs checked server logs for 137,210 domains in June 2026. Of the 28% that publish an llms.txt file, 97% received zero traffic to it in May 2026. Of the requests that did arrive, 96% came from bots and only 19.5% from named AI tools, with SEO audit tools accounting for more of the traffic than AI assistants. No AI bot was observed probing for llms.txt files that do not exist. Google’s guide says outright that it “ignores them” and that creating one “will neither harm nor help your site’s visibility or rankings in Google Search.” Build it if you like. Do not bill it as a GEO tactic.
Schema markup has no proven effect on AI citation. The only controlled study is a matched difference-in-differences analysis by Ahrefs in May 2026 covering 1,885 pages that added JSON-LD between August 2025 and March 2026, against 4,000 control pages. Result: Google AI Overviews down 4.6%, AI Mode up 2.4%, ChatGPT up 2.2%, with the last two statistically indistinguishable from zero. Microsoft’s Fabrice Canel has said schema helps Microsoft’s LLMs understand content, and that is a reasonable argument for keeping it. The claim that FAQPage schema gets you cited by AI has no primary evidence behind it at all.
Chunking your content for AI is explicitly rejected by Google: “There’s no requirement to break your content into tiny pieces for AI to better understand it.” The tension here is real and worth understanding. Retrieval genuinely is passage-level. Dejan AI reverse-engineered Google’s grounding budget across 7,060 queries, by observing the snippets supplied to Gemini through the search-grounded API, and found roughly 2,000 words allocated per query, with a median of 377 words selected per source page, and coverage collapsing from 61% for pages under 1,000 words to 13% for pages over 3,000. The conclusion was “density beats length.” Retrieval is chunk-level, but authoring for chunks is not what makes chunks retrievable. Relevance and information density are.
One more caveat that the industry systematically drops. The original GEO paper from KDD 2024 reported gains of 41% from adding quotations and 31% from adding statistics. Those gains were measured among documents already in the retrieval set. The paper describes how to win a larger share of a generated answer once you are already a candidate. It says nothing about becoming a candidate. A 2026 critical survey found that on one end-to-end benchmark, body-only optimization actually reduced top-10 presence after reranking by 16%.
Measuring AI Visibility From the Gulf
Measurement improved substantially in 2026, and the tooling that matters most for MENA is the least discussed. Three surfaces now report something real, and each has a limitation you need to state out loud before anyone builds a dashboard on it.
Google Analytics 4 added an official AI Assistant channel on 13 May 2026. Recognised AI assistant referrers are now auto-assigned medium ai-assistant and channel “AI Assistant” in Default Channel Group reports. This obsoletes the whole genre of custom regex channel-group tutorials from 2024 and 2025, though those remain useful for backfill and for assistants Google does not recognise. Google names ChatGPT, Gemini, and Claude and has not published the full hostname list it classifies. OpenAI appends utm_source=chatgpt.com to referral URLs, which is the only vendor-confirmed referral tag available.
Search Console added generative AI performance reports on 3 June 2026, covering AI Overviews and AI Mode across Search and Discover. The limitation is severe: there is no clicks metric. Impressions only. Google says it is “continuing to work with website owners to understand what insights and data would be most helpful to inform their strategies.”
Bing Webmaster Tools offers the single most actionable measurement artifact in the market and almost nobody in the region uses it. Its AI Performance report, in public preview since 10 February 2026 and expanded in June, includes Grounding Queries, defined by Microsoft as “key phrases the AI used when retrieving content.” That is the rewritten fan-out query an AI system actually used to reach your page. Neither Google nor OpenAI exposes this. Given that fan-out ranking is the strongest correlate of citation anyone has measured, a free report showing you the actual fan-out queries is worth more than most paid AI visibility tools.
The dark traffic problem is worse than most dashboards admit. Google AI Mode sidebar links pass no referrer at all, so that traffic lands as direct and is structurally uncountable in GA4. Native mobile app sessions from ChatGPT, Perplexity, and Claude frequently arrive with no referrer. Previsible’s study of 6.77 million LLM-driven sessions across 166 GA4 properties, published July 2026, found ChatGPT accounting for 92.4% of trackable LLM referral traffic, and then added the most important caveat in the study: it measures standalone LLM referrals only and excludes AI discovery inside Google’s own results, which the authors say “almost certainly represents a larger volume of AI-driven traffic than all standalone LLM platforms combined.”
A May 2026 ChatGPT interface change replaced footnote citations with clickable brand names, and homepage referrals jumped from a 26 to 32% band to about 60% of referral traffic, where they have stayed. Page-level attribution broke at that moment. If your Arabic landing page data looks strange since May, that is why.
A 90-Day Plan for Arabic AI Visibility
Days 1 to 15, eligibility. Audit robots.txt for OAI-SearchBot, Claude-SearchBot, PerplexityBot, Meta-WebIndexer, and Amzn-SearchBot. Confirm the Search Console generative AI control is set to Include if it has rolled out to your property. Verify no Arabic template is producing pages above 4 MB. Confirm sitemap lastmod values are accurate and enable IndexNow.
Days 16 to 30, baseline. Turn on the GA4 AI Assistant channel and Search Console generative AI reports. Register Bing Webmaster Tools and start collecting Grounding Queries, which need history before they become useful. Run 40 to 60 real Arabic prompts, in dialect and MSA, across ChatGPT, Google AI Mode, Gemini, and Perplexity, and record who gets cited. This is your only Arabic baseline, because no public dataset exists.
Days 31 to 60, the question layer. Take your top 10 commercial pages and rebuild the question layer in the register your customers use. Gulf dialect phrasings and Arabizi variants in H2s and FAQ entries, MSA in the answers. Replace generic global figures with Gulf-specific ones: SAMA, GASTAT, CST, TDRA, and DataReportal numbers rather than US benchmarks. This is the information gain that Google describes as “a unique viewpoint that stands out.”
Days 61 to 90, off-site. Build the review profile, because 1 to 13 reviews is worth 52 percentage points of citation rate. Prioritise brand mentions over link building, because mentions correlate at 0.664 and links at 0.218. Map your fan-out queries from Bing’s Grounding Queries report and build content against the sub-queries rather than the head term.
What Other Articles Will Not Tell You About Arabic GEO
Three things follow from the research above that are not being said anywhere else, and the first one is the reason this whole article exists.
The 0.6% Advantage. Arabic’s underrepresentation is usually framed as a problem for Arabic speakers, which it is. It is also the least competitive citation environment in any major market. In English, a query has millions of candidate documents and the marginal value of one more well-structured page approaches zero. In Arabic, the candidate pool for a specific commercial query is often measured in dozens. The same scarcity that degrades Arabic answers is what makes a single high-quality Arabic page disproportionately likely to be the one retrieved. Publishers who move now are competing against a nearly empty corpus. That window closes as the corpus fills, and it is filling slowly: Arabic’s share of Common Crawl has moved less than two tenths of a percentage point in six years.
The MSA Paradox. Every major Arabic-capable model is optimised to produce a register that has zero native speakers. Google states Gemini understands 16-plus dialects and answers in Modern Standard Arabic. Benchmarks show dialect accuracy trailing MSA by 4.2 points. The result is that Arabic AI answers are, structurally, formal answers to informal questions. The brands that win are not the ones that write the most formal Arabic. They are the ones that bridge the two registers on the same page, matching the dialect at retrieval and the Fusha at extraction.
The measurement vacuum is a content opportunity. Three large studies have measured how AI citations behave across languages. None of them measured Arabic. Profound’s 3.25 billion citation study includes Arabic markets but reports platform mix rather than source language. That means any regional brand willing to run a structured prompt study, publish the methodology, and release the data becomes the primary source on Arabic AI citation behaviour. Original data is the one asset in this category with no substitute, and in Arabic the field is empty.
Frequently Asked Questions
Does optimizing for AI search require different work from SEO?
Google’s official position, published in its generative AI optimization guide in May 2026, is that it does not: “optimizing for generative AI search is optimizing for the search experience, and thus still SEO.” What the industry calls generative engine optimization is, on Google’s surfaces, SEO with a different set of eligibility checks. The practical differences that do exist are those controls, such as OAI-SearchBot access and the Search Console generative AI setting, and content structure that answers sub-queries rather than head terms. In Arabic markets there is one further difference that matters more than either: register, meaning the gap between how people ask and how models answer.
Should a Gulf brand publish in Arabic or English?
Both, but not as mirrors of each other. English content competes in a corpus where models are strongest and competition is heaviest. Arabic content competes in a corpus that is 0.6% of the web where models are weakest and competition is thinnest. The mistake is treating Arabic as a translation layer over the English site. Native Arabic content, about Arabic-market realities, is what benchmarks consistently show models handling worst and therefore what is most valuable to supply.
Is ChatGPT Search still powered by Bing?
Not as a blanket statement in 2026. OpenAI runs its own crawler, OAI-SearchBot, and research by the agency RESONEO published in August 2026, covering 1,249 answers and 88,000 search results, identified an in-house index alongside several retrieval pipelines that vary by account tier and query type. URLs from that in-house index showed only 1.5% overlap with Bing’s top 20 for the same queries, though RESONEO’s other observed pipelines overlapped far more with Google. OpenAI describes its own setup only as “sometimes partners with other search providers,” naming Microsoft and Shopify in its privacy documentation without confirming what supplies what. Treat any confident claim about ChatGPT’s index, including this one, as provisional.
Does llms.txt help with AI visibility?
There is no evidence that it does. Ahrefs checked server logs across 137,210 domains in June 2026 and found that 97% of published llms.txt files received zero traffic in May 2026. Google’s documentation states that Google Search ignores them and that publishing one “will neither harm nor help” visibility. No major AI company has confirmed that its retrieval systems read third-party llms.txt files.
How do you track AI referral traffic from the Gulf?
Use the GA4 AI Assistant channel introduced in May 2026, Search Console’s generative AI performance report for impressions, and Bing Webmaster Tools for Grounding Queries. Then account for what none of them capture: Google AI Mode sidebar links pass no referrer, native app sessions often pass none, and Search Console reports impressions with no clicks metric. Any Arabic AI visibility baseline in 2026 has to be supplemented with manual prompt testing, because no public dataset covers Arabic queries.
Which AI assistants actually work in Arabic?
Google AI Overviews launched in MENA and in Arabic globally in May 2025. Google AI Mode launched in MENA in English in August 2025 and in Modern Standard Arabic on 8 October 2025. Gemini has supported Arabic since May 2024, understanding 16-plus dialects and replying in MSA. ChatGPT is available across most Arab League states, including Saudi Arabia, the UAE, Egypt, Qatar and Kuwait, though OpenAI has published no Arabic-specific announcement for ChatGPT Search. Perplexity’s API accepts Arabic as a search language filter, but the company has published no Arabic interface or MENA launch.
Working With Voctos
Voctos builds AI visibility programmes for brands in Saudi Arabia, the UAE, and the wider Arab market, combining technical eligibility work, native Arabic content, and prompt-level citation tracking across ChatGPT, Google AI Mode, Gemini, and Perplexity. The mechanics that apply in any language are covered in the GEO guide; this article is the part that only applies in Arabic. The Arabic baseline nobody publishes is the one worth building first.
Don't miss the chance to
make your website more visible!
Initial consultation and
audit of the current situation
Read also
Our cases
All casesTrusted by






















































































































