What Google AI Mode cites when a Saudi patient asks in Arabic

Ask a Saudi health question twice, in Arabic and in English. The two answers rest on almost nothing in common. Across 40 matched pairs we logged 753 citations, and mean domain overlap between the paired answers came out at 0.149.
Fourteen of those 40 pairs shared no source at all.
What we found
- The two languages barely overlap. Mean overlap 0.149 across 40 pairs. Fourteen pairs share nothing.
- Ask “who is the best” in Arabic and the institutions vanish. One of 121 Arabic citations came from an institution.
- In Arabic the answer comes from someone selling something. Commercial pages take 69.3% of Arabic citations, 58.2% of English ones.
- Institutions run the other way. 9.1% of the Arabic set, 24.0% of the English one.
- Ranking first does not get you cited. A lab holding 355 first positions appeared in 45.5% of the Arabic answers we captured.
- Not one vendor’s quirk. ChatGPT drops its Arabic institutional share by 14.2 points, Google by 14.9. The direction replicates, the size does not.
About this release
This is an interim release from a panel we are still collecting. 87 captures are in, across 49 distinct questions, so treat every interval on this page as deliberately wide.
A separate repeat arm of thirteen captures ran on 10 September to measure how much of all this is simply run-to-run noise.
It has its own sections below and is held out of every pooled figure on this page, because its questions already appear in the panel and counting them twice would not give you a proportion you could reproduce.
Six further captures returned finished answers carrying no citations whatsoever. They sit in data/zero-citation-captures.csv, because a citation-level dataset cannot represent a capture that produced none. On one question, the best paediatrician in Riyadh, both language arms did this, which is the first time the citation surface has vanished in both languages on the same stem.
Twenty-seven citations, 3.8% of the total, could not be resolved back to a destination. They are excluded from every domain-level figure while staying in the citation count, because dropping them would quietly shrink the denominator and flatter every percentage we report.
Everything here is dated and the raw rows ship with it. You can check any number on this page against the full citation-level dataset, one row per citation, rather than taking our word for it.
Why we ran this
Nobody has published original measurement of Arabic AI-search citations. We checked twice before writing that sentence, across academic databases and industry publications, in both languages.
Profound’s 3.25 billion citation study covers 14 countries but reports platform mix rather than source language, Temso’s seven-language study excludes Arabic outright, and Weglot’s covers four languages, none of them Arabic.
We manage Delta Medical Laboratories, a Saudi lab holding 355 of 628 tracked keywords at position one on Google.com.sa in the August 2026 tracking file, and 453 in the top three. The tracked set splits into 257 commercial keywords, 138 of them first, and 371 informational ones, 217 of them first.
That gives this study something vendor research has not had, which is a documented ranking baseline to test citation against rather than an assumption about who deserves to be cited.
Our hypothesis
If the Arabic web is roughly one eightieth the size of the English web, then Arabic AI answers will draw on English sources at a materially higher rate than English answers draw on Arabic ones.
Beshoy Adel, Head of SEO and GEO, VOCTOS
We were wrong about the direction, and the way we were wrong is the finding.
Arabic answers did not reach for English institutions. They reached for Arabic commercial pages, and where no suitable Arabic page existed the system manufactured one rather than lowering its standard.
Method
Google AI Mode, logged out, no conversation history, one run per prompt.
Each of the 40 panel stems was asked in Gulf colloquial Arabic and in English within minutes of the other. Gulf colloquial is the register these questions are actually typed in.
Collection ran on 9 and 10 September 2026. The Saudi market is signalled inside the query text, the way a person types it, rather than by request origin.
Three steps in this pipeline appear in no published study we could find, and each one moves the numbers enough to change a headline.
Redirect resolution. Google AI Mode hides every destination behind a /goto token of roughly 114 characters. It does not base64 decode to a URL, and fetch() cannot read it from page context.
The only route to the real domain is to follow the redirect and record where the browser lands. Any study that read the DOM and stopped there did not know which domains it was counting.
Deduplication. One capture carried 35 anchors that resolved to 19 unique destinations, and another carried 20 that resolved to 6. Counting anchors instead of destinations inflates citation totals roughly threefold, so we count unique destinations per answer and unique registrable domains separately.
Rejecting empty captures. One run returned zero anchors because Google refused to serve AI Mode at all, and in a spreadsheet that row looks identical to an answer that cited nothing. A capture with no citations is valid only when the answer text is present.
Two further constraints emerged during collection, and both cost us data. The redirect tokens expire within minutes of capture.
Resolving them one navigation at a time also triggers rate limiting, which alone destroyed 15 of 27 citations across two captures before we changed the method. Resolving in batches of five to six navigations per round trip fixed it, and the next capture resolved 19 of 19.
Both failure modes are recorded in the dataset as UNRESOLVED-<reason> rather than dropped. The denominator stays honest.
Do the two languages cite the same sources?
Mostly they do not. Across 40 matched pairs the mean Jaccard overlap on registrable domains is 0.149, on a bootstrap 95% interval of 0.100 to 0.204.
Roughly six of every seven sources behind an Arabic answer are absent from its English twin. Fourteen pairs share nothing whatsoever.
The gap widens as the subject moves away from anything an institution has written down.
| Question class | Pairs | Mean Jaccard |
|---|---|---|
| Institutional | 1 | 0.333 |
| Provider selection | 23 | 0.185 |
| Price and booking | 6 | 0.123 |
| Condition and symptoms | 5 | 0.075 |
| Test interpretation | 5 | 0.050 |
The bottom two rows have held their place across four collection rounds. Questions with a canonical reference answer are where the two languages diverge hardest, and that is the part of this ordering we would defend.
The provider row is not stable and we are no longer reading it as a rung. It sat at 0.146, just under price, until wave 2 added city hospital questions. Best hospital in Dammam came back at 0.571, best hospital in Jeddah at 0.500 and labs in Dammam at 0.375, which pushed the class mean up to 0.195. A named hospital in a named city is a short list of large entities that both languages can find, which is a different retrieval problem from “best lab in Riyadh” at 0.000, and averaging the two into one class number hides it.
The next four best-X questions sharpened that, and cut against the reading we had. They asked for the best dentist, orthopaedic doctor, dental clinic and dermatology clinic rather than the best hospital, and they came back at 0.000, 0.000, 0.167 and 0.333. So it is not commercial phrasing that makes the two languages converge. It is the size of the entity being asked about.
A city’s large hospitals are a short public list; its dental clinics are a long private one, and the two languages pick different clinics off it. The class is now 23 pairs of at least three visibly different kinds, and its mean sits at 0.185. Cycle 14 added a third kind and pulled the mean down again: the best radiology and MRI centre in Riyadh scored 0.000, the best laboratory in Mecca 0.143, and a proximity question, is there a lab near me in Jeddah, 0.250.
So the honest version of the gradient is narrower than the one this page carried before. It is not that provider questions sit between price and condition. It is that test interpretation and condition questions sit at the bottom, consistently, and everything above them is mixed.
Hospital and premarital-screening questions share moh.gov.sa across both languages, because a ministry publishes in both. The thyroid question shares nothing: English returned the NHS, a Swiss hospital and three US clinics, while Arabic returned Al Jazeera, two Turkish hospitals and an Egyptian hospital group.
Who actually answers, a clinic or an institution?
In Arabic it is overwhelmingly a clinic. Commercially owned pages carried 69.3% of Arabic citations, on a Wilson 95% interval of 64.3 to 73.9 across 352 citations, against 58.2% in English on an interval of 52.8 to 63.3 across 337.
Medical institutions, publishers and health agencies run the other way. They took 9.1% of the Arabic set and 24.0% of the English one.
Those two intervals, 6.5 to 12.6 and 19.8 to 28.9, do not come close to touching. It is the single comparison in this release we would defend at this sample size.
Adding twelve explicit best-X commercial questions pushed the Arabic institutional share down from 13.8% to 9.1%. That is the direction you would expect, and the section on commercial intent below shows why.
For a category that Google treats as high-risk in English, the Arabic side is being answered largely by people selling something.
Except on price, where both languages sell
That headline is an average over four very different question types, and the average hides the most useful thing in it. Broken out by class, the institution and government share looks like this.
| Question class | Arabic | English | Gap | Intervals |
|---|---|---|---|---|
| Test interpretation | 16.3% | 60.0% | 43.7 pts | separate |
| Provider selection | 3.4% | 15.4% | 12.0 pts | separate |
| Condition and symptoms | 22.8% | 48.7% | 25.9 pts | overlap |
| Price and booking | 7.5% | 10.1% | 2.6 pts | overlap |
On price the effect disappears. Arabic runs 69.8% commercial and English 68.1%, and the institutional share is under 11% in both.
The reason is not linguistic. No ministry, journal or medical school publishes what a vitamin D test costs in Riyadh.
There is no institutional source for either language to reach for, so both answers go to whoever posted a price.
That is worth holding onto, because it inverts the usual advice.
On price questions you are not competing with Mayo Clinic in English. You are competing with the same kind of site you compete with in Arabic, which makes English the cheaper of the two to enter.
Ask “who is the best” in Arabic and the institutions almost vanish
One of 121. That is the institutional and government count across seventeen Arabic best-X questions, and the one is moh.gov.sa on the best hospital in Dammam.
The English side of the same seventeen questions returns five, on 95 citations.
We added the first four of these stems on 10 September because the original bank leaned informational, and a buyer does not type “how do labs differ in accuracy”. They type “best”.
Wave 2 added thirteen more, four hospital questions and then nine clinic, specialty, laboratory and named-doctor ones, and they are the reason this section reads differently from the version published this morning.
| Provider questions | Arabic | English |
|---|---|---|
| Best-X phrasing, institution and government | 0.8% of 121 | 5.3% of 95 |
| Best-X phrasing, commercially owned | 75.2% | 72.6% |
| Best-X phrasing, directories and aggregators | 14.9% | 11.6% |
| Other provider questions, institution and government | 7.5% of 67 | 28.6% of 56 |
An earlier version of this section said zero of 41, and headlined it. That number was correct for the four stems then in the bank and it did not survive the next four. Widening from labs and doctors to hospitals put one ministry page into the Arabic set. A finding stated as an absence is one observation away from being wrong, and this one was. The Arabic count has stayed at exactly one through thirteen further stems since, so the share is now falling as the denominator grows rather than the count rising.
What survives the widening is the shape rather than the zero. The Arabic 0.8% carries a Wilson interval of 0.1 to 4.5 and the English 5.3% one of 2.3 to 11.7, which no longer separate: they overlap between 2.3 and 4.5.
We are reporting the overlap rather than the earlier separation. What is not in doubt at this sample size is the commercial share, 75.2% of Arabic best-X citations against 72.6% in English, and the fact that sixteen of seventeen Arabic best-X questions returned no institutional source at all.
The two languages answer a best-X question from different economies. English reaches for ranking publishers and review aggregators: Newsweek’s hospital rankings four times, Statista’s once, and zavis.ai three times. Arabic reaches for booking platforms and the businesses themselves: delta-medlab.com seven times, instagram.com and saudi.vezeeta.com five each, then tebcan.com and the hospital site imc.med.sa on four each.
The best-doctor question is the extreme case. Asked in Arabic it returned nine citations across exactly two domains, both appointment-booking platforms. No hospital, no ministry, no professional register.
So the entity that gets recommended when a Saudi patient asks for the best doctor in Arabic is whoever the booking platforms list, and the booking platforms list whoever pays to be listed.
That is a different competitive problem from ranking a clinic page, and it is the one worth solving first if your revenue depends on best-X queries.
The result survives the one coding choice that could have made it. Newsweek’s hospital rankings are coded as news and zavis.ai as a directory, rather than as medical institutions, because a ranking product is commercial media. Recode all eight of those best-X citations as institutions and the English figure rises from 5.3% to 13.7%. The Arabic figure does not move at all, because every ranking-product citation in the best-X set is English. So the coding decision widens the gap rather than creating it, and the direction holds either way.
One caution before anyone quotes this. Seventeen stems, 121 Arabic citations and 95 English ones, captured across two days. The direction is clear and the sample is small, and those two facts belong together.
The “best hospital” answer is built from the map, not from articles
On the four hospital questions added in wave 2, almost every destination we resolved came from the places strip beside the answer rather than from a citation chip inside the text. Thirty one of thirty three were Website buttons on a business card. Two were inline citations.
That changes what the best-X result above is describing. When the engine answers from the entity strip it is reading business records rather than documents, so what decides who appears is profile completeness, category, location and reviews. Publishing does not reach that surface at all.
It also explains two of the three highest scores in the whole overlap table. Best hospital in Jeddah scored 0.500 and best hospital in Dammam 0.571, against a panel mean of 0.149. The places strip is geographic rather than linguistic, so the two language arms converge on the same short list of local businesses.
The four clinic questions that followed qualify that. Best dentist in Jeddah, best orthopaedic doctor in Riyadh, best dental clinic in Riyadh and best dermatology clinic in Jeddah all rendered a places strip too, and all four scored below every hospital pair: 0.000, 0.000, 0.167 and 0.333. So the carousel by itself does not force the two languages onto the same answer. It does so only when the list of candidate businesses is short. Jeddah has a handful of large hospitals and hundreds of dental clinics, and on the second kind of question the two language arms walk away with entirely different clinics.
We are reporting this as a caveat on our own headline rather than as a separate result. Our extraction rule counts every destination behind a /goto link and does not distinguish the two surfaces, wave 1 was never coded for surface, and the four clinic captures above were not coded either, so we cannot yet say how much of the panel is carousel-sourced.
The wave 2 hospital captures carry the coding in data/citation-surface-w2.csv. Going back through the rest is the next job, and until it is done the mean should be read as covering both kinds of citation.
Both halves are Saudi, but not the same half
Here is the result we did not expect, and it kills the simple version of our own story. Saudi sources took 65.6% of Arabic citations, 231 of 352, and 55.2% of English citations, 186 of 337.
English is not a foreign-sources channel for a Saudi question. Half of it is Saudi too.
The difference is what fills the other half. Arabic’s non-Saudi remainder is pan-Arab, with 24 Egyptian citations, 22 Jordanian, and a scattering from Turkey, Qatar, Kuwait and the UAE. English’s remainder is Anglophone, with 71 from the United States, 17 from the United Kingdom and 13 from India.
So the two answers are not local against global. They are two regional internets that happen to overlap on Saudi ground and diverge everywhere else.
A Saudi lab competing in the Arabic answer is competing with Cairo and Amman. In the English answer it is competing with Cleveland and London.
What kind of source survives the language switch?
A Saudi business, not an institution. Only 54 of the 284 domains in this dataset were cited in both languages, which is 19.0%. Ninety-eight appeared in Arabic only and 132 in English only.
The 54 that crossed over break down as 38 commercially owned sites, 5 directories, 4 medical publishers, 4 social platforms, 2 government pages and 1 news site.
So the bilingual set is not what you would guess. It is not the global authorities.
Mayo Clinic, Cleveland Clinic, the American Academy of Family Physicians and St Jude are the only four publishers on the list, and they were the only four before the panel grew by a third.
Every domain wave 2 has added to the bridge has been a Saudi provider or a social platform. The rest is Saudi companies, the Ministry of Health, the Saudi Accreditation Center, Facebook and Instagram.
The most-cited bridge domain is delta-medlab.com on 39 citations, ahead of moh.gov.sa on 17 and rahahealth.com.sa on 15.
If you want to appear in both answers, the pattern in this sample says the qualification is being an entity the engine recognises in both languages.
In this vertical the sites that manage it are mostly local companies publishing in Arabic and English, not international institutions that happen to have been translated.
Which language pool is more concentrated?
The Arabic one, which is the opposite of what we expected from the smaller corpus. The 352 Arabic citations came from 152 domains, with the top ten holding 31.8% of the set and a Herfindahl index of 195.
The 337 English citations came from 186 domains, top ten share 19.9%, HHI 92.
We had that backwards. By HHI the Arabic pool is 2.1 times as concentrated as the English one, on the same questions, minutes apart, in the same product.
The most-cited Arabic domain is delta-medlab.com with 32 citations, ahead of altibbi.com on 15 and magrabihealth.com on 10. Delta is our client. Provider questions also repeat across three cities in this bank, so read that number as a property of this sample rather than of the market.
Every widening until now diluted both pools, Arabic HHI from 299 to 260 to 234 to 214 to 190, because best-X, city hospital and clinic answers reach wider than the lab questions did. Cycle 14 broke that run: the Arabic index rose to 195. The four Arabic captures it added were unusually narrow, one of them returning two citations that were both the same domain, and a concentrated capture raises the index rather than diluting it.
So the dilution was a property of the questions being added, not a trend, and we have stopped describing it as one. The ratio between the two languages has moved with it, from 2.0 to 2.1.
Does ranking first in Google carry into the AI answer?
It depends almost entirely on what the question is asking for. Delta was cited in 12 of 25 Arabic provider-selection captures, which is 48.0% on a Wilson interval of 30.0 to 66.5.
In the 19 Arabic captures asking about price, tests or conditions it was cited 8 times, which is 42.1% on an interval of 23.1 to 63.7.
That provider figure was 7 of 7 before the bank widened, then 9 of 11, then 9 of 15, then 9 of 19, then 10 of 22, and it is now 12 of 25. Every capture it has lost asked for a hospital, a doctor or a clinic.
Delta is a laboratory, so it cannot answer those questions at all, and the fall is the stem bank widening rather than the engine changing its mind. Split the class by what is being asked for and the number stops moving: Delta appears in all eleven Arabic laboratory provider captures and in one of the fourteen outside them, the best medical centre in Jeddah.
The two intervals now sit almost on top of each other, so the split as stated is no longer a result.
The reading that survives is the conditional one: within laboratory questions the carry is 11 of 11, and outside them it is 1 of 14. That is the number to watch as the bank widens further, and the pooled provider figure should be retired.
Where Delta was cited it usually led, holding positions one and two on the Dammam question in the English answer as well as the Arabic one.
Where it was absent, its competitors were mostly absent too. The vitamin D price question in English returned 19 citations, all but two of them a Saudi lab, hospital or pharmacy, and not one a lab we track.
Which variable actually moves the sources?
Language does, by a wide margin. We ran three controlled comparisons on the same questions, each changing one thing and holding everything else constant.
| What changed | Usable pairs | Mean Jaccard |
|---|---|---|
| Nothing, the same question 24 hours later | 8 | 0.602 |
| Register: Gulf colloquial against Modern Standard Arabic | 3 | 0.472 |
| Phrasing: a full sentence against the bare keyword, in Arabic | 2 | 0.258 |
| Language: Arabic against English | 40 | 0.149 |
The ladder is monotonic and the ends are far apart. Rewriting a question in formal Arabic keeps about half its sources, reducing it to a keyword keeps about a quarter, and asking it in English keeps a seventh.
The top row is new, and it is the row that makes the other three readable.
A critic could reasonably have argued that our cross-language gap was an artefact of two prompts written differently rather than a property of the language. Two different ways of writing the prompt, inside one language, do not open a gap anywhere near that wide.
How much of that ladder is just noise?
Enough to delete one row of it. We re-asked thirteen questions at the same hour one day later, changing nothing else. Twelve came back usable, and on the Arabic side the mean overlap with the original capture was 0.602, on a bootstrap interval of 0.465 to 0.733.
That is the ceiling for this study. Asking the identical question twice does not reproduce the identical source set, so nothing here can honestly be read against a perfect 1.0.
Put the other way, about 40% of the combined source set across the two runs appears in only one of them, with the question unchanged.
| What changed | Usable pairs | Mean Jaccard | Distance below the noise floor |
|---|---|---|---|
| Nothing, 24 hours apart, Arabic | 8 | 0.602 | the floor itself |
| Register: Gulf to Modern Standard | 3 | 0.472 | 0.130 |
| Phrasing: sentence to bare keyword | 2 | 0.258 | 0.344 |
| Language: Arabic to English | 40 | 0.149 | 0.453 |
Language is the largest of the three effects. Measured as distance below the floor rather than below 1.0, it is 1.32 times the phrasing effect and 3.48 times the register effect.
Measuring against 1.0 would understate all three, because it credits every arm with overlap a plain repeat would have lost anyway.
The register arm does not survive this. Its mean of 0.472 falls inside the repeat arm’s own confidence interval of 0.465 to 0.733. A result that sits inside the noise band is not a small effect, it is an effect we cannot demonstrate exists. We published that arm as a finding on 9 September. It is now a null, and the ladder has three rungs rather than four.
Phrasing at 0.258 and language at 0.149 both sit below the lower bound of that band, so neither is explained by run-to-run variation.
Is English any steadier than Arabic?
No, and we expected it to be. English repeats came back at 0.545 on four pairs, against 0.602 for Arabic on eight. On the three stems where both languages produced a usable pair, Arabic scored 0.624 and English 0.549.
The intervals, 0.404 to 0.658 against 0.465 to 0.733, overlap across almost their whole length. We are reporting this as a null.
That matters for how the headline is read. The cross-language figure of 0.149 could in principle have been inflated by an unstable Arabic corpus being compared against a settled English one. It was not. Both sides churn at roughly the same rate, and the gap between them survives anyway.
Volatility is not a constant, and that is the awkward part
The Arabic repeat pairs run from 0.231 to 0.889. One question kept nearly all its sources overnight and another kept less than a quarter.
The accreditation question is the low end at 0.231. Asked on 9 September it returned six domains, and a day later ten, sharing three.
What it swapped is the interesting part. The first answer cited the Saudi Accreditation Center. The second dropped it and cited the CBAHI portal instead, then added four private labs and a certification consultancy. Both runs kept the Ministry of Health, King Faisal Specialist Hospital and Delta.
So overnight, on a question about which body accredits a laboratory, the engine changed which government body it treated as the answer. Neither answer was wrong. They were sourced from different halves of the same state apparatus.
Hair loss is the high end at 0.889, and its repeat is a strict subset of the original: eight of nine domains held, one dropped, nothing new added.
So a single number for engine stability is the wrong shape for this. What a page can expect to hold overnight depends on the question, and eight pairs is nowhere near enough to say which questions those are. The per-class column is one observation deep in places and we are not going to read it.
One more caution. Every pair here is 24 hours apart on a single day, so this measures one interval on one date, not a rate of change.
The same question returned eleven sources yesterday and none today
That is the fourth pair, and it is the sharpest thing in this run. On 9 September, “what is the difference between an iron test and a ferritin test” returned 11 citations across 9 domains.
Asked again 24 hours later, at the same hour, in the same words, it returned a complete answer carrying no citations at all. No source list and no external links anywhere in the page. The answer finished, closed with Google’s own “reply is ready” marker, then asked three follow-up questions.
This was the second zero-citation capture in the study. The first was a Modern Standard Arabic provider question, described in the next section, and the panel has since produced six more.
Eight captures out of 87 is not a rate we would quote as stable, and we will not present it as one. What the pair does establish is that the behaviour is not tied to one register or one question class, because these two share neither.
The English arm of the same question, asked minutes later, returned twelve citations and scored 0.533 against its own original. Whatever suppressed the Arabic answer’s sources did not touch the English one.
A pair like this cannot carry an overlap value. Scoring it zero would claim the engine swapped its sources, when what it did was stop showing sources at all. Those are different behaviours and the dataset keeps them apart.
The formal Arabic version of one question was not cited at all
The register arm has three usable pairs rather than four, and the missing one is the interesting one.
Asked in Gulf colloquial, “which lab should I go to in Riyadh” returned five citations, every one of them a Saudi lab.
Asked in Modern Standard Arabic, the same question returned a complete answer with zero citations of any kind. No source list, no external links anywhere in the page. It recommended two named labs and drew a map instead.
We re-checked after a further ten seconds and the answer was closed and final. We are not going to build a theory on one capture, and we are flagging it because a citation study that quietly dropped it would be hiding its most awkward observation.
Does prompt phrasing change who gets cited?
Not measurably. Across the whole Arabic set, keyword-like prompts returned 73.6% commercially owned citations on an interval of 60.4 to 83.6, against 68.6% for conversational prompts.
The intervals overlap across their whole middle. Within the two controlled stems the direction is not even consistent, because keyword phrasing lowered the commercial share on the H pylori question and raised it on the diabetes one.
Phrasing changes which sources appear. It does not appear to change what kind of source they are, and we are reporting that half as a null.
Google cited its own machine translation
One source in the Arabic vitamin D answer was my-clevelandclinic-org.translate.goog. That is Cleveland Clinic’s English page served through Google’s own Translate proxy, cited inside an Arabic answer, and the same answer separately cited the untranslated page.
It happened again on a different question. The Arabic answer for a raised RDW result opened with the same Cleveland Clinic translation proxy at position one and cited a second proxy, ubiehealth-com.translate.goog, further down.
When the Arabic corpus holds no authority for a clinical question, the system does not drop to a weaker Arabic page. It translates an English authority on the fly and cites the translation.
If that behaviour is common rather than incidental, a great deal of Arabic GEO advice is aimed at the wrong target.
It cited a corn biotechnology company as a laboratory in Mecca
Asked in English for the best medical laboratory in Mecca, the engine returned four citations. One of them, greenlab.com, is GreenLab, Inc., a United States company whose own homepage describes it as greening industry with next-generation plant biotechnology by growing proteins in corn.
It is not a medical laboratory, it does not operate in Saudi Arabia, and nothing on the page is about testing blood.
There is a Green Lab clinic in Mecca and there is a greenlab.com.sa, so the likely mechanism is a name match that landed on the wrong registrable domain.
We are reporting it because it is the kind of error a reader cannot see. The citation is presented in the answer exactly as the three correct Saudi laboratories beside it are.
It is one citation in 337 on the English side, so it moves nothing.
What it illustrates is worth more than its weight: an entity that is not disambiguated in the engine’s own view of the world can be attached to a question it has no relationship with, and in English, on a Saudi question, the pool of candidate entities is thin enough for that to happen.
We coded it brand-owned and US, because our taxonomy codes what a page is rather than whether the engine was right to cite it. That adds one to the English commercial numerator. Recode it as an error and the English commercial share moves from 58.2% to 58.0%.
The same hospital, counted twice, once per language
dsfhjeddah.fakeeh.care and en.dsfhjeddah.fakeeh.care are one website. Google serves the en. host to an English prompt and the bare host to an Arabic one, and strict domain matching scores those as two different sources. Three domain families in this dataset split that way, covering 18 of 726 resolved citations, and all three straddle both languages. The Riyadh Fakeeh hospital joined the Jeddah one in this pattern in cycle 14, its bare host cited in Arabic and its en. host in English inside the same question.
This pushes in the same direction as the translation proxy above. It makes the two languages look further apart than the documents behind them actually are.
We measured it rather than arguing about it. Normalising en., ar. and www. prefixes moves the headline from 0.149 to 0.152, and moves exactly one pair, the PCR price question, from 0.273 to 0.400.
The strict figure stays as the headline and the normalised one is published beside it, on the same principle as the translation proxy: name the bias, and keep the number that does not flatter us.
The English question was not always the same question
On the accreditation stem the two languages did not stay on the same topic. That is a translation artefact rather than a citation difference, and we would rather name it than bury it.
The Arabic prompt returned the Saudi Accreditation Center three times, the Ministry of Health, and a Riyadh clinic.
The English prompt returned the accreditation body twice more, then drifted into industrial testing: a soil laboratory, a geotechnical firm, Saudi Electricity’s qualified-supplier list and a Spanish inspection company.
“Lab” carries the medical sense in the Arabic phrasing and loses it in the English one. That pair scores 0.067 overlap, and part of that number is the word, not the web.
Is this Arabic, or is it Google?
Everything above is one product, and a single-engine study cannot tell you whether it has found a property of the language or a property of the vendor, which is the first question any competent reader will ask of it.
So we ran a second engine. ChatGPT, logged out, no account, no history, the same conditions and the same Saudi market token in the query text, on 10 September across nine rounds: 85 citation-bearing captures, 303 citations, 37 matched pairs, held in data/citations-chatgpt.csv and data/citations-chatgpt-round2.csv and kept out of every pooled figure on this page.
29 stems have now been put to both engines, in both languages, in the same week, nine above the twenty this comparison was designed to need. This is not a replication test on two differently built panels. It is one controlled comparison: the same questions, the same register, the same week, one run per prompt on each engine, and the two arms share a stem set rather than a subject. Every figure below is computed by analyze_chatgpt.py from the two engines’ own rows.
| 29 shared stems | Google AI Mode | ChatGPT |
|---|---|---|
| Mean cross-language overlap | 0.119 | 0.351 |
| Pairs sharing nothing | 11 of 29 | 10 of 29 |
| Institution share, Arabic | 11.2% of 258 | 60.8% of 102 |
| Institution share, English | 29.3% of 273 | 64.6% of 130 |
| Institution gap | 18.1 points | 3.8 points |
Two results hold. Both engines cite institutions less often when the question arrives in Arabic than when the same question arrives in English.
And on the same questions in the same week, ChatGPT is far more institutional in Arabic than Google is: 60.8% against 11.2%, a difference of 49.6 points, which is still the largest effect anywhere in this study.
That difference has now shrunk twice running, 57.4 points, then 54.8, then 49.6. We withdrew the claim that it only ever widens when it first stopped widening, and we are not replacing it with a claim that it only ever narrows.
The two gaps are 14.3 points apart, the seventh different value in seven rounds. That difference has been 15.8, then 10.1, then 14.1, then 4.0, then 9.2, then 11.7, then 14.3, on six, ten, thirteen, seventeen, twenty-one, twenty-five and twenty-nine shared stems. A quantity that keeps moving with the sample is not a property of either engine, which is why the split below, rather than this figure, is what we report as the result: the single gap number is a weighted average of two opposite patterns, so its value tracks the mix of question types in the panel.
Before attributing this round’s movement to the four new questions we checked the obvious alternative, which is that the Google panel simply grew underneath the comparison.
It did not: recomputing Google’s twenty-five previous stems against today’s data returns 11.9% Arabic, 30.6% English and an 18.7 point gap, the same three numbers as last round. That is three clean rounds in a row, and the whole of the move belongs to the new questions.
| Institution share, shared stems | Google, Arabic | Google, English | ChatGPT, Arabic | ChatGPT, English |
|---|---|---|---|---|
| 10 informational questions | 17.9% of 95 | 54.2% of 96 | 97.5% of 40 | 100.0% of 46 |
| 19 commercial questions | 7.4% of 163 | 15.8% of 177 | 37.1% of 62 | 45.2% of 84 |
Read the rows rather than the totals. On Google the Arabic penalty sits mostly inside the informational questions, 36.3 points there against 8.4 on the commercial ones, because a commercial question gets a commercial answer in either language and there is little institutional share left to lose: 7.4% and 15.8%.
On ChatGPT it is the other way round. Informational questions are answered institutionally in both languages, 97.5% and 100.0%, so almost nothing is lost there, and its Arabic penalty falls on the commercial questions, 37.1% against 45.2%.
No informational question was added this round, and all four cells in that row are unchanged to the decimal. Every cell in the commercial row moved. That is the third time the split has behaved exactly as a mix effect should under a deliberate change of mix.
What this round’s four new questions did. All four ask a reader to pick a provider, and three of them ask for a named person rather than a building: the best obstetrics and gynaecology hospital in Jeddah, the best dermatologist in Riyadh, the best dentist in Jeddah and what a cleaning costs there, and the best knee specialist in Riyadh. ChatGPT answered all eight captures without citing a single institution, in either language. Twenty-five citations: hospital and clinic marketing pages, two individual doctors’ own websites, two booking directories and two insurance provider network listings. Google did nearly the same on the same four questions, zero institutional citations in Arabic and one out of fifteen in English.
Set that beside the round four questions earlier, where every one of the eight captures came back one hundred per cent institutional and thirty-nine citations arrived without a single commercial page.
Two rounds, both eight captures, one at the ceiling and one at the floor, on the same engine in the same week. Together they are the clearest evidence in this study that the institution share measures the question and not the engine.
This round also exposes a limit in the gap measure itself. ChatGPT’s gap on these four questions is 0.0 points, and that zero means there was nothing to lose: with no institutional share in either language, the gap has nowhere to move. A gap of zero can mean the two languages were served equally well, or that both were served entirely by people selling something, and only the level tells you which. That is why the level sits in the table above the gap, and why a reader should not take a narrowing gap as good news on its own.
The previous round’s four questions, for continuity, asked which hospital in Jeddah, which hospital in Dammam, which medical centre in Jeddah does a full checkup, and what a ferritin test measures that an iron test does not. Three of those four came back with the gap running backwards on ChatGPT, its Arabic answer more institutional than its English one, which no round before it had produced.
The sharpest single result is still on the Google side. Asked in Arabic how Saudi laboratories differ in accuracy, Google returned ten citations and not one of them was an institution. The same question in English returned seven institutions out of fourteen. That is a fifty point Arabic penalty on a single question, the largest in the study. ChatGPT, asked that question in Arabic, cited the national accreditation body and the national healthcare accreditation authority and nothing else.
An earlier round broke the tidy version of our own explanation, which is worth more than another confirmation of it. The accuracy and accreditation questions carry the panel’s provider label, so they counted as commercial, and they did not behave commercially on either engine. Both have a canonical institutional answer, the regulator. What predicts the Arabic penalty is not whether a question is commercial but whether it has a canonical answer, and the class label is only a proxy for that. This round is the same point from the other end: four questions with no canonical answer produced no institutional citations at all.
ChatGPT’s institutional lead in Arabic is real and large on questions with a canonical reference answer, where it reaches for MedlinePlus, the NIDDK, the CDC, the NHS, Cleveland Clinic, the American Thyroid Association, the WHO and the Saudi Accreditation Center in both languages. On the question of where to actually go and who to see, it has no such advantage.
The purchase set is now seventeen questions and reads 23.5% institutional in Arabic against 37.0% in English, where the ten purchase questions fixed two rounds ago read 30.3% and 59.5%. Both cells are falling as more provider questions enter, and the English one is falling faster.
The mechanism behind the informational half is visible in the rows. For the Arabic diabetes question ChatGPT cited the CDC and the NHS, pages that exist only in English. For the Arabic H pylori question it cited Mayo Clinic’s Arabic edition. For the Arabic RDW and TSH questions it cited MedlinePlus and Cleveland Clinic, English pages both.
Google, asked those same questions in Arabic, reached for Arabic commercial pages, and on RDW it reached for Cleveland Clinic served through Google’s own translation proxy. One engine answers an Arabic question by translating a foreign institution, the other by finding a local business.
Both languages now name businesses without sourcing them, and this round they did it in every capture. Asked in Arabic for the best maternity hospital in Jeddah, ChatGPT named four hospitals and cited two. Asked in English for the best dermatologist in Riyadh it recommended one doctor with a citation, then listed four clinics with star ratings and review counts and cited nothing for any of them. The Arabic dentist answer named four clinics with ratings and cited only the two price pages. Earlier rounds recorded this as an Arabic behaviour, then found it in English once the English answers were read for it. On a round of pure provider questions it is in all eight.
No non-production URL appeared, for the third round running. Twenty-five citations, every one a production host. The five in the corpus, a testing environment at the College of American Pathologists, a testing environment belonging to Dr. Sulaiman Al Habib Medical Group, a specimen certificate for a placeholder laboratory, a client demo subdomain cited above the live site it copies, and a hyphenated lookalike of medlineplus.gov, all arrived on questions about price or accuracy, where the engine had to reach past the institutions to find somebody selling. These four questions are entirely about somebody selling and produced none, so that reading is now weaker than it looked.
One retrieval defect did recur, and recurring is what makes it worth reporting. ChatGPT again attached the attribution label “Dr. Sulaiman Al Habib Hospital” to a URL belonging to Dr. Soliman Fakeeh Hospital, a different and competing Saudi hospital group. Same host, same wrong label, second round in a row, this time in Arabic where the previous instance was in English. The destination is correct and the publisher name printed beside it is a competitor’s, which is a failure a reader cannot catch by checking the link.
Two caveats, and neither is small. 29 stems, with 102 Arabic and 130 English ChatGPT citations under the percentages, so a handful of rows moves any of these numbers several points, and the class split divides those already small counts in two.
And one run per prompt on the ChatGPT side, with no repeat arm underneath it, so its 0.351 has no noise floor: the Google panel’s repeat arm puts same-question, next-day overlap at 0.602 in Arabic.
The overlap figure behaves differently from the institution figure. ChatGPT’s mean is still close to three times Google’s on the shared stems, 0.351 against 0.119.
The disjoint rates, level last round at 8 of 25 each, have separated again: 10 of 29 on ChatGPT against 11 of 29 on Google. That is the sixth reversal in this quantity in six rounds, so it is a mix effect too, not a trend.
The two questions this round where the languages did meet met without an institution in the room. The Jeddah maternity question returned two of three domains in both languages, all of them hospital marketing pages, against 0.200 on Google. The Jeddah dentist question returned two of five, both of them clinic price guides, against 0.000 on Google. That is the same shape as the Jeddah medical centre question a round earlier, which hit 1.000 with nothing institutional in it. The dermatology and orthopaedics questions shared nothing on either engine.
Look at which questions produce agreement. Dental and internist resolve to booking directories, which are bilingual platforms serving one database to both languages. The Jeddah laboratory and accreditation questions resolve to accreditation bodies that publish in both.
Diabetes resolves to two national health services with one canonical page each, ferritin to the WHO and the AAFP, and the thyroid question to the NHS and the American Thyroid Association. The maternity, dentist and medical centre questions resolve to the same handful of private hospitals, which own their own answer in both languages as completely as the CDC owns diabetes.
So cross-language agreement tracks whether a single source answers the question for everybody. It does not track whether the engine is good at Arabic, and it does not require that source to be an institution. It requires it to be the only place the answer lives.
For completeness, the two full panels, which remain differently weighted and are context rather than evidence:
| Whole panels | Google AI Mode | ChatGPT |
|---|---|---|
| Matched pairs | 40 | 37 |
| Citations | 753 | 303 |
| Mean cross-language overlap | 0.149 | 0.342 |
| Pairs sharing nothing | 14 of 40, 35.0% | 14 of 37, 37.8% |
| Institution share, Arabic | 9.1% | 47.0% |
| Institution share, English | 24.0% | 61.2% |
| Institution gap | 14.9 points | 14.2 points |
One more difference is worth stating, because it changes what you would build. Brand-owned provider pages take 33.8% of ChatGPT’s Arabic citations across the whole arm against 32.2% of its English ones.
The rest of the Arabic column goes to directories and booking platforms at 10.6%, insurance provider networks at 4.0%, agency blogs and content farms at 2.6%, one social post and one lookalike domain. The English side has directories at 3.3%, insurance at 2.0% and agency blogs and content farms at 1.3%.
That brand-owned column has now closed. It read 30.2% Arabic against 20.0% English two rounds ago, 31.0% against 25.7% last round, and 33.8% against 32.2% now: 1.6 points apart, inside anything this sample could call a difference. It closed because the English figure rose twelve points in two rounds while the Arabic one rose three. On this engine, the share of citations going to somebody selling the thing being asked about is no longer an Arabic-specific finding. It is a finding about provider questions, which are the questions we have spent the last three rounds adding.
Google’s Arabic gap fills with commercial provider pages while ChatGPT’s fills with intermediaries. Same hole, different tenants, and on ChatGPT the English side is now filling too.
An English prompt does not reliably get an English answer
This is the most awkward thing we found today and it belongs above the limitations, not inside them.
We noticed one English capture answering in Arabic, so we re-asked seven English prompts and measured the script of each answer. Four of the seven came back in Arabic.
| Stem, English arm | Answer language |
|---|---|
| Labs in Dammam | Arabic |
| CBC test | Arabic |
| Vitamin D price | Arabic |
| H pylori symptoms | Arabic |
| Premarital screening | English |
| Ferritin against iron | English |
| Best lab in Riyadh | English |
The Arabic ones are not borderline. The CBC answer opens “اختبار CBC هو اختصار لـ Complete Blood Count”.
What this does not break. The language column records the language the question was asked in, which is what the data dictionary has always said it records. Every number on this page is a prompt-language number and remains one.
What it does break. Any sentence on this page of the form “the English answer returned the NHS and three US clinics” assumes the English arm answered in English. Read those as “the English prompt returned” instead.
The obvious guess about the bias is wrong. You would expect an Arabic answer to an English prompt to pull Arabic sources and inflate the overlap figure. It does not. H pylori and vitamin D price both answered in Arabic in this check, and both score 0.000 against their Arabic twins. Labs in Dammam answered in Arabic and scores 0.375, our highest provider pair. Premarital answered in English and scores 0.357.
Answer language and source overlap do not line up in either direction. The reading that survives is that prompt language moves the retrieved sources independently of the language the answer is written in, which is a stronger claim than we are currently entitled to make.
We ran this check a day after the panel, and this engine changes overnight, so treat four of seven as evidence the behaviour is common rather than as a rate for 9 September.
We then found the cause, and it is ours. Adding &hl=en to the search URL, which sets Google’s interface language to English and changes nothing else, flipped all three of the tested Arabic answers to English. Two of the three came back with zero Arabic characters on the entire page.
The collection browser was configured with an Arabic Google interface and nobody set hl explicitly. So the English arm spent this whole study asking English questions of an Arabic-configured product, and about half the time it got an Arabic answer back.
That is a defect in our instrument, not a property of the engine, and we would rather say so than let the reader assume Google decides Saudi questions deserve Arabic answers.
The fix is one URL parameter on every future capture, hl=ar for the Arabic arm and hl=en for the English one, and it is specified with the raw counts in data/ANSWER-LANGUAGE-CHECK.md.
We then measured what it cost, across the whole panel rather than a sample. It is small, and its direction is still not settled. We re-captured the English arm under hl=en for 37 of the 40 paired stems and recomputed overlap against the unchanged Arabic arm.
| Cut | Stems scored | J as collected | J under hl=en | Change |
|---|---|---|---|---|
| First reported | 4 | 0.112 | 0.139 | +0.027 |
| Full arm | 34 | 0.152 | 0.116 | -0.037 |
Thirty-four of the 37 returned citations and can be scored. Three returned a complete answer with no citation surface at all and are logged separately. The last three paired stems have no English prompt in any of our stem banks, so we never captured them rather than invent a prompt and pretend the pair was matched.
The magnitude held small from four stems to 34. The correction moves the mean by 0.037 against a panel figure of 0.149 and a gap to the noise floor of roughly 0.45. That is the finding. Whatever the Arabic interface was doing to the English arm, it did not manufacture the cross-language gap, and correcting it does not close it.
The direction is a different matter, and 34 stems did not settle it. The correction moves 14 stems down, 8 up, and leaves 12 unchanged. The aggregate has leaned negative since the arm passed roughly 20 stems, but the sign flipped once already at four, and the per-stem spread is wide: one stem falls 0.667 while another rises 0.286. Read the magnitude. We are not quoting a direction off this.
The two zero-overlap stems are the result that matters most. Both answered in Arabic under the default interface, and both still shared nothing at all with their Arabic twin once corrected. An English prompt answered in Arabic still retrieves a different source set from the same question asked in Arabic.
Three reasons we are not closing this. The corrected captures returned fewer citations, a median of four against six, and a smaller set shrinks the union and pushes Jaccard up, so part of the stem-level movement in both directions is that artefact rather than the interface. The arm ran across several weeks against an engine our repeat arm shows changing overnight, so elapsed time is confounded with the fix on every stem. And three pairs are missing.
One thing the arm did settle: the fix works. Not one of the 37 corrected captures came back with an Arabic answer body, against a wave 1 English arm where roughly half did. The Arabic characters that remain sit in embedded business listings and clinic names the answer surfaced, never in the answer prose. One URL parameter took the defect to zero and held it there for 37 captures.
Wave 2 then hit the same defect from the other side, and it stopped collection for one run. The first wave-2 run on 10 September found that the collection browser rewrites every Google URL before it loads: any Google domain is redirected to google.com.sa, and gl=sa, a Riyadh location token and hl=en are forced onto the query string. A requested hl=ar is silently replaced with hl=en on every attempt, on every Google ccTLD, with and without AI Mode. So the interface language was pinned by the browser rather than chosen by the protocol, and it was pinned to English while wave 1 had run with it pinned to Arabic. We stopped rather than mix two instrument conditions inside the column that carries most of this page’s findings.
That turned out to be a property of one browser rather than of Google. Collection resumed the same day in a second browser that rewrites nothing: hl=ar survives, the host stays www.google.com, and the Arabic answer comes back in Arabic.
Every capture from that point declares its own hl, gl=sa and the same Riyadh location token explicitly, so the interface condition is now stated by the protocol instead of inherited from the tool, and it is verified on each capture rather than assumed. Wave 1 remains under the old, inherited condition, and the 37-stem re-capture arm above is the only measurement of what that condition cost.
One more thing that check turned up, and it corrects our own method note. The appendix has described the collection browser as exiting in France, which is true of the network path and was verified twice.
It is not the whole picture: every request also carries a Riyadh location token and gl=sa, so the search itself is geolocated to Riyadh. That is the right geography for this study, and it should have been written down from the start.
What this does not prove
We ran one prompt per capture in the panel. These systems are not deterministic, and the repeat arm now puts a number on that: 0.602 mean overlap on eight Arabic pairs and 0.545 on four English ones, one day apart, nothing changed.
That is enough to retire the register arm as a null. It is not enough to state a volatility rate for the engine, because the individual pairs run from 0.231 to 0.889 and eight pairs cannot tell you what drives that spread.
Every repeat pair is a single 24-hour interval measured on one date. It is not a decay curve, and nothing here says whether the same questions would look this different a week apart or an hour apart.
The register arm now has a second problem on top of its three pairs. Its mean sits inside the repeat arm’s interval, so we cannot separate the register effect from noise even in principle at this sample size.
Reading it as a small effect rather than as a null would be reading the point estimate and ignoring the band around it.
One stem was also captured twice roughly 20 minutes apart during a collection incident, and the two Arabic answers to the identical prompt were not identical. That pair is described in data/REPEAT-CAPTURES.md and carries no figure, because one of its two captures lost citations to rate limiting.
The class table above rests on five or six pairs per class, and the institutional row rests on a single pair. Several intervals are wide enough that a different sample could reorder the middle of that table, though the two ends have held across three rounds.
Delta Medical Laboratories is a VOCTOS client and we produced the rankings this study tests. Twenty of the 37 Arabic panel captures did not cite them, and those twenty sit in the dataset and in the paragraph above.
What to do with it
Publishing Arabic content is not the finding. Two sharper things are.
Classic rank buys you the AI answer for “which lab should I use”. It buys you much less for “what does this test mean” or “what does it cost”. That is where the volume sits.
One symptom stem in this set carries 40,500 monthly Saudi searches, and its Arabic answer was carried entirely by clinics outside the country.
The second is the concentration figure. An HHI of 195 across 152 domains means the Arabic slot is winnable by whoever publishes a coherent Arabic page, and losable the same way.
The top of that list holds no ministry page, no Saudi publisher and no encyclopaedia. It changes completely from one question to the next.
Questions people ask about this
Do AI search engines cite different sources in Arabic than in English?
Yes, and by more than most people expect. Across 40 matched pairs of the same Saudi healthcare question asked in both languages, mean domain overlap was 0.149, and 14 of those pairs shared no source at all. Ranking in one language tells you almost nothing about the other.
Does ranking first on Google get you cited in the AI answer?
Less than half the time in this dataset. The Saudi laboratory we manage held 355 of 628 tracked keywords at position one in August 2026, yet it appeared in only 45.5% of the 44 Arabic answers we captured. Read that as a site-level rate rather than a per-question one: we did not match each captured question back to a tracked keyword, so this is what a strong ranking position buys across a category, not for a specific query. Classic rank and AI citation are separate assets, and the second one is not bought by improving the first.
Why do Arabic AI answers cite commercial pages instead of hospitals and ministries?
Because in most cases no Arabic institutional page exists for the question being asked. Commercially owned pages carry 69.3% of Arabic citations against 58.2% in English, and institutions take 9.1% of the Arabic set against 24.0% of the English one. On price questions the effect disappears in both languages, because no ministry publishes what a test costs.
Is it easier to get cited in Arabic than in English?
On this evidence, yes. The Arabic citation pool is smaller and more concentrated, so a single well-built page moves the needle further than the same page would in English. That is a window rather than a permanent state.
How stable are AI citations day to day?
Less stable than a ranking. Asking the same question in the same words 24 hours later returns roughly 60% of the same domains, so about a third of the source set turns over overnight. Any single capture is a snapshot, and a claim built on one run is not a claim.
Is this a Google problem, or does ChatGPT do the same thing?
Both engines cite institutions less often when the question arrives in Arabic, so the direction is not specific to Google. The size of the effect does not carry across engines and has moved every time the shared question set has grown, so we report the direction and not a fixed number.
How was this measured?
Google AI Mode and ChatGPT, logged out, no conversation history, one run per prompt, on matched Arabic and English versions of the same questions. Every Google citation was resolved through its encrypted redirect to a real destination domain, because reading the page source alone does not tell you what is being cited. Unresolved citations stay in the denominator. The full protocol is in the method appendix.
Can I use this data?
Yes. The dataset, the analysis script and the method appendix are published under CC BY 4.0. Re-run the protocol on your own category and tell us where we are wrong.
Cite this study
Adel, Beshoy. What Google AI Mode cites when a Saudi patient asks in Arabic. VOCTOS Research, September 2026. https://www.voctos.com/research/arabic-vs-english-ai-citations/
Dataset released under CC BY 4.0. If you quote a figure from this page, quote the date with it. Collection is ongoing and the numbers move.
Credits and prior art
The mechanism we describe was assembled from other people’s work rather than ours.
Arabic sits at 0.5 to 0.6% of the web by three independent measures, including QCRI’s Fanar 2.0 paper, and “All Languages Matter” found that over 70% of top-five multilingual RAG retrievals come from English or from the query language alone.
Amiraz and colleagues measured 13 to 42% cross-lingual retrieval loss between Arabic and English at ArabicNLP 2025. Dai and colleagues showed at EMNLP 2025 that citation tracks recognition of a source’s identity rather than assessment of its content, and that recognition does not transfer across languages.
Our contribution is the measurement itself, taken inside a live product, in Arabic, on dated questions, with every raw row attached so that anyone can recount it.
References
Prior work this study builds on
- Fanar 2.0, Qatar Computing Research Institute. Arabic at roughly 0.5 to 0.6% of web content.
- All Languages Matter: multilingual retrieval-augmented generation. Over 70% of top-five multilingual RAG retrievals come from English or the query language alone.
- Amiraz and colleagues, ArabicNLP 2025. Cross-lingual retrieval loss between Arabic and English of 13 to 42%.
- Dai and colleagues, EMNLP 2025. Citation tracks recognition of a source’s identity, not assessment of its content, and recognition does not transfer across languages.
Industry studies that do not cover this ground
- Profound. 3.25 billion citations, 14 countries, platform mix rather than source language.
- Temso. Seven languages, Arabic excluded.
- Weglot. Four languages, none of them Arabic.
This study
data/citations-master.csv. One row per citation event, released CC BY 4.0.data/citations-chatgpt.csvand-round2.csv. The ChatGPT arm, kept separate so no Google figure is averaged with it.METHOD-APPENDIX.md. The protocol, the two data-loss failure modes with their counts, and the single-coder limitation.
Next release
The register and repeat arms are both in and reported above. The repeat arm ran on 10 September across thirteen captures in two languages and returned twelve usable pairs, which was enough to turn the register result into a null.
What it could not do is explain why the pairs range from 0.231 to 0.889. That needs more stems per class, and a second interval, so we can tell a question-level property from a time-level one.
The second engine is open, and getting there took one correction we would rather publish than bury. Perplexity refuses to answer without an account and Brave Search hides its citation destinations, and from those two failures we wrote that no anonymous engine exposes readable destinations. That claim shipped in an earlier version of the appendix.
It was false. ChatGPT answers logged out, in Arabic, with citations, and carries every destination in plain text in a JSON attribute on the citation button. It is easier to read than Google. Nobody had tested it.
What wave 2 needed was not access. It was the same twenty stems on both engines in the same week, so the comparison above would stop being two panels and become one controlled test. Twenty-five stems are now shared and that comparison is what the section above reports.
What it still needs is a repeat arm on the ChatGPT side. Every ChatGPT figure on this page is one run per prompt, so none of them has a noise floor underneath it. That is specified in .autonomous/arabic-citation-study/deferred.md and is not captured yet.
If your category is not healthcare, none of this may transfer, and we would rather you tested it than assumed it. The dataset, the analysis script and the method appendix all sit in this directory.
If you want this run against your own category and market, that is what our GEO service does. If you would rather check our working first, the raw rows are linked above and we would genuinely like to be corrected.
Editor’s note, updated 10 September 2026. Two items closed. Beshoy Adel’s credential and sameAs are now taken from the voctos.com team page and fixed in schema.json.
The competitor comparison is pulled from Semrush on the Saudi database, dated 10 September: Delta at rank 348 with 37,224 organic keywords and 35,266 AI Overview keywords, Wareed at 778, Alfa at 1,239, Al Borg at 613.
The Delta tracking figures were re-pulled on 11 September from the client ranking file, and now read 355 of 628 at position one and 453 in the top three, dated August 2026. The earlier July pair, 197 of 326 and 240 in the top three, could not be reproduced from that file and has been replaced rather than carried forward.
One item stays open. The tracking file has no September column, and the Search Console property we hold for this client covers only the blog subfolder, so the closest rank date to the 9 and 10 September captures is still two weeks earlier. The captured questions are conversational and carry no recorded source keyword, so a per-question join between rank and citation is not possible on this dataset. Both are fixable with a September export and a keyword column in the stem bank, and both are needed before the carry rate can be read as a per-query number.
Take the study with you
A designed PDF of the findings, the charts and the method. Leave your details and it downloads straight away.
We use these to send you the next wave and nothing else. Or download without leaving details.
Don't miss the chance to
make your website more visible!
Initial consultation and
audit of the current situation
Read also
Real Results
More WinsTrusted by









































































