Home BusinessCommentLost in translation

Lost in translation

Who gets to represent Lebanon in AI searches?

by Mohamed Soufan

When a user asks an AI system about Lebanon, the answer may be markedly different simply on basis of which language was used. Behind it is another decision that receives much less attention, namely, which sources the system chooses to surface, and whose information is consequently placed in front of the user.

That decision is becoming more consequential as AI becomes another route to information. A June 2026 Pew Research Center report on the impacts of AI use on Americans found that 42 percent of U.S. adults surveyed use AI chatbots to search for information.

For Lebanon, this raises a particular question. If someone asks the same question about the country in English and Arabic, will the AI system draw on the same institutions, publications and public sources? Or can changing the language also change who gets to represent Lebanon? My research led to the finding that there are, at the time of this writing, two quite different sets of information about Lebanon, depending on where you look.

 

Different language, different sources

In an August 2026 study, Language-Dependent Source Representation in AI Search: A Cross-Platform Audit of Lebanon, I tested whether language changes the sources AI search systems cite when answering questions about Lebanon. I submitted 40 questions to ChatGPT Search, Google Gemini Search, Perplexity and Microsoft Copilot Search. Each question was asked in both English and Arabic, with the two versions designed to ask the same thing rather than simply mirror one another word for word. The questions covered eight areas, from banking and infrastructure to health, education, tourism and public institutions. In total, the audit produced 320 responses.

The difference was substantial. Among responses with classified citations, the average Lebanese-source share was 31 percent in English, compared with 57 percent in Arabic. In other words, changing the language of an otherwise equivalent question was associated with a markedly different level of Lebanese representation in the sources AI systems chose to cite.

Because the same questions were repeated across four platforms, the study also compared English and Arabic at the question level rather than simply pooling every response together. On that measure, Arabic queries produced a 23 percentage-point higher share of Lebanese sources on average. The difference was statistically significant and appeared in the same direction on all four platforms.

The finding was not confined to one AI system or one topic area; it appeared across all four platforms and across all eight subject areas tested. What matters next is not simply how much representation changed, but which sources gained or lost visibility when the language changed.

Which sources become visible?

The difference was not only in how often Lebanese sources appeared, but in which kinds of sources AI systems turned to.

In English, international and multilateral organizations accounted for 21 percent of classified citations, compared with 12 percent in Arabic. Wikipedia and World Bank properties were particularly prominent: English Wikipedia appeared in 42 of the 160 English responses, while World Bank properties generated 72 appearances across the 160 English responses. Other frequently surfaced sources included the ILO, PMC/NCBI, ScienceDirect and Britannica.

Arabic queries produced a different mix. Government and public institutions accounted for 25 percent of classified citations, compared with 16 percent in English, while news and media sources more than doubled their share, from 10 percent in English to 24 percent in Arabic.

Looking at individual Lebanese sources makes the shift more concrete. In English, frequently cited sources included Banque du Liban, L’Orient Today, the Investment Development Authority of Lebanon and the Lebanese Center for Policy Studies. In Arabic, the list shifted toward institutions such as the Lebanese Army and Ministry of Economy, alongside Lebanese business-news outlets. An-Nahar appeared among the ten most frequently cited Lebanese sources in Arabic, but not among the ten most cited in English.

The result is two noticeably different information environments around Lebanon. Perhaps running this study in French would reveal a third environment. In this audit, English responses drew more heavily on international institutions and general reference material, while Arabic responses drew more heavily on Lebanese public institutions and domestic media. The question, then, is not simply whether AI can find information about Lebanon. It is which institutions are placed in the position of explaining Lebanon to the user.

Why the visibility gap matters

For Lebanese institutions, being online is no longer the same as being visible. As AI systems become another route to information, the sources they choose to surface can shape which institutions reach different audiences and which remain less visible. The issue is no longer only whether reliable information exists online, but whether AI systems surface it in the answers users receive.

This matters especially for Lebanese media and public bodies. A newsroom may publish substantial reporting about Lebanon and appear regularly when users ask questions in Arabic, yet have much less presence when similar questions are asked in English. The same risk applies to ministries, universities and research organizations. Their information may exist and be publicly accessible, but that does not guarantee that an AI system will surface it for audiences searching in another language. Of course, there is the constant caveat that AI developments are constantly in flux, and what is true one week may be invalidated the next. But what the study points to is a visibility problem that conventional measures of web presence do not fully capture.

In a September 2026 Executive analysis on AI readiness in Lebanon, Jamile Youssef raised a related problem. The article noted that government information remains fragmented and much of it has not been consistently digitized, while archives of Arabic knowledge and culture that are inaccessible digitally risk remaining outside the source material available to AI systems. My findings suggest an additional complication: digitizing information is necessary, but it may not be sufficient. Even among information that AI systems can already retrieve and cite, visibility can change substantially with the language of the query.

That has consequences beyond journalism. A Lebanese researcher trying to make local evidence discoverable internationally, a ministry publishing public information, or a business producing data about its sector may all reach different audiences through AI depending on the languages in which their material can be found and surfaced. For institutions thinking about digital visibility, the task is therefore becoming multilingual: it is not only a question of putting reliable information online, but of understanding whether that information remains visible when people ask about Lebanon in Arabic, English or eventually other languages.

What single-language AI audits can miss

The study does not establish why these differences occur. Several explanations are plausible. Lebanese material may be more abundant, better indexed or more readily matched to Arabic-language queries; AI systems may interpret Arabic queries as more locally oriented; or their retrieval systems may simply rank different kinds of sources depending on language. The data show the pattern, but they do not identify the mechanism behind it.

A February 2025 study comparing AI search retrievals of political information across four languages published in the academic journal New Media & Society, similarly found substantial differences in source use and attribution when auditing Microsoft Copilot across five languages during Taiwan’s 2024 presidential election.

There is another reason for caution. This was a snapshot of four live systems on one day, with one response collected for each question, language and platform combination. The results therefore describe the systems as they behaved during that collection period, and the size of the gap may change over time. The study also classified the institutional origin of sources, not the language of the individual pages being cited.

Even with those limitations, the practical lesson remains. If an organization audits an AI system only in English, it may conclude that Lebanon is mainly being represented through international institutions and reference sources. In this study, conducting the same audit in Arabic revealed a substantially more domestic information environment. A single-language test therefore should not be assumed to describe how an AI system represents a multilingual country. For Lebanese institutions, that changes what an AI visibility strategy should look like. Media outlets, universities, public agencies and businesses should not ask only whether AI systems can find them. They should ask in which languages they can be found, for which kinds of questions, and alongside which competing sources. As AI becomes a more common gateway to information, multilingual visibility will increasingly shape whose knowledge is present when a country—or when any given subject—is explained to the world

You may also like