Multilingual AI detection: how cross-language AI text detection works

Summary:

  • Multilingual AI detection must distinguish human and machine writing across different languages, domains and generating models, and performance in one language does not automatically transfer to another.
  • Detectors use fine-tuned multilingual models, perplexity-based measures, token weighting and ensembles, but training coverage, tokenisation, translation and mixed authorship all affect what they can reliably detect.
  • Strong results are possible in defined evaluations, but a benchmark score is not a universal accuracy guarantee: false positives against second-language and translated writing mean results still need context and human judgement.

An AI detector may perform well on English essays yet struggle with Ukrainian academic writing, Arabic news articles or a document that moves between two languages. The difficulty is not simply recognising different words. Detectors must distinguish human and machine writing across different linguistic structures, writing conventions, subject areas and generation systems – often without representative training examples for every combination. Research benchmarks have repeatedly identified generalisation to unfamiliar languages, domains and generators as a central challenge.

Translation adds another layer. Consider a researcher who writes an article in Turkish and uses software to translate it into English. Compare that with someone who asks an AI system to compose an English article from scratch. Both final documents contain machine-produced wording, but their intellectual origins and writing processes are different. A useful detection system – and anyone interpreting its results – needs to distinguish those questions.

This is the territory of multilingual AI text detection, also discussed in research as cross-lingual AI-generated text detection, multilingual machine-generated text detection and LLM-generated text detection. Understanding the differences between these terms is the first step towards understanding what the technology can realistically establish.

What is multilingual AI detection?

Multilingual AI detection attempts to identify machine-generated writing in more than one language. Cross-lingual detection usually places a stronger emphasis on transferring detection capabilities between languages, including situations where the target language was absent from the detector’s task-specific training data. This distinction is reflected in benchmarks that separately examine multilingual performance and generalisation to unseen languages.

The terminology can be organised as follows:

Term What it emphasises
Multilingual AI text detection Detecting AI-generated writing across several languages.
Cross-lingual AI text detection Transferring detection knowledge between languages, particularly into languages not represented in detector training.
Machine-generated text detection, often abbreviated to MGT detection The broader classification task of distinguishing machine-produced and human-written text.
LLM-generated text detection Detection focused specifically on output from large language models.

These are overlapping research terms rather than separate product categories. For a general audience, multilingual AI text detection is a useful umbrella label; cross-lingual detection is more precise when the question concerns transfer between languages. The terminology used in major shared tasks includes both “multilingual” and “machine-generated text detection”. See, for example, SemEval-2024 Task 8.

“Multilingual” does not necessarily mean “trained for every language”

A detector could be trained on labelled English, Spanish and German examples and then evaluated on those same three languages. That is a multilingual system, but it does not demonstrate that the detector will work reliably on Urdu.

There is also an important ambiguity around zero-shot. A statistical detection method may be described as zero-shot because it does not require a separately trained human-versus-AI classifier. A supervised detector can nevertheless be evaluated for zero-shot cross-lingual transfer when it receives no labelled detection examples in the target language. The detector has been trained; it simply has not been trained for that particular language. PAWN, for example, distinguishes its trained detection network from zero-shot statistical approaches while evaluating transfer to held-out languages.

AI detection is not cross-language plagiarism detection

The two tasks can appear together in a writing-checking service, but they answer different questions.

Cross-language plagiarism checking looks for relationships between a document and potential source material in another language. AI detection instead evaluates evidence that wording was machine-generated. A text can therefore contain copied human writing without displaying an AI-generation signal, or contain AI-generated prose without matching an identifiable source.

Even within AI detection, there are several different tasks. Binary classification asks whether a document belongs to a human-written or machine-generated category. Source attribution attempts to identify the generating model. Localisation attempts to identify where machine-generated writing occurs within a document. SemEval-2024 Task 8 separated these problems into different subtasks rather than treating them as interchangeable capabilities.

Consequently, success at classifying whole documents does not automatically establish that a system can identify individual AI-written sentences or name the model responsible.

How multilingual AI detection works

The main architectural families include fine-tuned multilingual models, cross-lingual adaptation, combinations of models or features, and analysis of token-probability distributions. These approaches can overlap: a single detector may combine a multilingual encoder with statistical features and language-specific decision thresholds.

Statistical and stylistic analysis

One approach examines measurable characteristics of writing rather than searching for matching sources. These can include vocabulary patterns, repetition, grammatical features and other aspects of linguistic style.

Such signals can contribute to classification, but their presence does not establish a universal distinction between human and machine writing. A SemEval-2024 study combining linguistic-stylistic features with pretrained language models found that stylistic information had predictive value, although its systems remained below the organisers’ baselines. That is a useful distinction: a feature can contain information without being sufficient for reliable detection on its own.

A related family uses perplexity and other probability-based measurements. Perplexity expresses how predictable a text is to a particular language model. It is sometimes used on the assumption that machine-generated writing will be unusually predictable.

The difficulty is that predictability is not synonymous with AI authorship. It depends on the scoring model, the text and the evaluation setting. Work at the 2025 GenAI Content Detection workshop combined a fine-tuned XLM-RoBERTa classifier with perplexity and other features and highlighted perplexity’s limitations in multilingual and multi-generator detection. A single perplexity threshold should not be treated as a language-independent authorship test.

Fine-tuned multilingual language models

A common supervised approach starts with a pretrained multilingual model and adapts it using examples labelled as human-written or AI-generated.

The model converts text into numerical representations that encode contextual information. A classification component then learns which patterns distinguish the labelled categories. Systems such as XLM-RoBERTa provide a starting point for multilingual detection, but the training process still needs representative examples of the writing the detector will encounter.

Research at the 2025 GenAI Content Detection workshop examined RoBERTa and XLM-RoBERTa systems and found that practical training choices – including input length, training duration and class balance – affected results. Choosing a well-known multilingual model is therefore only part of building a detector.

Multilingual encoders have also featured prominently in language-focused evaluations. In the 2026 AbjadGenEval shared task, systems using models such as XLM-R and DeBERTa-v3 were important participants in detecting AI-generated Arabic and Urdu news articles.

Parameter-efficient adaptation and cross-lingual transfer

Updating every parameter of a large model can be expensive. Parameter-efficient fine-tuning adapts a model through a smaller set of trainable components rather than retraining its entire internal structure.

For detection, the attraction is practical: researchers can investigate adaptation to new languages or tasks while limiting training requirements. However, reduced training cost does not itself demonstrate improved reliability.

The KInIT system at SemEval-2024 combined parameter-efficient fine-tuning with language identification and other detection signals. It also investigated per-language classification thresholds, recognising that the same numerical decision boundary may not be appropriate across every language.

This separates two questions that are easily confused: whether the model produces useful evidence, and how much evidence should be required before a document is flagged.

Learning which tokens matter most

Simple probability-based methods often summarise a document using an average. More elaborate systems attempt to distinguish informative parts of the text from less useful ones.

The Perplexity Attention Weighted Network, or PAWN, proposed by Miralles-González and colleagues, combines next-token probability information with contextual representations. It learns how much weight to give different tokens instead of assuming that each contributes equally to the detection decision. See Not all tokens are created equal: Perplexity Attention Weighted Networks for AI-generated text detection.

The authors reported a mean macro F1 score of 81.46% in their nine-language cross-validation setting using a LLaMA-based backbone (macro F1 averages the F1 score equally across the human and machine classes; the metrics are explained further in the accuracy section below). The result supports the usefulness of learning how to aggregate probability signals, rather than relying on a single average.

However, PAWN is not a training-free detector: its smaller detection network is trained even though the underlying language-model backbone can remain frozen. Its reported performance also belongs to the paper’s particular evaluation, not to every possible language or writing context.

Feature fusion and ensembles

Other approaches combine information from several sources.

The MLDet system presented by Agrahari and Singh used language-specific embeddings and a fusion mechanism intended to build representations that transfer across languages. This approach tries to retain useful linguistic information while learning features that are not tied entirely to one language.

An ensemble instead combines predictions from several models. The LuxVeri system, for example, combined RemBERT, XLM-RoBERTa and multilingual BERT for its multilingual task, weighting their contributions using inverse perplexity. It reported a multilingual macro F1 score of 0.7513 in that shared-task evaluation.

The rationale is that different models may compensate for one another’s weaknesses. But agreement is not automatically independent corroboration. If several components rely on similar training data or signals, they may also share the same errors. An ensemble therefore needs evaluation as a complete system, rather than being assumed reliable because several models contribute to its answer.

Why performance varies between languages

Unequal data coverage

Cross-lingual transfer cannot be assumed simply because a model accepts text in many languages. Detection training needs relevant examples of both human and machine writing, and the mixture of languages used for training can affect performance.

The effect can be large. In the 2025 GenAI Content Detection shared task, the Nota AI team reported that a multilingual classifier fine-tuned on all of the training languages reached a macro F1 of about 0.98 on non-English development data, whereas the same approach trained only on English samples fell to about 0.80 on the same data – evidence of some cross-lingual transfer, but also of a substantial loss.

The CEAID benchmark examined these issues for Central European languages, including different training-language combinations and evaluation across domains and generators. Its authors found that supervised detectors fine-tuned in the relevant Central European languages performed best in their evaluation and were the most resistant to the tested obfuscations.

The practical implication is more specific than “non-English detection does not work”. It is that direct evidence for the target language is more informative than an assumption that English performance will transfer.

A claim of support for a language should therefore be separated from evidence of validation in that language.

Tokenisation changes what the model sees

Language models typically process tokens: units that may be whole words, parts of words or other character sequences.

Equivalent information can require different numbers of tokens in different languages. Ahia and colleagues examined this issue across 22 languages and found substantial inequalities in tokenisation, with consequences for cost and the amount of information that can fit within a fixed token allowance. Their research concerns language-model tokenisation generally, rather than directly measuring AI-detector errors. Detection researchers have nevertheless identified the same mechanism as a concern: the PAWN authors list tokenisation among the challenges of multilingual detection, noting that backbone language models require many more tokens to encode text in less-represented languages, especially those written in other scripts.

For detection, an important inference follows. A detector with a fixed token limit may examine different amounts of actual content in different languages. A passage that fits comfortably within the limit in one language may be truncated in another.

This does not prove that tokenisation alone causes detection failures. It does mean that a comparison should report preprocessing and input limits rather than assuming that every language receives an equivalent examination.

Morphology, spelling and scripts create different test conditions

Research on multilingual detection frequently identifies word formation, diacritics, orthographic variation and script-specific perturbations as potential sources of difficulty. These are useful aspects to investigate, but they should not become shortcuts for explaining every difference in performance.

For example, a lower score on one language does not, by itself, establish that its grammar caused the difference. The datasets may also differ in topic, length, generation model, editing history or the quality of their labels. The DetectRL-X evaluation, discussed further below, illustrates the point: when its detectors were tested on languages included in training, average performance varied by only a few percentage points between languages its authors grouped as low, medium and high morphological complexity, and they reported no clear correlation between linguistic complexity and detection performance. In cross-lingual transfer, however, degradation did grow with complexity and with distance between language families. The same linguistic property can matter in one evaluation setting and hardly at all in another.

A stronger evaluation would compare these factors systematically. It could test standardised and non-standardised spelling, native-script and transliterated writing, or documents containing more than one language.

Consider a hypothetical Arabic paragraph containing English technical terminology. Treating the entire passage as an uncomplicated example of either language may conceal important differences between it and the detector’s training material. Code-switching should be an explicit test condition, not something assumed to be covered by separate Arabic and English accuracy figures.

Writing domain and text length matter

Language is only one source of variation. An academic abstract, a customer review and an informal social-media post are different detection settings.

The M4 benchmark examined multiple generators, domains and languages. Its findings included poor generalisation to unseen domains and generating models, often with machine-generated text incorrectly classified as human-written. A detector may therefore appear strong until the writing context changes.

Short and informal writing also requires its own evidence. The MultiSocial benchmark was developed around multilingual social-media content rather than assuming that results from longer, more formal documents would apply. Its dataset contains 472,097 texts across 22 languages and five social-media platforms. Its evaluation showed that fine-tuned detectors could work well in this setting, while the platforms selected for training influenced performance.

The conclusion is not that short text is automatically undetectable. It is that success on essays does not validate performance on a brief comment, and success on one social platform does not necessarily transfer to another.

New generators create moving targets

A detector may encounter writing from newer or unfamiliar generating systems. This is a form of distribution shift: the material being evaluated differs from the data on which the detector learned.

That is distinct from catastrophic forgetting, which can occur when updating a model causes it to lose previously learned capabilities. A detector performing badly on a new generator has not necessarily “forgotten” anything; it may never have learned that generator’s patterns.

The size of the effect varies between evaluations. In the DetectRL-X benchmark discussed below, unseen writing domains degraded the tested detectors more than unseen generators did, although both remained sources of error. Whether a new generator is a serious problem for a given detector is therefore an empirical question, not something that can be assumed in either direction.

Translation creates several different detection problems

“Can an AI detector recognise translated text?” is too broad a question. At least three situations need to be considered separately.

Human-written text translated by a machine

A person may produce the ideas, research, argument and original wording, then use machine translation to create another-language version.

The final wording has been produced through a machine process, but that does not establish that the original work was composed by an AI system. Indeed, machine-translation detection is itself a distinct research task. García-Romero and colleagues developed a method using a multilingual translation model to distinguish human translations from machine translations.

A detector trained to separate fully human-written documents from fully AI-composed documents may not be answering this more specific question.

This is also a reason to be cautious about translating an unsupported-language document into English solely to run an English-language AI detector. The additional translation step changes the object being assessed. Weber-Wulff and colleagues explicitly included machine-translated human writing as a separate category in their evaluation of detection tools, rather than treating it as equivalent to originally English human writing. The distinction mattered: across the fourteen tools they tested, overall accuracy was 96% for human-written English documents but fell by around 20% for human-written documents that had been machine-translated into English, and the authors reported that the risk of false positives increased dramatically for the machine-translated category. Their conclusion was that second-language writers who translate their own work are at particular risk of false accusation.

AI-generated text translated into another language

Here, the original content is machine-generated, but translation changes its surface form.

The ESPERANTO research examined back-translation and found that it could reduce the true-positive rate of the tested detectors. The evaluation covered nine detectors, including open-source and proprietary systems, and also investigated a method intended to improve robustness against this kind of transformation.

The broader lesson is about evidence preservation: retaining a text’s meaning does not necessarily preserve the statistical signals used to detect its origin.

It would nevertheless be wrong to turn this finding into a universal rule that translation always defeats detection. The outcome depends on the detector, translation process and evaluation conditions.

Human–AI collaboration and mixed authorship

Real writing workflows can involve outlining, drafting, translation, proofreading, expansion and substantial human revision. A single human-versus-AI label may conceal these different contributions.

The 2026 DetectRL-X benchmark includes eight languages, six domains and text generated using four commercial LLMs. It also evaluates operations such as polishing, expansion and condensation. Importantly, it examines a three-way distinction between human-written, LLM-generated and human-written but LLM-refined text, rather than relying entirely on a binary classification. Its findings show that these more realistic distinctions create additional difficulties.

Localising contributions is harder still. SemEval-2024 included a task to identify a transition from human to machine writing, but that is a more constrained problem than reconstructing repeated, interleaved human and AI revisions. A document-level classification should not be mistaken for a verified map of who wrote every sentence.

False positives and writers using a second language

Multilingual fairness is not only about detecting text written in different languages. It also concerns people writing in a language that is not their first.

In a widely discussed 2023 study by Liang and colleagues, seven detectors were evaluated using, among other material, 91 TOEFL essays written by non-native English writers. The mean false-positive rate for those essays was approximately 61%. The study connected this problem to the treatment of relatively predictable linguistic expression and showed that changes in rhetorical complexity could affect detection outcomes.

The scope of that result matters. It describes particular detectors and samples evaluated in 2023. It is not a measured error rate for every detector available in 2026, every non-native writer or every educational context. The comparison groups also differed in characteristics beyond first-language background.

Its lasting significance is methodological: a fairness evaluation should include relevant groups of genuine human writers, not merely demonstrate that a detector separates polished human articles from generated essays.

For an institution evaluating a system, appropriate human examples might include second-language writing, straightforward prose, subject-specific terminology and authorised language assistance. These categories should be tested, rather than presumed safe from false positives.

What multilingual benchmarks actually demonstrate

Several benchmarks show how the field has expanded beyond English-only evaluations:

Research benchmark Scope What it helps evaluate
MULTITuDE, 2023 74,081 texts in 11 languages, using eight multilingual LLMs Transfer across languages and generating models.
MultiSocial, 2025 472,097 texts across 22 languages and five social-media platforms Detection in multilingual, short and informal writing.
AbjadGenEval, 2026 Human-written and generated Arabic and Urdu news articles Detection in two Arabic-script languages, using varied generators.

These figures describe the respective datasets and tasks; they are not equivalent measures of detector quality.

AbjadGenEval reported leading F1 scores of 0.93 for Arabic and 0.89 for Urdu. Those results demonstrate that strong performance is possible outside English in a defined evaluation. They do not establish the same performance on student essays, dialectal messages, translated documents or unrelated languages sharing the same script.

Likewise, it would be misleading to place an accuracy figure from one paper beside a macro F1 score from another and declare a winner. The test sets, class proportions, text lengths, generating models and definitions of “AI-generated” may all differ.

A benchmark result is evidence about a specified experiment, not a universal accuracy certificate.

How to judge multilingual AI-detection accuracy

The following are practical evaluation questions arising from the distinctions above.

Ask what the performance metric measures

Accuracy is the proportion of all examples classified correctly.

Precision asks how many of the documents flagged as AI-generated actually belong to that category.

Recall asks how many of the genuinely AI-generated documents the detector finds.

F1 combines precision and recall. However, “macro F1” needs further explanation: averaging equally across human and machine classes is not the same as averaging equally across languages. A result can be class-balanced while still concealing weak performance in a less-represented language.

An overall score should therefore be accompanied by per-language results and separate false-positive and false-negative measurements.

Examine false positives at realistic prevalence

A hypothetical example illustrates why a low false-positive rate is not enough on its own.

Suppose a detector examines 10,000 documents, of which 100 genuinely belong to its target AI-generated category. Assume it detects 90% of those documents and incorrectly flags 1% of the 9,900 human-written documents.

It would produce 90 correct flags and 99 false flags. Fewer than half of its 189 positive results would be correct, despite its 90% recall and 1% false-positive rate.

These are illustrative assumptions, not measured performance for a particular product. They show why the proportion of positive results that are correct depends partly on how common the target behaviour is in the population being checked.

Test genuine generalisation

A useful evaluation should distinguish performance on familiar material from performance on something genuinely new.

For multilingual detection, that means asking whether the test language appeared in detector training, whether the generating model was familiar, and whether the documents came from the same domains and time period as the training examples.

Dataset splitting also deserves scrutiny. As a methodological safeguard, translations, paraphrases and lightly edited versions derived from the same original document should remain together rather than being scattered across training and test sets. Otherwise, the experiment risks rewarding familiarity with related content instead of genuine transfer.

Separate classification confidence from authorship percentages

A displayed score of “80%” is not self-explanatory.

It could represent a classifier score, a calibrated probability estimate, the proportion of assessed passages that crossed a threshold, or another product-specific measure. These interpretations are not interchangeable.

In particular, confidence that a document belongs to a category does not establish that the same percentage of its words was written by AI. Nor does a document-level probability validate every highlighted sentence. The interpretation should follow the system’s documented output definition and the evidence supporting that definition.

Require versioned, language-specific evidence

An informative evaluation should identify the detector version, test date, languages, text types, input limits and decision threshold.

Language-specific thresholds deserve particular attention. The KInIT work provides an example of investigating calibration by language rather than assuming that one threshold will serve every setting equally well. But any calibration should be evaluated on separate data, not adjusted until it produces attractive results on the final test set.

Can watermarking provide stronger evidence?

Most methods discussed so far infer origin from characteristics of the finished text. Watermarking takes a different approach by introducing a detectable signal during generation.

Kirchenbauer and colleagues’ original text-watermarking framework uses controlled token-selection preferences to create a statistical pattern that can later be tested. This is different from simply judging whether prose resembles known AI output.

Cross-language use still creates difficulties. He and colleagues examined the cross-lingual consistency of text watermarks and found that translation could substantially weaken the tested schemes. They also proposed a method intended to improve resistance to that problem.

Watermarking therefore offers a potentially useful source of evidence, but only under the conditions supported by the particular scheme. A missing watermark does not prove human authorship: the text may have come from a system that did not use that watermark, or undergone transformations that affected it.

More broadly, theoretical and empirical work by Sadasivan and colleagues links detection difficulty to overlap between human and machine text distributions and examines vulnerabilities under textual transformations. This does not establish that every detector is useless. It does explain why universal certainty from text alone is a much stronger claim than good performance on a particular benchmark.

Using multilingual detection responsibly

The evidence supports a proportionate role for multilingual AI detection: an additional analytical signal whose relevance depends on the document and the conditions under which the detector was validated.

A sensible review process should begin with whether the language and writing type are supported by meaningful evidence. An unsupported language, an unusually short passage or a heavily translated document may justify an inconclusive outcome rather than a forced binary verdict.

The next question is what behaviour actually matters. Fully generated work, authorised translation, proofreading and human revision are different processes. In an educational setting, the existence of AI assistance and a breach of assessment rules are separate propositions; research on academic-integrity detection explicitly recognises that AI use is not automatically misconduct.

As a practical safeguard, any consequential review should consider relevant drafts, revision records, sources and the writer’s explanation alongside the detector output. Repeatedly submitting the same document to several tools should not substitute for that examination, especially when the independence of their methods is unknown.

Conclusion

Multilingual AI detection is more than an English-language detector with a longer list of accepted inputs. It involves language coverage, cross-lingual transfer, tokenisation, domain variation, translation, mixed authorship and changing generation systems.

The research demonstrates both progress and limits. Carefully trained systems can perform strongly in defined multilingual settings, while generalisation and transformed text remain important challenges.

The most useful question is therefore not simply, “Does this detector work in this language?” It is:

What was it trained and tested on, what kind of AI involvement does it classify, how often does it make each type of error, and what conclusions can that evidence reasonably support?

PlagPointer · Plagiarism + AI detection

Highly accurate
plagiarism & AI detection.

Detect copied text, paraphrasing and AI-generated writing in a single scan.

From 18p per 250 words

Both checks included. No subscription. Credits never expire.

  • Over 99% AI detection accuracyA highly accurate detection engine with a low false-positive rate in provider testing.
  • Search 60 trillion web sourcesCheck for copied and paraphrased text, with 16,000+ open-access journals covered too.
  • Detailed, downloadable reportsSee matching sources and AI findings in a clear PDF report, saved to your account.
  • Your work stays privateNo notifications to your university or employer. No shared database uploads by default.

*Choose plagiarism detection, AI detection, or both for the same price. AI detection requires at least 350 words.

Leave a Comment

Find us on: