If a detector told you your Arabic text was AI-generated when you wrote every word of it yourself, the likely cause is not your writing ability. It is that you ran the finished draft through a rewriting tool before submitting. A study from the College of Applied Computer Science at King Saud University tested 400 human-written Arabic articles and 16,400 lightly edited versions of them. Changing just 10% of the words, with the meaning left intact, dropped Originality.AI's accuracy on Arabic from 92% to 12%.
The common mistake is assuming a detector answers the question "did this writer use AI?" It does not. It answers a much narrower one: do this text's statistical patterns resemble what language models produce? Rewriting tools work by that same logic — they swap your word for the most probable word in its place. Polishing three sentences can be enough to move your text from one bucket to the other. These results are probabilistic in every case, and they misfire more often for second-language writers, and more often still in Arabic.
What does "light polishing" actually mean in this study?
Not rewriting, not summarizing, not expanding. It means asking a model to change a set percentage of the words without touching the meaning or the length. The researchers built a dataset called Ar-APT: 400 human-written Arabic articles, each polished by ten different models at four levels (10%, 25%, 50%, and 75%), producing 16,400 samples.
To confirm the meaning survived, they measured semantic similarity between each original and its polished version. At the 10% level, every model held similarity above 92% except Qwen-3. So a human reader compares the two versions and finds them saying the same thing, while a detector reads them and calls one human and the other machine.
The working rule: if your words changed but your meaning did not, you are in the danger zone. That is precisely the situation of a student who wrote a paper themselves, then asked ChatGPT to "just improve the phrasing" at midnight before the deadline.
How far does detector accuracy fall after only 10% polishing?
The answer depends on which tool did the polishing, not on how much help you took. Here is Originality.AI on human-written Arabic articles polished by 10%. The figure represents how many stayed correctly classified as human:
- Polished by Mistral or Gemma-3: 12% — meaning 88% of genuinely human articles were flagged.
- Polished by Qwen-3: 18%.
- Polished by GPT-4o: 23%.
- Polished by Claude-4 Sonnet: 90% — almost unaffected.
The gap between 12% and 90% has nothing to do with the writer's honesty or the degree of intervention. It reflects which model's output the detector was trained on. At the 75% level with GPT-4o polishing, the figure fell to zero: not one article out of the hundred tested survived.
ZeroGPT started from a worse baseline. Before any polishing at all, it labeled 38% of human Arabic articles as machine-written. After 10% polishing by Mistral, only 31% remained correctly classified. See how tools diverge on Arabic text before you build a decision on a single score.
Did any detector come through the test intact?
No. The study tested four commercial tools and ten language models prompted to act as detectors. On raw Arabic with no polishing whatsoever:
- Originality.AI: 96% overall accuracy, 8% false positive rate — the strongest of the group.
- ZeroGPT: 80% accuracy, 38% false positive rate.
- Isgen: 57% false positive rate.
- Smodin: 80% false positive rate.
Using a chat model as a detector — a habit some instructors have picked up — performs worse than it looks. GPT-4o correctly called 99% of human texts human, but caught only 15.5% of machine-written ones. A detector that accuses nobody is not accurate; it is broken. Running the other direction, GPT-3.5 flagged roughly 94% of human texts.
The rule is blunt: do not ask a chat model "is this AI?" and do not let its answer decide anything about a person. Compare what each tool actually measures instead of trusting a bare percentage.
Why is Arabic harder for detectors than English?
Three concrete reasons. First, training data scarcity. High-quality human Arabic prose is far less abundant than English in detector training sets, so tools learn what "normal Arabic" looks like from a narrow sample and treat anything outside it as anomalous.
Second, morphology and dialect. Arabic generates dozens of forms from a single root, and Arabic writers shift between registers within one piece. That variety itself sometimes reads to a classifier as stylistic inconsistency.
Third, and strangest: diacritics. A study in the journal Information addressed this directly and found GPTZero managed only 62.7% accuracy on the AIRABIC benchmark, while models trained on diacritized text reached 98.4%. Marks above the letters, not the meaning of the sentence, can flip the classification. This edge case hits writers working with classical or religious material hardest. A deeper breakdown of Arabic detection challenges.
What should you do before you submit?
The goal is not to hide. The goal is to avoid manufacturing evidence against yourself that you cannot later explain.
- Keep a record of your writing. Google Docs version history or Word's tracked changes is the only evidence that holds up in front of a committee. A detector score does not.
- Disclose the help you took. If you used a tool for editing, say so according to your institution's policy. Policies vary enormously between universities, so read yours rather than someone else's.
- Do not run the whole draft through a rewriting tool at the last minute. Edit what needs editing yourself, even if it is slower.
- Ignore any result on text under 300 words. Short passages give a detector too little signal, and the output becomes noise.
- Check with a tool built for Arabic. Check your Arabic writing before you submit, and read the score as a prompt to review, not a verdict that closes the file.
What do these numbers not tell you?
The study is a preprint and has not completed peer review. The commercial tools were tested manually on 100 of the 400 articles, because the researchers could not obtain API keys to automate the process. And the human articles were drawn from material published in 2020 or earlier to guarantee they predate widespread LLM use — a genuine strength, but it means the style measured is journalism and blogging, not graduate coursework.
More importantly, detectors get updated. Today's figure may not hold in six months. And no detector is perfectly accurate, including ours. Everything these tools produce is a probability, and any institution that builds a penalty on a probability alone is building on sand. Read how we measure and what we acknowledge we cannot do.
Frequently asked questions
Does this mean detectors are useless?
No. Originality.AI identified 100% of fully machine-generated Arabic articles with no polishing applied. The failure runs in the other direction: lightly edited human text. The defensible use is as a signal that opens a review, not as evidence that closes a case.
My text was already flagged. What do I use to defend myself?
Version history, early drafts, and a request for an oral discussion where you explain your argument. Ask which tool produced the score, too — the figures above show the tool itself is a significant variable. How to handle a false accusation, step by step.
Does translating from English raise the score?
The study did not measure translation, so I cannot put a number on it. The same mechanism plausibly applies: any process that replaces your wording with more probable phrasing pushes the text toward the machine bucket. Treat it as a reasonable expectation, not a proven finding.
Do diacritics change the result?
According to one study, yes — and that is evidence of fragility, not a technique to adopt. The effect is unstable across tools, and using it to manipulate a score is neither honest nor reliable. Its real value is explaining why results swing so wildly on classical Arabic texts.
Written by
Founder of Sahihly
Founder of Sahihly. I build writing-quality tools for Arabic and English, and write about AI detection and its limits.
All articles by this writer →Try the detector and humanizer now — free, no account required.
Open the studioKeep reading
Falsely Flagged by an AI Detector, a Family Sued the University — and Won
The fear of a false AI-cheating accusation is no longer hypothetical. One family spent six figures in legal fees proving their son never used AI — and the university's detector was simply wrong.
ArabicDoes Turnitin Detect AI in Arabic? The Official Answer
Turnitin's AI detection officially covers English, Spanish and Japanese only. Arabic submissions get no report at all — which is not the same as being unchecked.