Perplexity and Burstiness

Table of Contents

Need Help With Your Academic Work?

Get expert, reliable support for assignments, essays, research, and editing — delivered on time and plagiarism-free.

 

Perplexity and Burstiness in AI Detection: Simple Meaning

You wrote your essay yourself, in your own words, over three late nights. You paste it into an AI checker just to be safe, and it comes back with a high AI score. A friend says it is because of perplexity and burstiness, and that you should add more of it. Nobody explains what that means, and your deadline is in two days.

Here is the short answer. Perplexity measures how predictable your word choices are to a language model. Burstiness is about how much your writing rhythm changes from sentence to sentence. Early AI detectors used both as clues, but the best known detector now says it works differently, so a score is a hint, not proof. This guide explains both ideas with real examples, shows you a free way to check your own writing, and tells you what to do if you are flagged. For the bigger picture, start with our guide on how AI detectors work.

What Perplexity and Burstiness Mean in Plain Words

What Is Perplexity?

Perplexity is a number that shows how surprised a language model is by the words you chose. A language model is the kind of program behind chatbots, and it works by guessing the next word in a sentence. If you write the word it expected, perplexity is low. If you write something it did not expect, perplexity is high. Hugging Face, which publishes widely used AI tools, defines perplexity in its documentation as the exponentiated average negative log-likelihood of a sequence. In simple words, it is a kind of average of how unlikely each word was.

Infographic explaining perplexity with next word probabilities for the phrase I like to drink hot and a chance to choices table

Here is a way to picture it. Take the start of a sentence: I like to drink hot. A model might give coffee a 40% chance, tea 35%, chocolate 15% and soup 1%. If you wrote coffee, the model saw it coming. If you wrote soup, it did not. For a single word, perplexity is simply 1 divided by the chance the model gave that word, so you can read it as the number of equally likely choices the model was torn between. A word given a 50% chance scores 2, a 25% chance scores 4, a 10% chance scores 10 and a 1% chance scores 100.

For a whole sentence the model averages across the words. As a made-up example, a four word sentence where the model gave 50%, 40%, 80% and 25% scores about 2.2. Another four word sentence where it gave 2%, 5%, 1% and 3% scores about 42.7. The probabilities here are invented to show the maths, not taken from a real tool.

One more fact that most guides skip: the Hugging Face team notes that perplexity scores are not comparable between models or datasets. That is one reason two detectors can show very different numbers for the same essay.

What Is Burstiness?

Burstiness is about rhythm. People often write in bursts: a long sentence, a short one, a fragment, then a long one again. A lot of AI writing feels more even. That everyday meaning is how most students and websites explain burstiness, as variety in sentence length and structure.

The original GPTZero went a step deeper. A research paper that describes it says GPTZero calculated burstiness as the standard deviation of the perplexity scores of each sentence. So it measured how much predictability jumps around, not only how long the sentences are. In our demo below we use the everyday meaning and call it a variety score, so we do not mix it up with any detector’s number. For more on how the two kinds of writing differ, read our guide on ai writing vs human writing.

 PerplexityBurstiness
What it looks atHow predictable each word isHow much the rhythm changes across sentences
Everyday meaningSurprise in word choiceVariety in sentence length and structure
Original GPTZero meaningPredictability measured with a language modelHow much sentence perplexity varies, as a standard deviation
Low value meansPredictable wordingAn even, steady rhythm
High value meansSurprising wordingA varied, uneven rhythm
Why honest writing can score lowCommon words, simple grammar, formal toneTemplates and repeated sentence patterns

See It on Real Paragraphs: Our Sentence Variety Demo

We wrote four short paragraphs on the same topic, sleep and study, in four different styles. Then we counted the words in every sentence and worked out a simple variety score: the standard deviation of sentence length divided by the average sentence length. A low score means an even rhythm. A high score means a varied rhythm. These are samples we wrote for this demo. They are not real student work and not detector output.

The Four Sample Paragraphs

Paragraph A, even and formal: Sleep plays an important role in the academic success of university students. Many students sleep too little because of heavy workloads and busy social lives. Research shows that a lack of sleep can reduce memory, attention and overall performance. Universities should therefore encourage healthy sleep habits among their students. Simple changes such as a regular bedtime can improve both health and learning outcomes. In conclusion, sleep is a key factor that every student should take seriously.

Sentence lengths in words: 12, 13, 14, 10, 14 and 13. Average 12.7, variety score 0.11.

Paragraph B, varied and personal: I pulled my first all-nighter in week three. It worked, sort of. I handed in the essay on time, then slept through the lecture I actually cared about and spent the next two days feeling like a cracked phone screen. Nobody warns you that tiredness compounds. A late night is not a one-off cost you pay once; it quietly bills you again on Thursday. So now I set an alarm to go to bed, which sounds ridiculous, and it has done more for my grades than any study app.

Sentence lengths in words: 8, 4, 28, 6, 18 and 25. Average 14.8, variety score 0.63.

Paragraph C, simple second-language style: Sleep is very important for students. Students must sleep eight hours every night. If students do not sleep, they feel tired in class. If they feel tired, they cannot study well. Good sleep helps students remember new information. Good sleep also helps students feel happy. For this reason, students should go to bed early.

Sentence lengths in words: 6, 7, 10, 8, 7, 7 and 9. Average 7.7, variety score 0.17.

Paragraph D, lab report style: Participants were asked to record their sleep duration for seven days. Sleep duration was measured using a wearable device. Test scores were collected at the end of each week. Scores were compared between the short sleep group and the long sleep group. A significant difference was observed between the two groups. These results suggest that sleep duration is associated with test performance.

Sentence lengths in words: 11, 8, 10, 13, 9 and 11. Average 10.3, variety score 0.15.

What the Numbers Show

Paragraph B, the personal story, scored 0.63. The other three scored 0.11, 0.17 and 0.15. The surprise is that the lab report style and the simple second language style came out almost as even as the formal paragraph, even though the people writing in those ways have very different reasons. A lab report follows a template. A beginner in English uses short, safe sentences. Neither is cheating. Both look flat on a rhythm measure.

Bar charts of sentence lengths in four sample paragraphs with variety scores of 0.11, 0.63, 0.17 and 0.15

Notice also what paragraph B has that the others lack: a specific story, odd comparisons and a personal opinion. That is what makes it varied. Variety is a side effect of having something specific to say, not a trick.

A warning before you read too much into this. Four tiny samples prove nothing about real detectors, which need hundreds of words and use far more than sentence length. Treat the demo as a teaching example, and use the method later in this guide on your own writing.

What AI Detectors Use Today: Myth vs Fact

GPTZero Moved Beyond Perplexity and Burstiness

GPTZero launched in January 2023 and, according to NPR, scored text using perplexity and burstiness. Today the company tells a different story. Its own explainer page says that as of autumn 2023 GPTZero no longer uses perplexity and burstiness for its AI detection, because it moved to a deep learning model, and that they remain one of seven indicators. GPTZero’s page on how AI detectors work adds that its model now looks at layers such as syntax, word choice and meaning. If you want to see how it performs in tests, read our review: is gptzero accurate.

What Turnitin Says

A developer article that quotes Turnitin’s FAQ says Turnitin states its model is not explicitly programmed to evaluate signals such as burstiness or perplexity, and instead learns statistical patterns from its training data. We could not open Turnitin’s FAQ page ourselves, so check its current wording before you quote it.

Why Two Detectors Give Two Different Scores

Different detectors use different models and different training data, and as we saw, perplexity numbers do not carry over between models. The percentage on screen can also mean different things in different tools. If you want help reading one, our guide on ai detection score meaning explains what the numbers and labels usually mean.

Timeline from 2023 to today showing GPTZero launch, the Stanford study, GPTZero's method change and current statements

MythFact
AI detectors only measure perplexity and burstiness.Early GPTZero did. GPTZero says it moved to deep learning in autumn 2023, and Turnitin says its model is not programmed to score them.
Burstiness just means sentence length.That is the everyday meaning. The original GPTZero measured how much sentence perplexity varied.
A low perplexity score proves AI wrote it.Common words and simple grammar also give low perplexity, including honest writing in a second language.
You can compare scores between tools.Perplexity is not comparable between models or datasets.
A detector score is proof.A score is a probability. Your drafts and notes are your evidence.

Why Honest Student Writing Can Look Flat to a Detector

The biggest risk for students is a false flag. In 2023, Stanford researchers tested seven GPT detectors in a study published in the journal Patterns. The detectors judged US student essays almost perfectly, but flagged an average of 61.3% of essays by non-native English writers as AI. All seven flagged 19.8% of those essays, and at least one flagged 97.8%. The essays flagged by every detector had significantly lower perplexity. One of the authors explained that common English words lead to low perplexity scores. The study dates from 2023 and did not single out one tool, and newer detectors may do better, but it shows why simple or careful writing can look flat. Our guide on ai detection for non-native English students covers this in more depth.

SituationWhy it can look flatWhat to keep or do
Writing in a second languageCommon words and simple grammar can lower perplexityKeep drafts, notes and any tutor feedback
Lab reports and templatesRepeated sentence patterns lower varietyKeep raw data, lab notes and the template your course gave you
A very formal toneFormal phrases are predictableKeep your outline and sources with page numbers
Very short textGPTZero warns that text under 100 words may be less accurateCheck the full draft, not a few lines
Heavy grammar tool editsSmoothing every sentence can even out rhythmKeep your first draft and the version history
Following a model essay too closelyBorrowed structure repeats a patternPlan in your own words and note where ideas come from
Six reasons honest student writing can look flat to AI detectors, with what to keep as proof for each

On grammar tools, a study on identifying and understanding AI-generated text reports that popular writing aids, even ones that are not mainly generative, can trigger false positives. That does not mean you should stop using them. It means you should keep your own first draft, so you can show how the work changed.

Check Your Own Sentence Variety for Free

You can repeat our demo on your own writing in two minutes with Google Sheets or Excel. This does not copy a detector. It is a mirror that shows whether every sentence in your draft has a similar length.

  1. Paste your text into a sheet with one sentence in each row of column A, from A2 down.
  2. In B2, count the words in the sentence, then copy the formula down the column.

=LEN(TRIM(A2))-LEN(SUBSTITUTE(TRIM(A2),” “,””))+1

  • Find the average sentence length in any empty cell.

=AVERAGE(B2:B20)

  • Find how much the length varies.

=STDEV.P(B2:B20)

  • Divide the second result by the first to get your variety score.

=STDEV.P(B2:B20)/AVERAGE(B2:B20)

Five step guide with spreadsheet formulas to measure your own sentence length variety in Google Sheets or Excel

How to Read Your Score

There is no pass mark, and no detector uses this exact number. A low score means every sentence is about the same length. A high score means your sentences vary a lot. Use it as a prompt to reread your draft aloud. If every sentence runs 12 to 14 words, ask whether some ideas deserve a short, direct line, and whether others need a longer explanation.

Please do not rewrite just to chase a number. Good writing varies because ideas vary. The best way to improve it for your reader is to add specific examples, names, numbers and your own reasoning where your assignment allows. And if the work is yours, the strongest protection is your evidence, which comes next.

If You Are Flagged: What to Do

Stay Calm and Ask for the Report

A flag is a signal for your instructor to look closer. It is not a verdict. Ask which tool produced the result, what the score means in that tool, and what your course policy says about detector results. Different tools give different numbers, so the exact report matters.

Build Your Authorship Evidence Kit

Gather six things: your outline, your reading notes with page numbers, your drafts, your version history, your source list and any feedback you received on drafts. Save them as you work, not after a flag, because history that was recorded at the time is more convincing. If the accusation has already been made, our guide for students who feel falsely accused of using ai walks through the steps.

Authorship evidence kit checklist for students: outline, notes, drafts, version history, source list and feedback

Use More Than One Signal

No single tool should decide your case, whether it helps you or hurts you. You can run your draft through our ai content detector as one more signal, and keep a dated screenshot of each result. Then bring the results together with your evidence kit when you talk to your instructor.

The Verdict: What to Take Away

Perplexity and burstiness are useful ideas. They explain why predictable, even writing can look machine-made to a program. But they are not the full story today. The detectors that made the terms famous now say they work differently, scores are not comparable between tools, and honest writing in a second language, in a lab report or in a formal tone can look flat. Treat a score as a hint, check more than one signal, and keep your drafts.

To go deeper into the whole process, read our complete guide to AI detection. It covers the full picture, from training data to the checks detectors run today.

Still Confused and Need Help?

A score does not tell you why your writing looks the way it does, and it does not help you improve it. Skyline Academic is one of the leading 1:1 live tutoring platforms for students, with a dedicated interactive LMS dashboard that keeps your work in one place. Here is what you get.

  • 1:1 personalized live tutoring for students with a tutor who helps you understand your result and strengthen your own writing
  • An interactive LMS dashboard with workshops, bootcamps, student progress tracking and assignment quizzes
  • An AI detection tool to review your writing before you submit it
  • A plagiarism checker to help you keep your work original

Worried about a flag or a high score? Book a session and bring your draft, your result and your questions. We will help you read the result, build your evidence and plan your next step. We always encourage students to follow their school’s academic integrity rules.

Frequently Asked Questions About Perplexity and Burstiness

What is perplexity in AI detection?

Perplexity measures how predictable your words are to a language model. If the model expected your words, perplexity is low. If your word choices surprised it, perplexity is high.

What does low perplexity mean in AI detection?

It means the wording was very predictable. Early detectors treated that as a hint of AI writing. But common words, simple grammar and formal tone also give low perplexity in honest writing.

What is burstiness in writing?

Burstiness is how much your rhythm varies, such as mixing long and short sentences. In everyday use it means variety in sentence length and structure. The original GPTZero measured how much sentence perplexity varied.

What is a good burstiness score?

There is no official good score. Numbers differ from tool to tool, and GPTZero says burstiness is no longer the basis of its detection. Do not chase a number. Focus on clear writing and keep your drafts.

Do all AI detectors use perplexity and burstiness?

No. Early versions of GPTZero did, but many detectors now use trained models. Turnitin says its model is not explicitly programmed to score them. Methods differ and change, so no single rule fits every tool.

Does GPTZero still use perplexity and burstiness?

GPTZero says it stopped using them for detection in autumn 2023 and moved to a deep learning model. It describes them as one of seven indicators in its current approach.

Does Turnitin use perplexity and burstiness?

Turnitin says its model is not explicitly programmed to evaluate them and instead learns statistical patterns from training data. Check Turnitin’s current FAQ for the latest wording.

Why does my own writing have low perplexity?

Common words, simple grammar and a formal tone are all predictable, and writing in a second language often uses them. A Stanford study found essays by non-native writers had lower perplexity. It is not proof of AI use.

Can I measure burstiness myself?

You can measure a simple version. Count the words in each sentence in a spreadsheet, then divide the standard deviation by the average. It is a mirror for your writing, not a copy of any detector.

How do AI detectors measure whether text was written by AI?

Older tools used statistical signals like perplexity and burstiness. Many current detectors use models trained on human and AI text. They give a probability, not proof, which is why you should keep your drafts.

Final Thoughts

Perplexity and burstiness are easy to explain and easy to over trust. They describe how predictable and how even a piece of writing is, and honest writers can look predictable and even for good reasons. Use the ideas to understand what a detector might be reacting to, check your own writing with the free method above, and keep your outline, notes and drafts as you go. If you ever need help making sense of a score, Skyline Academic is here to support you every step of the way.

Get Expert Academic Tips Straight to Your Inbox

Subscribe to get the latest tips, resources, and insights delivered straight to your inbox. Learn smarter, stay informed, and never miss an update!