You can usually tell if something was written by AI by reading for three things at once: rhythm, specificity, and stake. AI prose keeps a flat, even sentence length, stays abstract where a person would name a date or a dollar figure, and never risks an opinion it might have to defend. Run the text through a detector such as GPTZero for a second opinion, then confirm with process evidence like drafts or version history. No single signal is proof.
The Three-Layer Check
Read first, scan second, verify third.
Layer one is your own eye. Look for uniform paragraph lengths, tidy tricolon lists, and transition words doing work that logic should be doing.
Layer two is a detector. GPTZero classifies text against major models including ChatGPT, Claude, Gemini and Llama, and reports a probability rather than a verdict.
Layer three is provenance. Ask for the working file. A real draft has messy revision history, dead ends, and comments. A generated draft usually arrives clean and finished on the first save.
Understanding AI Writing: What the Machine Is Actually Doing
A large language model does not know anything. It predicts the next chunk of text based on statistical patterns in what it read during training. That single fact explains almost every tell you will learn to spot.
Why prediction produces smooth prose
Because the model picks likely continuations, it drifts toward the median sentence. Average length. Average vocabulary. Average structure. Human writing is lumpy by comparison: a four-word sentence lands next to a thirty-word one because the writer was excited, or tired, or making a point.
Why it produces confident vagueness
The model has no source in front of it unless you give it one. So it reaches for safe generalities. “Studies show engagement improves” is cheap to generate. “Open rates went from 18% to 26% over eleven weeks” is not, because it requires a real number from a real place.
What “AI-assisted” muddies
Most text you will examine is not pure. Someone generated a draft and edited it, or wrote a draft and asked a model to polish it. Detection tools were built for the first case. The second case, human ideas in machine phrasing, is where confidence should drop hardest.
The models most people are using
ChatGPT, Claude and Gemini dominate general writing. Each has house habits. Knowing them helps: one leans on em dashes and bulleted summaries, another on cautious both-sides framing, another on cheerful second person. Spend twenty minutes generating samples from each and the fingerprints get obvious.
Common Signs of AI-Generated Text
None of these is conclusive alone. Three or more together, in a piece with no verifiable specifics, is a strong signal.
- Sentences that hover between 15 and 22 words for pages at a time.
- Paragraphs that all run three to four sentences, like bricks.
- Lists that always arrive in threes.
- Vocabulary tells: delve, tapestry, testament, landscape, realm, underscore, pivotal, robust.
- Hedged openers stacked in a row: “It’s important to note”, “That said”, “Ultimately”.
- A summary sentence at the end of every section, restating what you just read.
- Named claims with no source, or a source that turns out not to say that.
- Perfect punctuation with zero typos across 2,000 words.
- No first-person detail: no time, no place, no cost, no name.
- A conclusion that recommends “staying informed” and nothing sharper.
The specificity test
Pick any paragraph and ask: could this sentence have been written by someone who had never done the thing? If yes for the whole piece, you are probably looking at generated text. Real practitioners leak details. They mention the client who cancelled, the tool that broke, the invoice that arrived late.
The stake test
Human writing takes positions that could be wrong. Generated writing hedges toward the consensus, because the consensus is the statistically likely continuation. If nothing in the piece could be argued with, that is information.
The rhythm test
Read a paragraph aloud. If your breathing stays perfectly even, that is a machine cadence. Humans vary. They interrupt themselves.
Tools for Detecting AI Writing
Detectors are useful as a second opinion and dangerous as a verdict. Use them to open a conversation, never to close one.
| Tool | Best for | What it reports | Worth knowing |
|---|---|---|---|
| GPTZero | Classroom and editorial screening | AI probability across ChatGPT, Claude, Gemini, Llama | Also offers writing-process replay to prove authorship |
| Grammarly’s AI detector | Writers checking their own drafts | Percentage of text flagged as AI | Sits inside a tool many writers already use |
| Scribbr | Students on a budget | Sentence-level highlighting | Free tier is generous; academic framing |
| Copyleaks | Institutions and publishers | Segment-level flags, multi-language | Built for bulk scanning workflows |
How the classifiers actually decide
Most detectors score two properties. Perplexity measures how surprising each word is given the ones before it. Burstiness measures how much sentence-to-sentence variation there is. Generated text tends to score low on both, because prediction smooths surprise out. GPTZero publishes accuracy claims for its detection model, and pairs the score with authorship verification features rather than presenting the number alone. Grammarly similarly returns a share-of-text figure instead of a binary label.
Reading a score correctly
A 90% result does not mean the text is 90% machine written. It means the classifier is fairly confident the pattern matches its AI training distribution. Those are different claims. Treat any score between roughly 30% and 70% as noise.
A worked example
An editor receives a 1,400-word contributor piece. The detector returns 82% AI. Before acting, she does three things. She checks the Google Docs version history and finds a single paste event. She searches two statistics in the piece and finds neither exists. She asks the writer to explain a claim in the fourth paragraph, and gets a rewrite instead of an explanation. Now she has a case. The detector alone was not one.
Limitations of AI Detection
This is the part most articles skip, and it matters more than the tool list.
False positives hit predictable groups
Writing that is plain, formulaic, or non-native-English tends to score as machine written, because it looks statistically unsurprising. Technical documentation, legal boilerplate and ESL student essays all get flagged more than they should. Accusing someone on that basis is a real harm.
Light editing defeats most detectors
Change a dozen words, split three sentences, add one personal aside, and scores often drop below the threshold. Paraphrasing tools are built specifically to do this. Any detector that could not be beaten in five minutes would also flag half of ordinary human prose.
Short samples are unreliable
Under roughly 300 words there is not enough signal. Treat a scan of a single paragraph as entertainment.
The target keeps moving
Detectors are trained on the output of models that already exist. Each new model release shifts the distribution, and accuracy claims made against last year’s output do not automatically transfer.
What to do with all that
Weight process evidence above scores. Version history, drafts, an interview, a five-minute conversation about the argument. Those are hard to fake and easy to check.
Ethical Considerations
In education
A detector score is an invitation to talk, not evidence for a grade penalty. Institutions that act on scores alone end up disciplining the students least able to defend themselves. Better policy: state what tool use is allowed, ask for process artefacts by default, and treat flags as one input in a documented review.
In hiring and publishing
Screening submissions is legitimate. Rejecting a candidate silently on a probability score is not, especially when the underlying model is opaque. Tell people you scan, and tell them how a flag will be handled.
For the person being accused
You are entitled to ask what tool was used, what score triggered the concern, and what evidence beyond the score exists. Keep your drafts. Turn on version history before you need it.
Practical Tips for Writers
If you write honestly and still get flagged, the fix is usually style, not innocence.
- Vary sentence length deliberately. Put a five-word sentence after a long one.
- Name things. Real dates, real prices, real product versions.
- Cut the throat-clearing openers. Start sentences with the subject.
- Kill one list per draft and turn it into prose.
- Add one thing only you could know: a mistake, a cost, a conversation.
- Drop the section-summary sentences. Trust the reader.
- Use the words you actually say out loud. If you have never said “tapestry”, do not write it.
- Take one clear position per piece.
- Keep drafts in a versioned document from the first line.
- If you used a model, say so, and say what for.
If you are already flagged
Do not rewrite in a panic. Gather your version history, your notes, your browser history for the sources you read. Then rewrite the three weakest paragraphs with concrete detail and resubmit both versions. Showing your working beats arguing about a percentage.
Product, Course, App and Platform Experience
Most detector platforms now behave like small suites rather than single-purpose scanners. GPTZero bundles detection with a plagiarism checker, a hallucination checker for fabricated citations, and integrations for Google Docs, Canvas and Chrome, according to its product pages. Grammarly folds detection into the editor writers already have open, so the check happens where the drafting happens, per Grammarly’s detector page. Scribbr and Copyleaks sit at opposite ends of the same market, one aimed at individual students, the other at institutional bulk scanning.
The practical difference is workflow, not accuracy. A teacher grading 90 essays needs LMS integration. A freelancer needs a browser tab. Pick for where the text already lives, and verify current pricing on the official site before you commit to a plan, since tiers and free limits change often.
If you want to understand the models producing this text rather than only the tools policing it, you can explore Coursiv AI lessons and learn how generation and detection actually work.
Decision Framework: What to Know Before Deciding
Before you act on any judgment about authorship, run this framework.
- How high are the stakes? A blog comment and a dissertation deserve different thresholds.
- How long is the sample? Under 300 words, stop.
- Is there process evidence? Version history outranks any score.
- Who is the writer? Non-native speakers and technical writers get false positives more often.
- What is your policy? If you have not published one, you cannot enforce one fairly.
- What happens if you are wrong? Design the process around that answer.
Use detectors to prioritise attention, not to assign guilt. The reliable workflow is: read for the human tells, scan for a second opinion, then ask for the working file. Two out of three pointing the same way is a conversation. One is nothing.
Conclusion and Next Steps
Telling AI writing from human writing is a reading skill first and a tooling problem second. Learn the tells: flat rhythm, vague specifics, absent stake. Then use a detector as a tiebreaker, knowing it produces false positives on plain and non-native prose. Finally, ask for provenance, because process evidence is the only signal that is genuinely hard to fake.
Start by scanning three texts you already know the origin of. A piece you wrote, a piece a colleague wrote, and a piece you generated. Compare what the tool says with what you know is true. That calibration will teach you more than any accuracy claim.
The two follow-ups worth reading next are how accurate are ai content detectors and does relying on ai hurt your skills.