New Breakthrough AI Detection. Our best yet for AI-written source code. Catches 90% of AI/GPT code with 1.3% false flags. 90% caught, 1.3% false flags. Read more Read more
New: MCP integration. Run plagiarism scans from Claude, Cursor or any AI assistant. Run scans from Claude or Cursor. Set it up

Code Intelligence Hub

Expert insights on AI code detection and academic integrity

AI-Generated Code Detection: The New Frontier in Academic Integrity
Featured

AI-Generated Code Detection: The New Frontier in Academic Integrity

As AI coding assistants become ubiquitous, learn how institutions are adapting to detect AI-generated code and maintain educational standards.

Codequiry Editorial Team Codequiry Editorial Team · Jan 5, 2026
Read More →

Latest Articles

Stay ahead with expert analysis and practical guides

Can a Python Submission Be Traced Back to a Java Repository? General 15 min
Alex Petrov Alex Petrov · 1 week ago

Can a Python Submission Be Traced Back to a Java Repository?

Two Python submissions scored 4% against each other and in the 70s against a Java gist from 2017. Cross-language plagiarism is the fastest-growing blind spot in academic integrity because translation destroys the text while preserving everything that matters. Here's what survives a translation, what detectors actually see, and where the false positives come from.

How Code Plagiarism Detection Went From Hashes to LLMs General 11 min
Marcus Rodriguez Marcus Rodriguez · 1 week ago

How Code Plagiarism Detection Went From Hashes to LLMs

Ottenstein's 1976 detector hashed student Fortran token streams, and most of what we run today is a refined version of the same idea. This is the fifty-year arc from line diffs to winnowing, AST matching, web crawling, and statistical AI detection, plus the failure mode that still bites: a 0% similarity score that tells you nothing about authorship.

Token and AST Normalization in Code Similarity Detection General 10 min
David Kim David Kim · 1 week ago

Token and AST Normalization in Code Similarity Detection

Line diffing under-reports copied code and over-reports similar-looking code. Here's what token normalization and AST fingerprinting actually compare, where each one breaks, and how to wire both into a CI pipeline or an academic submission workflow.

How AST Comparison Catches Refactored Code Plagiarism General 12 min
Rachel Foster Rachel Foster · 1 week ago

How AST Comparison Catches Refactored Code Plagiarism

Renaming variables and swapping a for loop for a while loop defeats simple text matching, but it rarely defeats structural comparison. This report walks through the obfuscation ladder, the algorithms that climb it, and the published detection rates behind the claims, including the cases where every engine still misses.

Interpreting Code Similarity Scores in Programming Courses General 9 min
Priya Sharma Priya Sharma · 2 weeks ago

Interpreting Code Similarity Scores in Programming Courses

Similarity scores are ranking signals, not verdicts. I'll walk through the distributions, thresholds, and triage rules I use when reviewing code similarity reports for 400-student courses, plus where AI-generated code fits in the same queue.

The Long Road to Refactoring-Resistant Code Plagiarism Detection General 6 min
Alex Petrov Alex Petrov · 3 weeks ago

The Long Road to Refactoring-Resistant Code Plagiarism Detection

A hands-on retrospective on how code similarity detection grew from naive line diffs to tokenization, ASTs, and fingerprinting. Follow a step-by-step Python prototype and a production workflow with Codequiry to catch refactored plagiarism in CS courses.

A Short History of AI-Generated Code Detection General 8 min
Dr. Sarah Chen Dr. Sarah Chen · 3 weeks ago

A Short History of AI-Generated Code Detection

A CS professor traces how AI-generated code detection grew out of MOSS-era token fingerprints, code stylometry, and a broken similarity assumption. The piece explains how modern detectors work, where they still stumble, and why stacked peer, web, and AI signals make the most defensible academic workflow.

A TA's Script for Sorting 300 Code Similarity Reports by Office Hours General 8 min
Rachel Foster Rachel Foster · 3 weeks ago

A TA's Script for Sorting 300 Code Similarity Reports by Office Hours

A teaching assistant at UC San Diego reduced a 312-submission similarity queue to a shortlist of 14 files in about two hours. The workflow relies on Codequiry's outlier scoring, a Python triage script, and a strict two-pass review rule. Here is the exact process, including the script and the thresholds she uses.

How Open Source License Scanning Evolved From SPDX to Policy Gates General 7 min
David Kim David Kim · 3 weeks ago

How Open Source License Scanning Evolved From SPDX to Policy Gates

License scanning went from hand-checking COPYING files to SPDX metadata and policy-as-code gates. This piece traces that arc, shows which gaps still let copied code slip through, and explains how to pair license scans with source matching for audit-grade evidence.

Why a 400-Student Intro Course Adopted Layered Code Checks General 11 min
Dr. Sarah Chen Dr. Sarah Chen · 4 weeks ago

Why a 400-Student Intro Course Adopted Layered Code Checks

One CS department's switch from a single similarity tool to a layered detection workflow changed what they could see in student code. Peer copying, web sources, and AI-generated submissions each required different signals, and combining them revealed more than any one check alone.

What 41,000 Code Submissions Reveal About Similarity Score Thresholds General 9 min
Priya Sharma Priya Sharma · 1 month ago

What 41,000 Code Submissions Reveal About Similarity Score Thresholds

Across 41,000 student submissions at a large public university, an 85% token-level similarity score between two students predicted confirmed misconduct 92% of the time in introductory courses. This guide walks through the exact calibration workflow, score distributions by language, and tiered review thresholds that worked for my assessment team.