AI-Generated Code Detection: The New Frontier in Academic Integrity
As AI coding assistants become ubiquitous, learn how institutions are adapting to detect AI-generated code and maintain educational standards.
Expert insights on AI code detection and academic integrity
As AI coding assistants become ubiquitous, learn how institutions are adapting to detect AI-generated code and maintain educational standards.
Stay ahead with expert analysis and practical guides
General
15 min
Two Python submissions scored 4% against each other and in the 70s against a Java gist from 2017. Cross-language plagiarism is the fastest-growing blind spot in academic integrity because translation destroys the text while preserving everything that matters. Here's what survives a translation, what detectors actually see, and where the false positives come from.
General
11 min
Ottenstein's 1976 detector hashed student Fortran token streams, and most of what we run today is a refined version of the same idea. This is the fifty-year arc from line diffs to winnowing, AST matching, web crawling, and statistical AI detection, plus the failure mode that still bites: a 0% similarity score that tells you nothing about authorship.
General
10 min
Line diffing under-reports copied code and over-reports similar-looking code. Here's what token normalization and AST fingerprinting actually compare, where each one breaks, and how to wire both into a CI pipeline or an academic submission workflow.
General
12 min
Renaming variables and swapping a for loop for a while loop defeats simple text matching, but it rarely defeats structural comparison. This report walks through the obfuscation ladder, the algorithms that climb it, and the published detection rates behind the claims, including the cases where every engine still misses.
General
9 min
Similarity scores are ranking signals, not verdicts. I'll walk through the distributions, thresholds, and triage rules I use when reviewing code similarity reports for 400-student courses, plus where AI-generated code fits in the same queue.
General
8 min
A University of Missouri-St. Louis instructor spent two decades watching code similarity tools change shape. The gaps between each generation of detection explain why instructors now run peer, web, and AI checks side by side.
General
6 min
A hands-on retrospective on how code similarity detection grew from naive line diffs to tokenization, ASTs, and fingerprinting. Follow a step-by-step Python prototype and a production workflow with Codequiry to catch refactored plagiarism in CS courses.
General
8 min
A CS professor traces how AI-generated code detection grew out of MOSS-era token fingerprints, code stylometry, and a broken similarity assumption. The piece explains how modern detectors work, where they still stumble, and why stacked peer, web, and AI signals make the most defensible academic workflow.
General
8 min
A teaching assistant at UC San Diego reduced a 312-submission similarity queue to a shortlist of 14 files in about two hours. The workflow relies on Codequiry's outlier scoring, a Python triage script, and a strict two-pass review rule. Here is the exact process, including the script and the thresholds she uses.
General
7 min
License scanning went from hand-checking COPYING files to SPDX metadata and policy-as-code gates. This piece traces that arc, shows which gaps still let copied code slip through, and explains how to pair license scans with source matching for audit-grade evidence.
General
11 min
One CS department's switch from a single similarity tool to a layered detection workflow changed what they could see in student code. Peer copying, web sources, and AI-generated submissions each required different signals, and combining them revealed more than any one check alone.
General
9 min
Across 41,000 student submissions at a large public university, an 85% token-level similarity score between two students predicted confirmed misconduct 92% of the time in introductory courses. This guide walks through the exact calibration workflow, score distributions by language, and tiered review thresholds that worked for my assessment team.