New Breakthrough AI Detection. Our best yet for AI-written source code. Catches 90% of AI/GPT code with 1.3% false flags. 90% caught, 1.3% false flags. Read more Read more
New: MCP integration. Run plagiarism scans from Claude, Cursor or any AI assistant. Run scans from Claude or Cursor. Set it up
James Okafor

James Okafor

Developer Advocate at Codequiry

James writes about code integrity for practicing engineers and helps teams wire Codequiry into their CI and review pipelines.

Articles by James Okafor

Does Convergent AI Output Look Like Peer Plagiarism to a Detector? General 10 min
James Okafor James Okafor • 2 weeks ago

Does Convergent AI Output Look Like Peer Plagiarism to a Detector?

Twenty-three of 412 submissions in a data structures course shared a Dijkstra implementation that differed by fewer than three tokens, and the peer similarity engine flagged all 23 as a single cluster. Nobody had copied anybody. Here is the mechanism behind convergent AI output, the identifiers and AST evidence that separated it from real collusion, and the three-pass workflow we used to keep the scores from contaminating each other.

What Cross-Language Code Plagiarism Detection Can and Cannot See General 11 min
James Okafor James Okafor • 2 weeks ago

What Cross-Language Code Plagiarism Detection Can and Cannot See

Cross-language code plagiarism detection compares normalized structure rather than raw text, which works when a translation was mechanical and fails when the student rewrote the algorithm. Here is what survives a Java-to-Python translation, what the token and IR approaches actually see, and how to run the check across a whole cohort without drowning in false positives.

Scanning 9,301 Python Files for Stack Overflow Copy-Paste General 10 min
James Okafor James Okafor • 1 month ago

Scanning 9,301 Python Files for Stack Overflow Copy-Paste

A practical, code-level guide to batch-scanning Python files for web-sourced code. We walk through token normalization, fingerprinting, uploading to Codequiry, interpreting web match URLs, and stacking an AI check on flagged files. Built for CS professors auditing assignments and engineering managers verifying contractor code.

A TA's Method for Refactoring-Resistant Plagiarism Checks General 9 min
James Okafor James Okafor • 1 month ago

A TA's Method for Refactoring-Resistant Plagiarism Checks

Line diffs collapse when students rename variables, reorder functions, and extract methods. Here's how token fingerprints and AST normalization catch refactored copies, and how to triage hundreds of similarity reports without drowning in false positives.

Teaching Web Code Plagiarism Detection With Real Student Cases General 7 min
James Okafor James Okafor • 1 month ago

Teaching Web Code Plagiarism Detection With Real Student Cases

Web code plagiarism hides in plain sight when students copy from Stack Overflow, GitHub, or tutorials and rename a few variables. This post shows how to teach detection as a skill, design assignments that surface copied web code, and use a source-aware checker like Codequiry to see the evidence.

ChatGPT vs Copilot vs Gemini Code Detection Benchmarked General 8 min
James Okafor James Okafor • 1 month ago

ChatGPT vs Copilot vs Gemini Code Detection Benchmarked

A head-to-head evaluation of AI code detection across ChatGPT-4o, GitHub Copilot, Claude 3.5 Sonnet, and Gemini 1.5 Pro. One pattern kept surfacing: text-only detectors miss refactored LLM code, while structural and multi-signal checks hold up.

Token Fingerprinting vs AST Matching on 1,000 Refactored Java Programs General 11 min
James Okafor James Okafor • 1 month ago

Token Fingerprinting vs AST Matching on 1,000 Refactored Java Programs

When students rename variables, extract methods, and reorder statements to hide copied code, which detection algorithm actually holds up? A controlled experiment pits winnowing, token-based matching, and AST structural hashing against a ladder of refactoring transformations — and reveals why single-technique checkers miss the cases that academic-integrity panels care about most.

Finding Stack Overflow Code in Student Submissions With Fingerprints General 9 min
James Okafor James Okafor • 2 months ago

Finding Stack Overflow Code in Student Submissions With Fingerprints

A study of 5,000 Java assignments from three US universities found that nearly one in four contained code blocks directly traceable to Stack Overflow answers — yet traditional similarity checkers missed them all. We applied token-sequence fingerprinting and a web index of 1.2 million programming snippets to surface hidden web plagiarism at scale.

How Perplexity-Based AI Code Detectors Actually Work General 11 min
James Okafor James Okafor • 2 months ago

How Perplexity-Based AI Code Detectors Actually Work

Perplexity-based detectors aren’t magic — they measure how surprising a sequence of code tokens would be to a model trained on real human code. This report breaks open the inner math, real false-positive rates from Stanford and Edinburgh benchmarks, and why the strongest detectors stack statistical signals with AST fingerprinting and web-source checks.

Bootcamps Are Automatically Scanning Every Submission for AI Code General 6 min
James Okafor James Okafor • 2 months ago

Bootcamps Are Automatically Scanning Every Submission for AI Code

Coding bootcamps that graduate thousands of developers a year are shifting from manual spot-checks to automated AI code detection on every single submission. Here’s a hands‑on guide to building that pipeline with Codequiry’s API — from single‑file scanning to a full CI‑grade batch checker you can run in a GitHub Action.

How Code Similarity Detection Grew From Diff to AI General 8 min
James Okafor James Okafor • 2 months ago

How Code Similarity Detection Grew From Diff to AI

From the early days of the Unix diff command to the rise of MOSS, JPlag, and AI-powered detectors, code similarity detection has undergone a quiet revolution. This retrospective traces the key technical milestones—tokenization, ASTs, fingerprinting, web-source matching, and the new frontier of AI-generated code—showing how each layer made plagiarism harder to hide. See how modern platforms like Codequiry unify these techniques into a single pipeline.

Running a Side-by-Side Accuracy Benchmark of Code Plagiarism Checkers General 8 min
James Okafor James Okafor • 2 months ago

Running a Side-by-Side Accuracy Benchmark of Code Plagiarism Checkers

We ran 2,400 real student Java assignments through MOSS, JPlag, Turnitin, and Codequiry with known ground truth. The F1 scores and false positive rates diverged by over 20 percentage points. Here’s what worked, what broke, and how to run your own benchmark.