New Breakthrough AI Detection. Our best yet for AI-written source code. Catches 90% of AI/GPT code with 1.3% false flags. 90% caught, 1.3% false flags. Read more Read more
New: MCP integration. Run plagiarism scans from Claude, Cursor or any AI assistant. Run scans from Claude or Cursor. Set it up
Dr. Sarah Chen

Dr. Sarah Chen

AI Detection Researcher at Codequiry

Sarah leads AI-generated code research at Codequiry, focusing on how large-language-model output can be reliably distinguished from human-written code.

Articles by Dr. Sarah Chen

How Cross-Language Code Plagiarism Detection Works General 17 min
Dr. Sarah Chen Dr. Sarah Chen • 1 day ago

How Cross-Language Code Plagiarism Detection Works

A student submits Python. Their partner submits Java. A line diff reports 0% identical text, and the plagiarism checker stays quiet. Cross-language copying is a semantic clone problem, and it needs a different kind of comparison than the token matching most tools provide. Here is how the detection actually works, where it fails, and what to put in your syllabus before next term.

When Does Copied Code Become an Open Source License Violation? General 15 min
Dr. Sarah Chen Dr. Sarah Chen • 1 day ago

When Does Copied Code Become an Open Source License Violation?

Dependency scanners read manifests. They cannot see the 300 lines someone pasted into a file with the header deleted. This piece walks through what actually constitutes an open source license violation, why SBOM tooling goes blind at exactly the wrong moment, and how provenance checks catch copied code before counsel does.

Where Should You Set the Code Plagiarism Score Threshold? General 12 min
Dr. Sarah Chen Dr. Sarah Chen • 2 days ago

Where Should You Set the Code Plagiarism Score Threshold?

A mid-size CS department got 41 similarity flags from a single assignment and no written policy for what any of them meant. This is the calibration exercise they ran, the AI cluster that confused everyone, and the starter-file mistake that produced 61 false 100% matches.

A Short History of AI-Generated Code Detection General 8 min
Dr. Sarah Chen Dr. Sarah Chen • 3 weeks ago

A Short History of AI-Generated Code Detection

A CS professor traces how AI-generated code detection grew out of MOSS-era token fingerprints, code stylometry, and a broken similarity assumption. The piece explains how modern detectors work, where they still stumble, and why stacked peer, web, and AI signals make the most defensible academic workflow.

Why a 400-Student Intro Course Adopted Layered Code Checks General 11 min
Dr. Sarah Chen Dr. Sarah Chen • 4 weeks ago

Why a 400-Student Intro Course Adopted Layered Code Checks

One CS department's switch from a single similarity tool to a layered detection workflow changed what they could see in student code. Peer copying, web sources, and AI-generated submissions each required different signals, and combining them revealed more than any one check alone.

AI Code Detector Comparison Across Codequiry, GPTZero, and Copyleaks General 10 min
Dr. Sarah Chen Dr. Sarah Chen • 1 month ago

AI Code Detector Comparison Across Codequiry, GPTZero, and Copyleaks

A CS professor ran 1,200 Java submissions through three AI code detectors. Codequiry caught 94% of known AI files and flagged only 3.5% of pre-LLM human code, while the other tools posted two to three times that false positive rate. The full numbers and methods are below.

How a CS Professor Spots Refactored Code Plagiarism in Java Labs General 11 min
Dr. Sarah Chen Dr. Sarah Chen • 1 month ago

How a CS Professor Spots Refactored Code Plagiarism in Java Labs

This is the workflow Dr. Sarah Chen uses every week to catch plagiarism that survives renaming, reordering, and refactoring. It layers token normalization, AST comparison, and a smart review queue, then walks through what to look for in a diff before talking to a student.

What Happens When a CS Course Runs Both MOSS and ChatGPT Detectors General 11 min
Dr. Sarah Chen Dr. Sarah Chen • 1 month ago

What Happens When a CS Course Runs Both MOSS and ChatGPT Detectors

Over 1,200 student submissions from a large public university’s introductory Python course were analyzed with Codequiry’s similarity engine and its AI code detector. The results show how traditional plagiarism tools miss a growing fraction of unauthorized work—and why layering AI detection changes what instructors actually see.

Writing Programming Assignments That Resist Plagiarism General 7 min
Dr. Sarah Chen Dr. Sarah Chen • 1 month ago

Writing Programming Assignments That Resist Plagiarism

We cut similarity rates from 43% to 7% in a Data Structures course not by policing harder but by rewriting the assignments themselves. Here's what worked, what broke, and where detection tools like Codequiry still earn their keep.

How AST-Based Similarity Catches Disguised Code Plagiarism General 15 min
Dr. Sarah Chen Dr. Sarah Chen • 2 months ago

How AST-Based Similarity Catches Disguised Code Plagiarism

Token-based plagiarism detectors match sequences of tokens, but smart students can evade them by renaming variables, reordering statements, and refactoring code. Abstract Syntax Tree (AST) comparison digs deeper into the structural DNA of a program, making it far harder to disguise copied code. Learn how AST-based detection works, why it catches what MOSS and JPlag miss, and where Codequiry’s multi-layered approach fits in.

Why CS Departments Are Now Running AI Checks After Plagiarism General 12 min
Dr. Sarah Chen Dr. Sarah Chen • 2 months ago

Why CS Departments Are Now Running AI Checks After Plagiarism

When Midwestern State University’s CS department discovered that MOSS alone missed nearly a third of suspicious submissions—many generated by ChatGPT—they implemented a two-stage detection pipeline. This is what they learned about running plagiarism checks first, then AI detection, and why the combination caught more than either tool alone.

Designing Plagiarism-Resistant Programming Assignments General 11 min
Dr. Sarah Chen Dr. Sarah Chen • 2 months ago

Designing Plagiarism-Resistant Programming Assignments

Copy-paste code and AI-generated solutions are flooding CS courses. Smart assignment design—parameterized prompts, incremental checkpoints, and open-ended specs—can make copying so cumbersome that students default to doing their own work. This deep dive builds a concrete, reusable playbook for instructors.