River
Y CombinatorBacked by Y Combinator

Education & TrainingFree

Assessment Design to Reduce Cheating

Send the assessment and the outcomes it targets, and get redesign options that keep the outcome, each priced in marking hours for your real cohort.

Start here

Almost everything written about assessment and cheating is about detection. Run the submission through a checker, watch for a tonal shift, compare against the student's earlier writing, look for citations that do not exist. Detection is reactive, it produces a probability rather than evidence, and it fails in a way that costs you more than the cheating does. Design is the durable response, and the reason design advice gets ignored is that it never says what the alternative costs to mark.

SOC 210 Research Methods is an invented example. 184 students, a 2,000-word take-home essay on a single shared prompt, marked in about 22 minutes each. That is 4,048 minutes, or 67.47 hours of marking, and it assesses all four course outcomes in one task. Every one of those four outcomes is outsourceable as the task is currently written, because nothing in it is specific to the student and nothing requires them to account for their own work.

Now price the alternatives for 184 students. A staged process portfolio, three formative submissions at five minutes each plus an eighteen-minute final, is 33 minutes a student, so 101.2 hours. That is exactly 50.0 percent more marking than the essay it replaces, which is why the most commonly recommended redesign is the one departments quietly abandon. A supervised in-class write-up of a pre-prepared analysis is 18 minutes plus 4.5 hours of invigilation, which is 59.7 hours, or 11.5 percent less than the status quo.

The option table has to carry hours, or it is not a decision

The result that changes the conversation is the oral. Ten minutes of oral defence per student plus a lighter twelve-minute read of the written work is 22 minutes a head, the same 22 minutes the essay already costs. The total is identical at 67.47 hours. What differs is the shape: 30.67 hours of it is timetabled contact that has to be scheduled and staffed, and 36.8 hours is marking that can be done at midnight. Same hours, very different feasibility.

Individualised task data is the other option worth costing properly. Give every student a different dataset from a bank and the analysis outcome stops being outsourceable, because the answer is specific to them. Marking drops to 20 minutes because the key is parameterised, so 61.33 hours, but building the bank is about 14 hours once. First year 75.33 hours, every year after 61.33. That front-loaded cost is invisible in every recommendation of this technique and is the reason it stalls.

None of this makes detection the fallback. Detectors misfire unevenly: one evaluation of several widely used GPT detectors found they consistently misclassify non-native English writing samples as AI-generated, whereas native writing samples are accurately identified. At a one percent false positive rate, 184 submissions produce roughly two wrongly accused students. Meanwhile accreditors already require institutions to establish that a student who registers in any course offered via distance education is the same student who academically engages in the course.

How it works

  1. Send the assessment

    The task as written, the outcomes it claims to measure, and your cohort and staffing.

  2. Test the exposure

    Each outcome checked against whether the task requires anything specific to the student.

  3. Generate the options

    Redesigns that keep each outcome measurable, with the mechanism that protects it named.

  4. Price each one

    Marking minutes times your cohort, contact time split out, and build cost shown once.

What you get

  • Each target outcome tested for whether the current task can be completed by somebody else
  • Redesign options that preserve the outcome, rather than generic advice to make assessment authentic
  • Every option priced in marking minutes multiplied by your real cohort, reported in hours
  • Contact time separated from marking time, because one has to be timetabled and one does not
  • One-time build costs shown separately from recurring cost, so year one and year two both appear
  • Which outcomes each option protects and which it leaves exposed, stated per option rather than in general

Common questions

Why not just use a detection tool?

Because a detector output is a probability, an integrity case needs evidence, and the errors are not evenly distributed. At a one percent false positive rate a 184-student cohort produces about two wrongly flagged students. The cost of that, to those students and to the department, is higher than the cost of the cheating it caught.

Is the marking cost really the blocker?

It is the usual one. A process portfolio for 184 students is 101.2 hours against the essay's 67.47, which is 50.0 percent more work for the same credit weighting. No amount of agreement about pedagogy survives that on a real timetable, which is why the hours belong in the comparison from the start.

How do you decide whether an option still measures the outcome?

Outcome by outcome, and it is reported rather than scored. A supervised write-up protects written communication and leaves question formulation exposed, because that part was prepared in advance. Individualised data protects the analysis outcome specifically. No single option protects all four, and saying so honestly is the point.

Does this work for an online or distance cohort?

It has to, and the constraint is different. Accreditors require institutions to establish that the registered student is the one engaging in the course, so identity is part of the design rather than an add-on. Options that depend on a bookable room get flagged as unavailable for the distance portion of the cohort.

Where do the marking time estimates come from?

Yours, wherever you can supply them, because departmental norms vary enormously. Where you cannot, the analysis uses a stated assumption and shows the sensitivity, so a two-minute error per student is visible as roughly six hours across a 184-person cohort rather than buried.

How does this relate to writing the rubric?

The rubric comes after the design decision. The assessment and rubric pack is where criteria and level descriptors get written, and a redesigned task usually needs a new rubric rather than the old one reworded. This tool stops at the design choice and its cost.

What if the problem is a question bank rather than an essay?

Different mechanism, same logic. The exam and question bank pack handles item construction and bank hygiene, and item analysis review tells you which items are actually discriminating. Item exposure is a bank management problem before it is a design problem.

Assessment Design to Reduce Cheating

Fill in the form and your workspace opens with the work already underway.