General
AI-Resilient Assessment: A Complete Guide for Universities

General

Imagine a provost discovering the moment the old assessment model had broken. Her faculty had spent three months redesigning a capstone course, replacing a final exam with a complex case analysis. The first cohort submitted their work. The submissions were polished, well-structured, and analytically sound. Every single one of them had been written by ChatGPT. The students had learned nothing. The grades reported exactly the wrong signal. And the faculty were left with an impossible choice: spend the semester policing AI use or admit that the assessment itself was the problem.
This scenario is playing out across higher education. The instinct is to reach for detection tools, policy updates, and honor code revisions. But the data tells a different story. A study published in the International Journal for Educational Integrity found that 50.9% of students have used or considered using generative AI for academic purposes, with rates climbing every semester. Meanwhile, reading compliance research by Burchfield and Sappington, documented across multiple institutions, confirms what instructors already feel: most students do not complete assigned readings, and the trend has been declining for decades. The traditional model was already fragile. AI simply exposed the cracks.
The alternative is not a better detector. It is a better assessment framework -- one that makes the learner's reasoning, decisions, practice, and reflection observable, so AI can support the experience without replacing the cognitive work the assessment is meant to measure.

An AI-resilient assessment is one that preserves its validity even when students have access to generative AI tools. It does not rely on the assumption that the student produced the work in isolation. Instead, it makes the process of thinking visible, verifiable, and resistant to outsourcing. The concept draws on a growing body of research and practice. The Ivey Business School white paper "From Chalkboard to Chatbot" (August 2026) provides one of the clearest frameworks. It argues that generative AI can remove the struggle through which durable learning is produced -- what the authors call productive friction (PDF pp. 2-3). An AI-resilient assessment does not try to eliminate AI. It restores the friction that makes learning stick.
A common misconception is that AI-resilient assessment requires banning AI tools. It does not. In fact, many of the most effective AI-resilient designs integrate AI deliberately -- as a research assistant, a feedback generator, a role-play partner, or a coaching tool. The key distinction is whether the AI is doing the thinking for the student or supporting the student's own thinking.
This distinction matters because blanket bans are both unenforceable and pedagogically counterproductive. Students will graduate into a professional world where AI is embedded in every tool they use. The goal of assessment should not be to simulate a pre-AI environment. It should be to verify that the student can do the cognitive work that AI cannot replace: framing problems, weighing trade-offs, making judgments under uncertainty, and defending decisions.
Detection-based approaches try to prevent AI use. AI-resilient assessment tries to preserve learning. These are fundamentally different objectives.
Detection asks: "Did the student use AI?" AI-resilient design asks: "Did the student learn what the assessment was designed to measure?" The first question leads to an arms race of detectors and evasions. The second leads to a redesign of the assessment itself.

For decades, the essay was the gold standard of higher-order assessment. A well-written essay required the student to research, synthesize, argue, and conclude. The format demanded cognitive effort, and the quality of the output was a reasonable proxy for the quality of the thinking.
That equation no longer holds. A student can paste a prompt into ChatGPT and receive a passable 2,000-word analysis in seconds. The output is polished. The structure is sound. The citations are plausible. But the student has performed none of the cognitive work the essay was designed to measure. The grade reports a signal that is now meaningless.
This is not a marginal problem. A 2024 study from the University of Reading found that AI-generated exam submissions went undetected 94% of the time and consistently scored higher than real student submissions. The detection gap is not closing -- it is widening.
The most insidious effect of AI on assessment is not cheating. It is the illusion of learning.
When a student uses AI to complete an assignment, they often receive a good grade. The grade creates the impression that learning has occurred. But the student has not built the underlying mental models, the reasoning pathways, or the judgment that the grade is supposed to represent. They have demonstrated the ability to prompt effectively, not the ability to analyze, evaluate, or create.
This distinction between performance gains (doing well on the assessment) and learning gains (building durable knowledge and skill) is the central challenge of assessment in the AI era. When AI can produce the performance without the learning, the assessment is no longer fit for purpose.
The market for AI detection tools has exploded. Turnitin's AI detection, GPTZero, Originality.ai -- universities are spending heavily on the promise that technology can catch technology. The data tells a different story. A peer-reviewed study in Springer found that AI detectors exhibit systematic inconsistencies. Research in Cell Press's Patterns journal confirmed that detectors are biased against non-native English writers, flagging their original work as AI-generated at disproportionately high rates. At least a dozen major universities, including Yale, Vanderbilt, and Northwestern, have banned or discouraged their use.
The deeper problem is structural. Detection assumes a cat-and-mouse game where the institution stays one step ahead. But LLMs improve faster than detection models, and students have every incentive to stay current on evasion techniques. The arms race resets every semester, and the institution loses every time.
The Ivey report maps learning tools across the learning journey and three outcomes: conceptual knowledge, specific skills, and judgment (PDF p. 5, Tables 2-3). Four differentiators emerge across the tools: cognitive challenge, faculty oversight, resilience against generic AI use, and volume and variability of practice (PDF p. 5). These differentiators translate into five design principles.
The most AI-resistant assessments do not ask students to produce a piece of text that an LLM can generate. They ask students to make decisions that require judgment, context, and trade-offs. A decision point forces the student to commit to a path, justify it, and live with the consequences. There is no single correct answer to copy-paste. The assessment captures the reasoning process, not just the output.
Generic prompts are vulnerable because they can be answered by generic AI. Assessments that embed students in a rich, specific context -- with stakeholders, constraints, incomplete information, and competing priorities -- are far harder to outsource. Navigating the specificity of the context is substantially harder for AI and more observable when attempted -- the context only exists inside the assessment.
When the only thing submitted is a final answer, AI can produce it. When the assessment requires students to show their reasoning -- through decision logs, annotated choices, or step-by-step justifications -- the cognitive work becomes visible and verifiable. The reasoning trail is what gets graded, not the polished conclusion.
One-shot assessments are easy to outsource because the student has one opportunity to produce output. Multi-stage assessments that require students to reflect on feedback, revise their work, and explain what changed and why create a process that is much harder for AI to replicate. The AI can generate a first draft. Explaining why the student chose to revise it requires a personal reasoning trail that is harder to fabricate.
The most productive role for AI in assessment is not as a substitute for thinking but as a coach that supports it. When AI provides feedback on a student's reasoning, challenges their assumptions, or suggests alternative perspectives, it strengthens the learning process rather than bypassing it. The key is that the AI operates on the student's thinking, not instead of it.
One of the most important insights from the Ivey report is that different learning outcomes require different assessment formats. No single format should be expected to measure all three equally well.
What a student knows about a topic. This is the domain of declarative knowledge: frameworks, theories, definitions, and models. Traditional formats like essays, exams, and multiple-choice tests can still assess conceptual knowledge effectively, provided the assessment is designed to test understanding rather than recall. The vulnerability is that AI can generate plausible-sounding explanations of concepts without genuine understanding. The fix is to require students to apply concepts to novel situations, not just recite them.
What a student can do with their knowledge. This is the domain of procedural knowledge: analysis, calculation, diagnosis, and execution. Skills-based assessments are more AI-resistant than knowledge-based assessments because they require the student to perform a task, not just describe it. Simulations, case analyses, and hands-on exercises are natural formats for this domain.
What a student chooses when the answer is not clear. This is the hardest domain to assess and the most valuable. Judgment requires the student to weigh incomplete information, navigate competing priorities, and make a decision with consequences. It is also the most AI-resistant domain, because judgment is context-dependent and personal. No LLM can replicate the specific reasoning that led a particular student to a particular decision in a specific scenario.
A common mistake is to treat a single assessment format -- typically the essay -- as a universal measure of all three outcomes. Essays can assess conceptual knowledge. They are weaker at assessing practical skills. They are almost useless at assessing judgment under uncertainty, because the student writes about what they would do rather than actually doing it.
The implication is clear: assessment design should start with the learning outcome, not the format. If the outcome is judgment, the format should be decision-based. If the outcome is skill, the format should be performance-based. If the outcome is knowledge, the format can be more traditional, but should still require application rather than recall.
Instead of: "Analyze the strategic options facing Company X and recommend a course of action."
Try: A branching case where students face three consecutive decisions, each building on the previous one. After each decision, they write a short rationale explaining what they chose, what information they prioritized, and what they would need to know to be more confident. The AI can generate the analysis. It cannot generate the sequence of decisions that the student actually made, because the sequence depends on the student's prior choices.
Instead of: "Read this 20-page case study and prepare a discussion."
Try: A simulation where students receive messages from virtual stakeholders over time. The first message presents a problem. The second reveals a complication. The third introduces a stakeholder with conflicting interests. Students must decide who to talk to, what to ask, and when to act. Information is disclosed only when the student asks the right question. The AI can summarize the case. Staging the stakeholder interactions makes outsourcing harder: AI can simulate a stakeholder, but the simulation makes the student's engagement observable rather than invisible.
Instead of: One negotiation role-play with a classmate.
Try: A series of AI-powered role-plays where students negotiate with virtual counterparts who have different personalities, constraints, and information. Each repetition builds on the previous one. The student's strategy must adapt to the counterpart's behavior. The AI can generate a script. Adapting convincingly to the student's specific moves across multiple rounds is harder to outsource -- the effort required to sustain the illusion across rounds becomes its own form of engagement.
Instead of: "Submit your final recommendation."
Try: A three-stage process: submit an initial recommendation with reasoning, receive AI-generated feedback that challenges assumptions, then submit a revised recommendation explaining what changed and why. The grade is based on the quality of the revision, not the initial answer. The AI can generate a plausible first answer. Explaining why the student chose to revise it requires a personal reasoning trail that is harder to fabricate.
Not every assessment needs to be AI-resilient through redesign. Some formats are structurally AI-resistant. Closed-book, in-person exams where students write by hand remain effective for assessing foundational knowledge. Live oral defenses, where students must defend their reasoning in real time to a faculty member, are among the most AI-resistant formats available. The cost is time and scale. The benefit is near-certainty that the work is the student's own.
The Ivey report recommends that faculty experiments should identify the learning activity at risk of being shortcut, the friction being restored or introduced, and the evidence that would demonstrate learning beyond satisfaction (PDF p. 7). The following checklist operationalizes that recommendation.
Start with the outcome, not the format. What should students be able to do after this assessment that they could not do before? Be specific. "Demonstrate critical thinking" is too vague. "Evaluate a proposed merger under conditions of incomplete information" is specific enough to design around.
Map the assessment task against what current AI tools can do. Can a student paste the prompt into ChatGPT and receive a passable answer in under a minute? If yes, the task is vulnerable. The question is not whether students will use AI -- they will. The question is whether the assessment can still measure learning when they do.
Once you have identified the shortcut, design the friction that restores the cognitive work. This might mean adding a decision point, requiring a reasoning trace, introducing stakeholder interaction, or staging information over time. The friction should be proportional to the learning objective. Not every assessment needs to be a multi-stage simulation. But every assessment should have at least one point where the student must commit to a position that AI cannot pre-generate.
What will you grade? If the answer is "the final product," the assessment is likely still vulnerable. Shift the evidence toward the process: the decision trail, the reasoning log, the revision history, the reflection. These artifacts are harder to fake and more informative about what the student actually learned.
Be explicit about what AI use is permitted and what is not. The Ivey report recommends preserving AI-free assessment or social-learning spaces where appropriate (PDF pp. 7-8). Some assessments should be AI-free. Others should be AI-supported. The key is that the boundary is transparent, justified, and tied to the learning objective.
The first version of any redesigned assessment will not be perfect. Pilot it with one section. Debrief with students. What did they learn? What did they find confusing? Did the assessment actually measure what it was designed to measure? Use the data to iterate. The Ivey report explicitly notes that its review is a review of 14 tools, not a universal ranking -- its judgments are subjective (PDF p. 9). The same humility should apply to any assessment redesign.
Simulations are one of the most effective formats for AI-resilient assessment because they are built around structured decisions under conditions of incomplete information. A well-designed simulation presents students with a scenario, introduces stakeholders with competing priorities, and requires them to make decisions with consequences. The simulation does not ask students to write about what they would do. It asks them to do it.
This is fundamentally different from a case study, which presents information retrospectively and asks students to analyze what happened. A simulation presents information episodically and asks students to shape what happens next. The distinction is the difference between analysis and judgment, and it is the reason simulations are structurally resistant to AI outsourcing.
Learning that transfers to real-world performance requires three conditions: the learner must make a choice, experience a consequence, and receive feedback in a timeframe where the connection is still live. Simulations deliver all three. The choice is the student's decision. The consequence is the scenario's response. The feedback is immediate and specific.
This is why flight simulators are the gold standard in aviation training and why case-based simulations have been used in business education for decades. The medium forces the learner to act, to commit to a path, and to live with the outcome. No amount of slide content can replicate what happens in the moment when someone has to choose with incomplete information under time pressure.
One of the most underappreciated advantages of simulation-based assessment is the data it generates. Decisions, hesitations, and pivots are logged where the simulation captures them. Faculty can see not just what students concluded, but indicators of how they reasoned their way to that conclusion. This visibility transforms the debrief from a general discussion into a targeted intervention. Instead of asking "What did everyone think?" faculty can ask "Why did 60% of you choose option B at the third decision point, and what information would have changed your mind?"
Platforms like LiveCase provide an environment for scenario-based, decision-rich learning. Students interact with AI-powered characters, receive partial information, make time-pressured decisions, and see consequences unfold. Faculty can author their own simulations or choose from a catalog of ready-to-run scenarios, including bestsellers published through Harvard Business Impact.
But LiveCase is one implementation option, not the definition of the category. The principles of AI-resilient assessment apply whether you use a simulation platform, a role-play exercise, a staged case analysis, or a live oral defense. The format should serve the learning objective, not the other way around.
For a deeper look at how decision-based assessments address AI cheating, see our earlier article on why decision-based assessments beat AI cheating. For a practical guide to measuring critical thinking beyond the essay, see how to measure critical thinking. And for an overview of how AI simulation tools fit into higher education assessment, read why AI simulation tools are the missing piece in higher education assessment.
An AI-resilient assessment is one that preserves its validity even when students have access to generative AI. It measures the learner's reasoning, decisions, and reflection rather than the polished output that an LLM can generate.
No. No assessment format is completely AI-proof, any more than any lock is completely pick-proof. AI-resilient means the assessment maintains its integrity despite AI use, because the cognitive work it measures cannot be outsourced to a language model.
It depends on the assessment design. Some AI-resilient assessments intentionally integrate AI as a coaching or research tool. Others are designed to be AI-free. The key is that the boundary is transparent, justified, and tied to the learning objective.
Shift the evidence from the final product to the process. Require decision logs, annotated choices, reasoning traces, and staged revisions. Grade the reasoning trail, not the polished conclusion.
The Ivey report recommends preserving AI-free assessment or social-learning spaces where appropriate (PDF pp. 7-8). Foundational knowledge checks, live oral defenses, and in-person collaborative exercises are natural candidates for AI-free assessment.
By observing the decisions a student makes when the answer is not clear. This requires a scenario with incomplete information, competing priorities, and consequences. The student's decision trail reveals their judgment more accurately than any essay could.
Start with one course, one module, one assessment. Identify the learning outcome most vulnerable to AI outsourcing. Redesign that single assessment using the principles above. Measure the results. Use the data to build the case for broader adoption.
Want to talk this through? Book a free 15-minute discovery call
Build your own simulation, bring in our Studio, start with a published case, or talk through your idea.
Ask about a demo, quote, or your simulation. We usually answer within a few hours.
Contact usCreate with AI or start from scratch. You keep full creative control.
Create an accountWork with our case writers on a classroom-ready simulation.
Explore StudioBrowse simulations authored with leading educators.
Browse catalogue
Author: Denis Duvauchelle
Elevate your AI skills for better learning 🌟 | AI Developer & Education Innovator | 50K + Executives / HigherEd success stories. He specializes in both research and implementation, and is dedicated to creating the best possible experience for educational simulations, both in terms of design and usage. With a focus on driving engagement and learning outcomes, Denis is committed to delivering innovative and impactful solutions for his clients. https://www.linkedin.com/in/desduvauchelle/
Published: 9/3/2026
Related posts








© 2021 — @Livecase