AI Compromised Traditional Assessment. Enter, Oral Exams—With AI.

Generative AI is making traditional assessments increasingly unreliable. An 800 year old solution is available ... Oral Exams.

Share
AI Compromised Traditional Assessment.  Enter, Oral Exams—With AI.
audio-thumbnail
Prefer to listen? Play the audio version of this article below.
0:00
/772.571625

Generative AI is disrupting … everything. It is changing education faster than our assessments can adapt, both positively and negatively. In my own classes, I’ve watched my take-home final exam averages rise steadily over the past three years. I would like to think it is solely because of improved pedagogy and delivery; however, the magnitude and timing of the increase are strongly correlated with the onset of generative AI. When reasoning can be outsourced to a model, we can lose the signal that exams and other assessments have been trusted to provide.

As educators, while we need to detect academic dishonesty, I believe we are much more interested in preventing it in the first place and amplifying genuine learning. That’s why I support bringing back one of the most effective assessments in history, the oral exam … and using AI to make it scalable.

Generative AI Is Compromising Assessment ... and Learning

In just the last few years, generative AI has improved dramatically. In 2024, my colleague Dr. Monnie McGee and I compared ChatGPT 3.5 and GPT-4 on statistics exam questions. At the time, one concern was equity: the paid model achieved about 82% accuracy on multiple-choice and free-response questions, while the free model was effective only about 50% of the time. (McGee & Sadler, 2025,2024)

Today, that gap is far less meaningful. Widely available free AI systems—including ChatGPT, Gemini, Grok, DeepSeek, and Claude—can answer many of these same questions with very high accuracy.

That progress holds enormous promise for improving and democratizing education. But it also threatens the reliability of traditional assessment tools—quizzes, exams, and papers—especially in online and flipped-classroom settings. And that threat is already being realized. A Turnitin survey published in April 2025 reported a striking, though perhaps unsurprising, result:

Another article from the AP (2025) featured a teacher who is quoted as saying,

“Teachers around the country say that student use of artificial intelligence has become so prevalent that to assign writing outside of the classroom is like asking students to cheat.”

Additionally, as I was writing this article, this literally popped up on my LinkedIn from a colleague of mine:

“One of my amazing faculty asked me how she should address rampant AI use in her (intro) Creative Coding course.”

The examples are ubiquitous, but the part that worries me even more than “cheating” is what happens when assessment stops reflecting understanding.  This enables students to move forward with gaps in their understanding that are likely to resurface in advanced courses, internships, and careers.

What I Saw in My Own Classroom

I can see strong evidence of this shift in my own teaching. Over the past three years, the average scores on my take-home midterm and final have increased swiftly and steadily.

I have taught our Foundations of Statistics course every semester since 2014 and have included the results from the final exam since 2019. We have three semesters a year and offer this course in each semester; the results in the plot are the averages of all the students that took the course that year (~60 students per year for a total of ~ 420 students).

The glaring point: final exam averages were consistently between 72 and 75 from 2019 to 2022 and made a significant jump in 2023 after the introduction of ChatGPT 3.0 in Nov 2022. While possible, this is unlikely to be a coincidence as responses tended to

  1. reference vocabulary, symbols and methods not mentioned in the course
  2. be in a format consistent with generated responses and inconsistent with student’s pervious work (use of bolded words, bullet points, grey dividing lines etc.)
  3. often include answers to questions that were not asked in the problem and often in detail and length that would have been extremely unlikely to be completed in the 3 hours allotted.

A Solution: The Oral Exam (The Viva Voce)

Long before bluebooks and take-home essays, universities assessed students orally. In the medieval period, a major form of summative assessment was the disputation (or disputatio): a live, structured exchange in which a student defended an idea, fielded objections and follow-up questions, and demonstrated understanding in real time. That tradition is a clear ancestor of modern oral examinations—often described as viva voce (“with the living voice”)—and it remains one of the most authentic ways to assess genuine understanding in an age where traditional written assessments can be compromised by generative AI.

And … academia is taking notice. In December 2025, The Washington Post (2025) profiled professors at the University of Wyoming, Vanderbilt, Illinois State, and UC San Diego who are reviving oral exams specifically to address the concerns about assessment raised by ChatGPT-style tools. As academic integrity researcher Tricia Bertram Gallant put it, oral assessments are “definitely experiencing a renaissance,” and University of Wyoming professor Andrew Dobson said he never imagined oral exams would be “dusted off and gain a second life.”

Putting the Oral Exam Into Practice

The first oral exam in DS 6371 was administered in the Summer of 2023 and consisted of two questions which spanned the student learning outcomes of the course including significance tests, experimental design, and linear regression. Students were asked an initial question and then, based on the student’s response, one or more follow-up questions were posed and the responses were graded based on the following rubric:

The exam was graded as pass/fail and a “pass” was required to pass the course.

What Students Thought

A lot was learned that first semester. Students almost uniformly reported high anxiety before the exam although also uniformly reported, after taking the exam, that they felt it was extremely useful, accurate and that they were glad they took the exam.

In order to reduce negative anxiety, students were given the bank of 20 questions from which the two questions would be drawn. Additionally, the students were given more control by allowing them to select one of the questions while the professor selected the other.  A comment from one of the students reflected the effectiveness of making the questions available ahead of time:

“Even though, I was anxious before exam I appreciated getting to know what questions were going to be used…”

Effects of this change and a reflection of the current pre-viva student state of mind are available from a nonscientific survey given to all 45 students (18 responses thus far) from the last three  DS 6371 cohorts. This survey revealed that this overwhelming negative anxiety had largely dissipated and given way to a neutral and possibly slightly positive feeling of the upcoming exam (average score =3.1)

Student-reported emotions before the oral exam.

Additional evidence of effectiveness and reduced anxiety was reflected in general feelings of the oral exam after the student had completed the exam.  Results below indicate that posttest, students largely moved from an ambiguous three into one of two groups (below).  The overall average score now reflected a strong positive signal with the average emotion score increasing to a 4.0 and exactly half the students reporting an “extremely positive” emotion about the exam. 

Student-reported emotions after the oral exam.

Did It Actually Work?

The early evidence is encouraging. In addition to the reduction in anxiety and the positive shift in sentiment, students reported that the oral exam accurately reflected their understanding and supported learning. Because this is a small, self-reported sample rather than a formal validation study, I view these results as intriguing signals rather than definitive proof.

Accuracy and Face Validity

Student perceptions provide meaningful evidence about the assessment’s face validity and whether it elicited the kinds of thinking and understanding it was designed to measure. In our survey, 83% of students agreed or strongly agreed that “The oral exam accurately assessed my understanding.” That’s an important signal for a high-stakes assessment: students experienced the oral exam as tracking what they knew (and didn’t know) in real time and in a way that traditional written exams can struggle to capture. 

83% of respondents agreed or strongly agreed that the oral exam accurately assessed their understanding.

Identifying Gaps in Knowledge

The strongest learning signal in the survey was metacognitive: 100% of respondents agreed or strongly agreed that “The oral exam helped me identify gaps in my knowledge.” Since the oral exam is customized for each student, follow-up questions based on student responses allow for direct and efficient inquiry into areas where students have answered incorrectly or seem less confident.

100% of respondents agreed or strongly agreed that the oral exam helped them identify gaps in their knowledge.

Deeper Learning

Finally, students reported that the oral exam didn’t just measure learning—it amplified learning. 94% agreed or strongly agreed that “The oral exam encouraged deeper learning than a traditional written exam would have.” That makes sense: preparing for an oral exam tends to shift studying away from memorizing steps and toward building coherent explanations, practicing reasoning out loud, and connecting ideas. 

94% of respondents agreed or strongly agreed that the oral exam encouraged deeper learning than a traditional written exam would have.

A particularly interesting possibility is that oral assessment changes behavior long before the exam itself. Homework, reading, projects, and practice become preparation for a future moment when students will have to explain what they know in their own words.

Students can still use generative AI—and I want them to. AI can be a revolutionary on-demand teaching assistant and collaborator. But if students know they will eventually need to explain the material themselves, they have a stronger incentive to remain active participants in the learning process rather than outsourcing the entire task.

The goal is not to stop students from using AI. It is to help them use AI as a collaborator rather than a crutch.

Three Strong Signals from Students 

Summary of the three strongest student-reported outcomes. These early survey results align with three years (nine semesters) of oral-exam experience and student feedback.

While the sample size is quite small, they are very much in line with my experience with oral exam results and students’ comments about them over the past 3 years (9 semesters).  The bottom line is that these early results suggest the oral exam functions as a more accurate and authentic assessment in the age of generative AI, one that improves and amplifies student learning outcomes and does so in a way that translates directly to interviews, stakeholder communication, and real applied work. 

The Problem We Haven’t Solved Yet: Scale

There is one enormous limitation: oral exams take time. They are difficult to administer even in smaller classes and become practically impossible at scale.

That challenge deserves its own discussion. In the next article, I will explore the scalability problem—and how AI can help make oral assessment more practical while giving students opportunities to practice, receive personalized feedback, and prepare before the official oral exam.

Where We Go From Here

We shifted to oral exams not to “catch cheaters,” but to amplify learning. When students know they will have to explain ideas in their own words, in real time, assessment stops being a game of AI output and becomes a meaningful measure of genuine understanding.

In my experience, the oral exam does not just measure mastery; it strengthens it. It helps students surface gaps, build confidence communicating what they know, and develop durable knowledge that carries into advanced courses, internships, interviews, and careers.

Ultimately, the goal is not enforcement. It is to help students learn deeply, fall in love with learning and with their own minds, and move closer to the goals and dreams that brought them here in the first place.


Disclosure: I am the founder and CEO of Elenkis, an AI-powered oral assessment platform focused on helping educators assess student understanding through interactive conversation.

About the Author

Dr. Bivin Sadler, is a clinical associate professor of statistics and data science at Southern Methodist University (SMU) and the founder and CEO of Elenkis. His work focuses on assessment, learning, oral exams, and the use of AI in education.

References

  • Turnitin — “Turnitin Releases New Findings and Key Insights from Focused Survey on Impact of AI in Education.”
    Turnitin survey release
  • Associated Press — “The Rise of AI Tools Forces Schools to Reconsider What Counts as Cheating.”
    AP News
  • The Washington Post — “Professors Are Turning to This Old-School Method to Stop AI Use on Exams.”
    Washington Post
  • McGee & Sadler — “Generative AI Takes a Statistics Exam…”
    Journal of Data Science
  • McGee & Sadler — “Equity in the Use of ChatGPT for the Classroom…”
    arXiv