Free toolsPricingMobile app
Log in
Get started freeGet started free
Ask Bo
  • Ask Bo anythingAnswers from your own lectures, cited
  • AI FlashcardsMake me a deck for chapter 4
  • Image OcclusionMask the labels on a diagram
  • Practice examsBuild a 20-question mock
  • Mind mapsShow how these ideas connect
  • Study guidesSum up the whole unit
  • AI SummarySum up Friday's lecture
  • AI QuizQuiz me on chapter 4
  • Cheat sheetsOne page for the final
Get started freeGet started freeLog in
  • Ask Bo anything
  • AI Flashcards
  • Image Occlusion
  • Practice exams
  • Mind maps
  • Study guides
  • AI Summary
  • AI Quiz
  • Cheat sheets
Free toolsPricingMobile app
  • How we test
  • We do not grade our own homework
  • What we found
  • How we actually decide
  • Why we run this again
  • What this means for you
Back to blog
10 Aug 20265 min read

Which AI Actually Writes Your Study Material? We Tested Five and Picked on Quality

We ran GPT-5.6, Gemini 3.1 Pro, Grok 4.5, Kimi K3 and DeepSeek V4 Pro on the same real lecture material, then had them graded blind. Here is what we found and what we run.

  • LHLuis Henrich-Bandis
Loading...
Ask Bo anything

Upload your lectures and notes. Bo reads them and answers from your own course material.

Practice exams from your course

Bo writes a full practice paper from the topics and lectures you choose.

Flashcards in seconds

Bo turns your slides and readings into ready-to-study flashcard decks.

Mind maps and summaries

See how your concepts connect, or get a clean summary of any lecture.

Your course, not the internet.

Features

  • Ask Bo
  • AI Flashcards
  • Image Occlusion
  • AI Exams
  • Mind Maps
  • Study Guides
  • AI Summary
  • AI Quiz
  • Cheat Sheets

Free tools

  • Flashcard Generator
  • Quiz Generator
  • Mind Map Generator
  • Study Guide Generator
  • PDF Summarizer
  • Slide flashcards

Compare

  • vs ChatGPT
  • vs Quizlet
  • vs Anki
  • vs YouLearn
  • All comparisons

Resources

  • Glossary
  • Answers
  • How it works
  • Why StudyPDF
  • Use cases

Company

  • Pricing
  • FAQ
  • Mission
  • Creator program
  • Enterprise
  • Contact
  • Changelog

Legal

  • Privacy
  • Terms
  • Imprint
© 2026 StudyPDFFree to start. No card required.

When an AI tool makes your study material, you are trusting a model you never chose and usually cannot see. So it is fair to ask which one we use, and how we decided.

The short answer: we currently run GPT-5.6 on its most capable tier for writing flashcards, quizzes and exams. The longer answer is how we got there, because we do not pick by price, by benchmark leaderboards, or by whichever model was in the news that month. We test them on your kind of material.

Here is the most recent round.

How we test

The hard part of comparing models is that almost nothing stays still. Different retrieval, different source text, different output. So we hold everything fixed except the model.

We took three real documents from real courses, deliberately different in shape:

  • A dense accounting glossary, wall to wall definitions
  • A sparse psychology slide deck, a few bullet points per slide
  • Long form marketing notes, actual paragraphs

Then we froze the source material and the instructions, and changed only the model. Thirty runs, one at a time, generating flashcards and quizzes from each document with the exact prompts the product uses.

Five models, from five different labs: GPT-5.6, Gemini 3.1 Pro, Grok 4.5, Kimi K3 and DeepSeek V4 Pro.

We do not grade our own homework

This is where model comparisons usually go wrong. If you score the outputs yourself, you find what you expected to find. If you let one AI score them, you have measured which output most resembles that AI's own writing style.

So the grading went to a panel of three judges from three different companies. Each judge saw all five sets side by side, with the labels shuffled, so there was no way to tell which model wrote which. They scored whether every card was genuinely supported by the source, whether the set covered the whole topic, whether the questions were clear and non repetitive, whether the difficulty matched what was asked for, and whether a real student revising would get value out of it.

Three of the judges were also contestants, which is a fair thing to be suspicious about. So we checked it: we compared the rank each judge gave its own output against the rank the other judges gave that same output. Nobody inflated themselves. DeepSeek actually graded its own set more harshly than the others did.

What we found

The result was not a landslide. Four of the five models finished in a tight group, close enough that we would not claim a winner between them on this evidence.

Grounding, which is the one that matters most: GPT-5.6 scored highest at 4.83 out of 5, with Gemini 3.1 Pro at 4.78 and Grok 4.5 and Kimi K3 both at 4.67. Grounding means every card is actually supported by your document rather than invented. A confidently wrong flashcard is worse than no flashcard, so this is the number we refuse to trade away.

Question quality: GPT-5.6 again highest at 4.78, Kimi K3 close behind at 4.67, then Gemini at 4.39 and Grok at 4.11. This measures whether each item tests one clear idea, is answerable, and is not a reworded duplicate of another card.

Coverage: here we lost. Grok 4.5 scored 4.50, Gemini 3.72, and GPT-5.6 came last of the top group at 3.28. Coverage means spreading across the whole document rather than clustering on the parts that were easiest to write about. That is a real weakness in what we ship today, and it is on our list.

One model was clearly behind: DeepSeek V4 Pro finished last on grounding, question quality and overall rank, and was also by far the slowest, taking around 80 seconds against roughly 23 for the fastest. It is a capable model in general. On this particular job it was not competitive.

How we actually decide

Quality is the filter. Speed is only the tiebreak.

DeepSeek was eliminated on quality, and no amount of it being cheap would have changed that. Between the four that were left, the honest read is that they are close, so we look at what a close call should turn on for a study tool: grounding and question quality, because those are what determine whether the deck teaches you the right thing.

GPT-5.6 was top on both. It also happened to be the fastest of the five, at around 23 seconds for a set against 30 for Grok, 45 for Kimi and 80 for DeepSeek. That is a nice bonus, not the reason.

If the order had been reversed, if the fastest model had been the one with the shakiest grounding, we would have shipped the slower one. We have made that exact call before.

Why we run this again

Every model in this comparison is newer than a year old, and at least one of them will be replaced before the next semester ends. A model choice made once and never revisited quietly becomes a model choice made by whoever shipped fastest in 2025.

So this is a recurring test, not a one off. Same three documents, same frozen instructions, same blind panel, new contenders as they arrive. When a new model wins on grounding and question quality, we switch, and the switch is boring because the harness already exists.

What this means for you

Two practical things.

Every card you generate is graded, before it ever reaches you, on whether it can be traced back to your own document. That is the property we optimise for above everything else, including speed.

And if a set ever feels too shallow or too narrow for what your exam actually asks, tell Bo to make it harder or to cover more of the material. The difficulty and scope settings are real, and coverage is the one dimension where we already know we have room to improve.