When an AI tool makes your study material, you are trusting a model you never chose and usually cannot see. So it is fair to ask which one we use, and how we decided.
The short answer: we currently run GPT-5.6 on its most capable tier for writing flashcards, quizzes and exams. The longer answer is how we got there, because we do not pick by price, by benchmark leaderboards, or by whichever model was in the news that month. We test them on your kind of material.
Here is the most recent round.
The hard part of comparing models is that almost nothing stays still. Different retrieval, different source text, different output. So we hold everything fixed except the model.
We took three real documents from real courses, deliberately different in shape:
- A dense accounting glossary, wall to wall definitions
- A sparse psychology slide deck, a few bullet points per slide
- Long form marketing notes, actual paragraphs
Then we froze the source material and the instructions, and changed only the model. Thirty runs, one at a time, generating flashcards and quizzes from each document with the exact prompts the product uses.
Five models, from five different labs: GPT-5.6, Gemini 3.1 Pro, Grok 4.5, Kimi K3 and DeepSeek V4 Pro.
This is where model comparisons usually go wrong. If you score the outputs yourself, you find what you expected to find. If you let one AI score them, you have measured which output most resembles that AI's own writing style.
So the grading went to a panel of three judges from three different companies. Each judge saw all five sets side by side, with the labels shuffled, so there was no way to tell which model wrote which. They scored whether every card was genuinely supported by the source, whether the set covered the whole topic, whether the questions were clear and non repetitive, whether the difficulty matched what was asked for, and whether a real student revising would get value out of it.
Three of the judges were also contestants, which is a fair thing to be suspicious about. So we checked it: we compared the rank each judge gave its own output against the rank the other judges gave that same output. Nobody inflated themselves. DeepSeek actually graded its own set more harshly than the others did.
The result was not a landslide. Four of the five models finished in a tight group, close enough that we would not claim a winner between them on this evidence.
Grounding, which is the one that matters most: GPT-5.6 scored highest at 4.83 out of 5, with Gemini 3.1 Pro at 4.78 and Grok 4.5 and Kimi K3 both at 4.67. Grounding means every card is actually supported by your document rather than invented. A confidently wrong flashcard is worse than no flashcard, so this is the number we refuse to trade away.
Question quality: GPT-5.6 again highest at 4.78, Kimi K3 close behind at 4.67, then Gemini at 4.39 and Grok at 4.11. This measures whether each item tests one clear idea, is answerable, and is not a reworded duplicate of another card.
Coverage: here we lost. Grok 4.5 scored 4.50, Gemini 3.72, and GPT-5.6 came last of the top group at 3.28. Coverage means spreading across the whole document rather than clustering on the parts that were easiest to write about. That is a real weakness in what we ship today, and it is on our list.
One model was clearly behind: DeepSeek V4 Pro finished last on grounding, question quality and overall rank, and was also by far the slowest, taking around 80 seconds against roughly 23 for the fastest. It is a capable model in general. On this particular job it was not competitive.
Quality is the filter. Speed is only the tiebreak.
DeepSeek was eliminated on quality, and no amount of it being cheap would have changed that. Between the four that were left, the honest read is that they are close, so we look at what a close call should turn on for a study tool: grounding and question quality, because those are what determine whether the deck teaches you the right thing.
GPT-5.6 was top on both. It also happened to be the fastest of the five, at around 23 seconds for a set against 30 for Grok, 45 for Kimi and 80 for DeepSeek. That is a nice bonus, not the reason.
If the order had been reversed, if the fastest model had been the one with the shakiest grounding, we would have shipped the slower one. We have made that exact call before.
Every model in this comparison is newer than a year old, and at least one of them will be replaced before the next semester ends. A model choice made once and never revisited quietly becomes a model choice made by whoever shipped fastest in 2025.
So this is a recurring test, not a one off. Same three documents, same frozen instructions, same blind panel, new contenders as they arrive. When a new model wins on grounding and question quality, we switch, and the switch is boring because the harness already exists.
Two practical things.
Every card you generate is graded, before it ever reaches you, on whether it can be traced back to your own document. That is the property we optimise for above everything else, including speed.
And if a set ever feels too shallow or too narrow for what your exam actually asks, tell Bo to make it harder or to cover more of the material. The difficulty and scope settings are real, and coverage is the one dimension where we already know we have room to improve.