What the grader does, where it is wrong today, and what we do about it. No accuracy percentage; the last section says why.
Checked against production code on .
In short
Outcome
Before you rely on it, you know where the grader errs in your favour, where against, and what to do in each case.
Cost
Five minutes to read. In a session, a typed fallback when a spoken answer keeps being misheard.
We guarantee
Every grade shows the reference answer. A card that cannot be heard after a repeat is not scored. The fast matcher is regression-tested against 4,743 labelled answers on every change.
We do not guarantee
An accuracy percentage. None is published, and this page says why rather than inventing one.
What you are trusting
You speak, a machine decides whether you knew it, and the schedule for that card moves on its word. A grade that is wrong in either direction costs you. A false "wrong" drags a known card back tomorrow. A false "right" lets a gap sit until the exam finds it.
So you should know how the decision is made and where it fails. This page is that list, checked against the code running in production on the date at the top, and re-checked every eight weeks.
How a spoken answer is graded
Before any grade: the four cases where the grader asks you to say it again. A transcript it cannot trust is never scored.
Fast matching first. Exact and normalized matches, number and formula forms, a list matcher that ignores order and duplicates, and, in English, correction of common mishearings ("talk to you" for Tokyo). Obvious right and obvious wrong answers resolve here in under a millisecond and the model never sees them.
The model for everything else. Anything the fast matcher is not sure about goes to an LLM grader with the question, the reference answer and your transcript. It scores the answer, gives partial credit and says what was missing. It is a judgment on meaning, not a keyword match.
Follow-ups. A middling score gets one clarifying question ("Did you mean X?"). A wrong answer that is close can get one guiding question before the grade. One follow-up per card.
Every grade shows the reference answer, so a wrong mark is visible to you the moment it happens.
Where it is wrong today
Misheard list items are marked wrong. When speech-to-text mangles one item in a list, "are gone" for Argon, "specific" for Pacific, the answer can be scored as missing that item. The model reads text, not sound; in context it often recovers the word, but not reliably. We tested a fix that asks the model about each item on its own. It was worse: it accepted 12 of 19 correct mangled-list answers where whole-answer scoring accepted 17 of 19, and its misses were confident, so no threshold rescues them. We are not building it. (48 hand-labelled answers, one run, one prompt wording.)
Sound-alike wrong lists can be accepted. The fast list matcher tolerates fuzzy and phonetic variation per item so that recognition noise does not fail you. The cost: "paris, bern, rome" passes for Paris, Berlin, Rome, and "photon, neuron, electrode" passes for proton, neutron, electron. Bern is not Berlin. This is the opposite error, a wrong answer confirmed, and it is live.
A short correct answer to a long reference can be marked wrong without the model seeing it. "photosynthesis" against a reference that is the full definition, or "canine, feline, avian" for "Dog, Cat, Bird": three terms or fewer with almost no overlap are rejected by the fast matcher. We cannot loosen that gate without also accepting "cellular respiration" against the same photosynthesis definition; the two are identical in shape to a term-overlap check. Open, not fixed. If you write your own cards, a reference answer that names the term avoids it.
"Not sure" means different things per speech provider. The say-it-again floor only works where the provider's confidence score reflects the words. Soniox and Google report that, and the floor is 0.65. Deepgram Flux reports whether it is confident the turn ended, which says nothing about whether the words are right, so there is no floor on it and a mishearing from Flux goes straight to grading. Whisper reports no confidence at all. Which provider served your session decides whether this safety net was there.
Mishearing tolerance is English-first. Homophone correction, phrase collapse and the phonetic near-miss check compare English sounds. In the other 60+ languages more answers go straight to the model grader, which reads text. Pronunciation scoring is language-agnostic.
The model grades text. Whatever the recognizer wrote is what the model reads. Context helps ("nitrogen, oxygen, are gone" is readable as argon) but is not guaranteed to.
Typed answers skip the cut-off and pronunciation checks. They were submitted whole and are graded as typed.
What Quizlar does about it
Asks before grading. Low confidence, cut-off, near-miss, unclear pronunciation: one repeat each, then the grade.
Leaves a card it cannot hear. Asked once more and still no usable transcript: the card is skipped with no rating and no review-history entry. Your schedule is untouched.
Prefers deferring to deciding. When the fast matcher was last re-tuned it was moved to hand more cases to the model instead of deciding them. The cost is model calls (17.7% of corpus cases deferred, up from 10.8%); the benefit is fewer confident wrong grades.
Guards against regression in CI. The fast matcher runs against 4,743 labelled spoken and typed answers on every change. A change that accepts a wrong answer, hard-rejects a right one, or defers more than the ceiling fails the build. That is a regression guard, not an accuracy figure: the corpus is the one the matcher was tuned on.
Shows the reference answer every time and keeps a typed fallback under every question.
Why there is no accuracy number
A figure that held would need a held-out set written after the last tuning, across subjects, accents, microphones and 60+ languages. Each held-out set we have written so far was used once to find a bug and then retired, because a set that informed a fix is no longer unbiased. Until a fresh one exists we do not publish a number. The counts above are what we have.
What this costs you: you cannot compare Quizlar to another grader on a percentage. What you can do is run a session on your own material and watch the reference answers.