gtfoo

Live product · Language learning

1 Percent More Fluent

Generated reading and listening in the language you're learning, pitched at a level you can actually read — and checked, not claimed.

Measured difficultyYour topics, your lengthContinuous 0–100 levelBehavioural calibrationNative-voice audioSpanish · Chinese · Indonesian

What's different

Vs. the usual language apps

  • The difficulty is verified against a frequency list before you see it — not a label a model claimed and nobody checked.
  • Unlimited material on any topic you name, held at your level — real articles are as hard as their subject happens to be, and app curricula run out.
  • You choose the length, so it fits the time you actually have — rather than being handed whatever the lesson or the article happens to be.
  • Long-form reading and listening from the same piece — exactly what bite-sized drilling apps don't give you, and they cap out where this starts.
  • Your level is one continuous number calibrated from behaviour, so it moves a little after every session instead of jumping between CEFR buckets.

The idea

13 years of Duolingo, and still no long reads

I'm learning Spanish so I can actually explore Latin America, and business Chinese for work — supporting Mandarin-speaking markets, where the vocabulary that matters is contracts and trade, not the weather.

I've used Duolingo for 13 years and it still does a lot right, but there's a ceiling: almost no long-form reading or listening, and no way to pick the topic — which is fatal when what you need is business Chinese and what you're given is a fruit stall. Reading Wikipedia in Spanish and Chinese failed the other way: the articles are long and as hard as their subject happens to be, so I'd stall a paragraph in and learn nothing.

Hence this: generate material on topics I actually care about, at the length I have time for, pitched at my proficiency level — scuba diving in the Yucatán one evening, a merchant fighting payment fraud the next. The name is the whole ambition: not fluency in a month, just one percent more each time I read.

v1

The smallest useful loop

The first version did one thing: you name a topic, it generates a piece in Spanish at your level, and you read it — tapping any word you don't know for a definition. The generate-measure-rewrite loop was there from the start, because a piece that hasn't been checked isn't worth showing.

But “at your level” needs a level, and asking people for their CEFR grade doesn't work — most learners genuinely don't know theirs, and self-report is systematically biased. So it opens with a quick-and-dirty yes/no word test: 90 seconds of “do you know this word?” across the frequency range.

That part isn't improvised — it's an established method from the research literature, the Yes/No vocabulary test, including the trick that makes it work: seeding the list with pseudowords — invented forms that obey Spanish spelling but mean nothing — and subtracting your over-claiming back out with the standard false-alarm correction. I did deviate from the textbook version in two places, and the section below is about why I had to.

Along the way

What it grew into

  • Difficulty is measured, then fixed

    Every generated piece is tokenised and checked against a frequency list before anyone sees it. Models land 12–15% of words outside the target band when merely asked; the checker is what turns a request into a promise. The tolerance is deliberately asymmetric — slightly too hard beats slightly too easy, because a hard piece costs a few taps while an easy one teaches nothing and quietly ratchets your level upward.

  • It can teach the words the piece is about

    Ask for a conversation about stablecoins and every word that makes it about stablecoins is rare — so the difficulty checker flagged them and the rewrite replaced them with everyday equivalents, deleting exactly what you wanted to learn. The model now declares its key terms up front; those are exempt from the difficulty budget and arrive at full strength with a gloss attached, while the prose around them stays at your level. You’re a beginner in the language and fluent in the subject, which is how professional language learning actually works.

  • Pinyin you can trust

    Chinese readings come from a dictionary that segments words first — so 银行 is yín háng but 行走 is xíng zǒu — rather than from the model, which once returned “dai e” for 大额. A wrong reading is the one error a learner cannot catch: you came to the pinyin precisely because you can’t read the answer off the characters, so it looks exactly like a right one.

  • The next piece is written while you read

    Finish a piece and its follow-on is already generating behind the review panel — topic derived from the key terms of what you just read, and up to six of your tapped words woven in. A word you had to tap is close to the definition of one you half-know, so it returns in a new context: words you needed across several pieces first, nothing from the last half hour, exempt from the difficulty budget because their presence is deliberate. It never writes ahead of a reader who isn’t following, though — at most one speculative piece exists at a time, so ignoring the suggestion costs nothing and taking it makes the wait zero.

  • Listening, spoken as dialogue

    Pieces can be heard as well as read — conversations voiced as actual dialogue, one voice per character, names not read aloud, and every voice a native speaker of the language: mainland Mandarin for Chinese, not an English voice doing its best. Each candidate was auditioned on a full-length article, because the voice that won the three-sentence audition fell apart across seventy seconds — a short audition proves nothing. Audio streams, so it starts in about 2.5 seconds rather than after the full synthesis, and every clip is cached by content hash.

Trade-offs

The judgment calls I'm proud of

Almost every hard decision here was the same shape: assume the model will drift, and put something deterministic in its way.

Never ask a model for a level

Request “B1 Spanish” and you get whatever the model feels like that day. So the app never asks. It sends concrete constraints — a vocabulary band, a mean sentence length, a permitted set of tenses — and then verifies the output rather than trusting it. Measure, don’t request: it’s the decision the whole product rests on.

Calibrate on behaviour, not self-report

The level moves on what you actually do: how many words you tapped for a definition (the honest signal — you aren’t consciously producing it), a short comprehension check in the target language, and a one-tap too easy / just right / too hard. Target is about 5% of words looked up; below that the text is wasted practice, above it comprehension collapses.

Guards that don't depend on the model complying

The band stops constraining a model as it widens, so the prompt states the budget as a target to hit rather than a cap to stay under, and names concrete words from just past the band — “be harder” alone does nothing, because the model can’t know where the band ends. Then the loop is closed from the other side: a session whose text measurably undershot its level can lower the reader’s level but never raise it.

Spaced repetition disguised as reading

The obvious feature is flashcards, and I still won’t build them — plenty of learners have a review system with years of scheduling in it, and the TSV export respects that. But the reading itself now does the reviewing: looked-up words return inside new pieces, in new contexts, which is where a half-known word becomes a known one. No decks, no due dates — the schedule is just the next thing you felt like reading.

Speech is the entire cost, so nothing is spoken speculatively

Text is nearly free; speech is effectively 100% of the running cost. So audio is only ever generated for a piece that already passed verification and that you explicitly asked to hear — a rejected draft must never reach the TTS call — and every clip is cached by content hash. That audio directory is a cache paid for in real money.

Issues faced

The measurements that measured the wrong thing

All three shipped, all three were wrong, and all three are now covered by tests — they taught me more than any feature did.

Cognates broke the vocabulary test

Rare Spanish words are disproportionately Latinate, so they're more transparent to an English speaker, not less — epinefrina, presidir, humanamente are all readable with no Spanish at all. Combined with a scoring quirk where the widest band carried 60% of the estimate from five items, this rated a genuine A2/B1 learner as C2. Three fixes: catch trials drawn per band, credit made monotonic so a band is never scored above the ones beneath it, and the scale capped where the corpus tail stops predicting anything.

The streaming release didn't stream

Audio streaming shipped verified at every layer — the provider streams to curl, Node streams from the provider, a route streams to the browser — and in production the first byte arrived minutes late, at connection teardown. Every probe had passed because every probe enqueued something on every pull. The real parser sometimes had nothing to hand on, and a ReadableStream pull() that resolves without enqueuing anything is never called again: not slow — deadlocked, by contract. The regression test now feeds the parser one JSON line split across forty-five reads and runs under a timeout, because the failure it guards is a hang, not a wrong answer.

Two benchmark runs said opposite things, and both were right

A scaffold for beginner-level generation failed its bench 2/9 against plain's 6/9 — then a rerun inverted it exactly. The model's median attempt at those levels lands at 2.2–2.8× the difficulty budget while the pass ceiling sat at 2.25×: the whole distribution straddled the line, so pass/fail at nine samples was measuring variance, not quality. The rates were decisive where the counts were noise — every failure over the ceiling, none under — and the decision came from medians: scaffold on, ceiling widened at the floor, and no reader placed into the zone where nothing works. When a benchmark flips on rerun, the interesting finding is rarely the benchmark.

Where it is now

Live, and the reading I actually do

Spanish, Simplified Chinese and Indonesian all work end to end — placement, generation, measurement, calibration and audio — each language keeping its own level, with sign-in by magic link or passkey to carry progress across devices. The loop closes by itself now: finish a piece and the next one is waiting, woven with the words you tapped; text starts appearing within a couple of seconds of asking, audio within about 2.5. And it says who it's for out loud: below roughly 780 passive words the generator measurably can't hold its own difficulty budget, so rather than pretend, the placement page tells a true beginner that a course will serve them better first.

It's live, and it's what I now read in the languages I'm learning: my topics, my length, my level — and no further updates on Lin's sandwiches.

Find your level in 90 seconds

Take the placement check, name a topic you actually care about, and read something pitched at you — in Spanish, Chinese or Indonesian.

Launch the app ↗