Updated August 4, 2026

How we curate Japanese vocabulary

Moji - Learn Japanese Daily for iPhone ships 7,435 vocabulary entries. The dictionary base is JMdict from the Electronic Dictionary Research and Development Group, the study order is derived from JPDB frequency data, and the stroke paths behind writing practice come from KanjiVG. None of those sources are shipped as-is. Every entry is then edited against a written standard: at most three friendly glosses of 30 characters or less, dictionary parenthetical noise stripped out, an example sentence with an English meaning on all 7,435 entries, a usage note only when there is something genuinely worth saying, and a kanji mnemonic only where a kanji story exists. Entries that get retired leave a tombstone behind so a catalog cleanup never wipes out study progress.

This page is the long version of that, with the numbers from each pass.

Where the data comes from

Three open sources do the heavy lifting, and each one does exactly one job.

JMdict is the dictionary base: headwords, readings, and the raw sense list. It is a decades-old community project and the same backbone most serious Japanese tools are built on.

JPDB frequency data decides ordering. Each entry carries a frequencyRank keyed on the exact word and reading pair, and 7,191 of 7,435 entries have one. Ranks are backfilled by script so the mapping is reproducible: if a headword is sitting on a rare orthography or an unusual reading, its rank is buried, and the fix is to correct the form and re-run rather than to hand-patch a number. More on what we do with those ranks in why word order matters.

KanjiVG provides stroke-order paths, which is what makes the kanji and kana writing practice able to tell you that stroke three went the wrong direction rather than just marking the whole character wrong.

What “curated” means for a meaning

A raw dictionary entry is written for someone who already reads Japanese. It stacks eight senses, hedges each one, and wraps half of them in parentheses full of field labels and register notes. That is the correct behavior for a dictionary and the wrong thing to put on a flashcard that a beginner looks at for four seconds.

So meanings are cut to at most three glosses, each 30 characters or less, in the words a friend would actually use. Parenthetical dictionary noise is removed. When a sense is technically attested but nobody would teach it at this level, it goes.

That trimming is not cosmetic. In one review pass over the highest-frequency band, meanings that were simply wrong for the entry as taught came out: 頭 (あたま, head) was carrying “hair”, 夜 (よる, night) was carrying “dinner”, 一度 (いちど, once) was carrying “temporarily”.

Why 1,737 entries have no usage note

Usage notes are the small print on the back of the card: the particle trap, the register warning, the partner verb. 5,698 entries have one. The other 1,737 are empty on purpose.

The test a note has to pass is whether a Japanese friend would actually lean over and tell you this. A friend tells you that 結婚 takes と and not を. A friend tells you that うん is fine with friends but you say はい to your boss. No friend has ever leaned over to explain that an adverb typically precedes the phrase it modifies. Concrete nouns like 猫 (ねこ, cat) and 傘 (かさ, umbrella) stay blank, because there is nothing to say and a note that says nothing is a tax on the learner’s attention.

The house rules behind those notes are specific: one job per note, no grammar jargon at all, always a short reusable Japanese phrase with its English right beside it, and a hard stop at two sentences. The result reads like this:

Takes に, not を: 友達に会う (meet a friend). Meeting someone in Japanese is moving toward them, so を is always wrong here.

There is also a fatigue rule. Once a pattern has been taught on its flagship words, it does not get repeated on every remaining word of that class. The pattern is taught; the rest stay blank.

Why only kanji words get mnemonics

5,300 entries carry a hand-written mnemonic. Every one of them is a kanji word. Kana-only words do not get mnemonics at all, because a sound pun bolted onto a word with no kanji story competes with the word at recall time and the learner ends up keeping neither.

The format is fixed: name the components with a gloss for each, then one concrete image that fuses them into the meaning. The image has to be an image, not the definition restated, and the components named have to actually appear in the character. Folk imagery is allowed when it is framed as what the character looks like, never as invented etymology.

日 (sun) + 月 (moon) — the sky’s two lamps hung side by side, pouring out light together.

That is 明るい (あかるい, bright). The full method, including why false component glosses get flagged, is in the kanji mnemonic method.

Blank is allowed here too. Transparent loanwords, derived forms taught through their base word, and function words with no workable hook stay empty. A forced mnemonic is worse than none.

What the review passes actually changed

Standards on paper are cheap. Here is what two real passes over the catalog produced.

The top-300 frequency audit. A flag, verify, rewrite, verify pass over the 300 highest-frequency words applied 116 corrections across 85 entries. Grammar fields were fixed (できる re-typed and marked as happening on its own, transitivity dropped from the 十分 and 無理 na-adjectives). Readings were corrected to the standalone forms the examples actually use, with pronunciation and frequency rank re-derived to match: 光 to ひかり, 中 to なか, 後 to あと. Around 50 mnemonics were normalized to the house format, and false component claims were corrected, including 気, which is 气 plus メ and not 米. The verification step is the part worth quoting: it rejected 7 flags as subjective, caught 8 issues the first pass missed, corrected 4 proposed rewrites, and dropped 1 before anything was applied.

The first-600 sentence rewrite. Example sentences for the first 600 words in the study order were held to a budget of at most one word the learner has not been taught by that point, with katakana loanwords exempt. 48 sentences were rewritten. Of the 62 cards originally flagged, 14 turned out to be false positives from a crude matcher and were kept as they were, which is its own small point about review: the flag list is a starting position, not a verdict. Derived data was re-swept in the same pass so nothing downstream drifted out of sync. The reasoning behind the budget is in example sentence difficulty.

Frequency data gets audited by hand too. Because JPDB ranks each written shape separately, a pairing pass hands a kana word’s rank to its kanji homophones so common kana-usual words are not buried. That pass prints a cascade report on every run, and every cascade needs a human call. Skipping that audit is how 酔う (よう, to get drunk) once came to sit at rank 144.

A structural validator runs after any edit. It is deliberately narrow: it checks the things that are certain and catastrophic, meaning duplicate ids, duplicate or missing study-order values, keys the app’s decoder would silently drop, retirement-ledger breakage, and whether a card’s word, reading, romaji, and part of speech all describe the same form. Taste is not automated, on purpose.

Retiring an entry without erasing your progress

Catalogs change. Duplicates surface, a headword turns out to be the wrong form to teach, two entries collapse into one. The problem is that a learner’s review history is attached to the card, so quietly deleting an entry would quietly delete whatever they had built on it.

So entries are never just deleted. Removing one requires adding a row to a tombstone ledger naming the retired id, a survivor id, and the reason. On the next launch, the app merges the retired card’s study state into its survivor and then removes it. Chains get flattened at authoring time so a survivor is always a live entry, and the validator refuses a ledger that breaks any of those rules. The ledger currently holds 9 entries, most of them duplicate headwords folded into the entry that should have carried them.

What this does not claim to be

Moji’s catalog is not a dictionary and is not trying to become one. 7,435 entries is a curated teaching set, chosen and ordered for a beginner working toward JLPT-level comprehension, not a complete inventory of the language. When you need every attested sense of a word, use a dictionary. JMdict is right there, and it is free.

Curation also means the standards above cost coverage. Holding the untaught-word budget across the first 600 entries took a rewrite pass. Refusing to write a usage note unless there is a real fact leaves 1,737 cards with a blank field. Those are the trades, and they are the intended ones.

If you want to see what the ordering and the review passes add up to in practice, the app is free to try, and the whole comprehension path, flashcards with mnemonics and example sentences included, is on the free tier: Download Moji on the App Store

Sources