Where the words come from

Every definition, reading and count in this app was compiled by somebody else, mostly by volunteers, and mostly over decades. This page says who, and under what terms.

The dictionary behind every word you tap

JMdict — The meaning of a word, its parts of speech, and the written forms and readings a word can take — the body of every word sheet in the reader.

JMdict is the property of the Electronic Dictionary Research and Development Group, and is used under licence from the Electronic Dictionary Research and Development Group. JMdict. Creative Commons Attribution-ShareAlike 4.0, with EDRDG's additional conditions.

KANJIDIC2 — Which school year a kanji is taught in, and whether it is on the 常用漢字 list. That is what the difficulty rating on the catalogue is partly built from, and what the kanji chart on a book's page draws.

KANJIDIC2 is the property of the Electronic Dictionary Research and Development Group, and is used under licence from the Electronic Dictionary Research and Development Group. KANJIDIC2. Creative Commons Attribution-ShareAlike 4.0, with EDRDG's additional conditions.

UniDic (unidic-cwj) — The list of words this app can have knowledge about at all, and the reading each one takes. It is also where the pitch accent in your exported cards comes from. NINJAL's work, from the National Institute for Japanese Language and Linguistics.

Built with UniDic, from the National Institute for Japanese Language and Linguistics (NINJAL). UniDic (unidic-cwj). Modified BSD (UniDic is triple-licensed; this is the arm taken).

SudachiDict — Where one word ends and the next begins. Japanese is written without spaces, so nothing in this reader is tappable until a tokeniser has said what the words are — including which compounds come apart, and into what.

Tokenised with Sudachi and SudachiDict, by Works Applications. SudachiDict. Apache License 2.0.

JMdict and KANJIDIC2 are asked for on every screen that shows them, which is why the word sheet in the reader links here. The wording is EDRDG’s own: they ask to be credited as used under licence from the Electronic Dictionary Research and Development Group.

The corpus behind “how often this word turns up”

Tatoeba Japanese sentence export. The rank beside a word — “rank 1–1,000” — is counted from this, and it is read from the running server rather than typed here, so it names whatever corpus this copy actually holds.

Tatoeba Japanese sentence export. Short, learner-contributed conversational sentences, largely translated from English. Not prose. 2.5M words counted, 32,980 words ranked. CC BY 2.0 FR, Sentences from Tatoeba (tatoeba.org), used under CC BY 2.0 FR. Frequencies are counts computed from that corpus, not Tatoeba data.

A word the corpus never saw has no count, and is never shown as a rare one. Any corpus small enough to ship ranks a fraction of the words in a dictionary, so “not counted” is the ordinary answer rather than a bottom rung, and the sheet says the words “not counted”.

The books

Not here, on purpose: each book in the catalogue carries its own credit on its own page, pointing at the 図書カード of the edition it was set from — who transcribed it, who proofread it, and the terms it is offered under. One line here saying where the texts “generally” come from would be a weaker claim than the one every book already makes. A book you imported yourself is yours and is credited to nobody.

What we hold, and why is the other half of this: this page is about the data we did not write, that one is about the data you make.

yomu · where the words come from