What we hold

Yomu stores what you read, when you read it, how long each page was on screen, every word you looked up and how well you remember each one. That is a detailed picture of a person, so this page says exactly what is kept, why it is kept, and how to take it away or destroy it. It is generated from the same list the deletion works through, so it cannot quietly fall behind the database.

The short version

What you read and learn

Your books

Signing in

What the server records about itself

How long

For as long as you have an account, and no longer. Nothing above is deleted on a schedule, and that is deliberate rather than lazy: your event log is the only record of your own learning, it cannot be reconstructed, and a product that silently dropped a year-old review would be corrupting the very thing it exists to keep.

Five things do expire on their own, without you asking:

Who else sees it

That is the complete list. There is no analytics provider, no third-party error reporter, no advertising network and no data broker, and none of the words on this page can be made true by anything other than the code.

The person who runs this server can see some of it, and here is exactly what. Whoever runs a database can read it — that is true of every product there has ever been, and a page that implied otherwise would be lying. What is worth stating is the part that is built rather than merely possible: there is one owner-only screen, and it shows your name, when you registered, how many words and books and reading events are yours, when you last had a book open, and the titles of the books on this server together with how well they parsed. It deliberately does not show the text of a book, an email address, or anything you typed — not even to the owner, and a test searches the whole screen for seeded book text and fails if a word of it appears. There is no way to sign in as you from it, and no way to delete your account from it. On a copy running on your own computer, the person that screen belongs to is you.

This server does record its own faults, and that is not a third party. When something here breaks while you are using it, the server writes down which screen failed, what the error was and your account id — so that the person who runs it finds out from a dashboard rather than from you, and so that a screen failing for one person can be told apart from a screen failing for everybody. Nothing leaves this machine, nothing you typed or were reading is in it, and each fault is one row with a count on it rather than one row every time it happens. They are deleted after thirty days — from the live database on the day they turn thirty days old, and out of the last backup holding them within 12 months, which is the same bound everything else on this page has and for the same reason. That is the whole of it, and it is listed table by table above like everything else.

Taking it, and leaving

Settings has both. The export is a single JSON file: your cards, your schedule, your known words and the complete event log (every review, every word you looked up, every page reading credited) and, for every word in it, the written forms, the reading, the meaning, the pitch accent with its source and the frequency rank with its corpus and licence. It is a vocabulary, not a history of numbers. It does not include the text of your books, which you imported from your own files, and there is no Anki .apkg yet.

Deletion is deletion. Everything marked above as deleted goes in a single database transaction, and the deletion refuses to report success while a single row that names you is still there. It checks the database’s own catalogue afterwards, so a table added in a future version is covered before anybody remembers to think about it. If anything were left, the whole thing is undone and you are told, rather than being told you are gone when you are not.

There is no grace period and no thirty-day window, on purpose: an account that can still be restored is an account that has not been deleted. Take the export first. It is offered on the confirmation screen for that reason.

Backups are the one place a deletion does not reach, and that gap is bounded. A deletion empties the live database immediately and completely. It does not, and cannot, reach into copies that were already taken. A backup is a photograph of an earlier moment, and rewriting one is not a thing a database can do. So an account deleted today is still inside the copies taken before today. What we can promise is when that stops being true: the last copy holding you ages out within 12 months, because the retention rule keeps everything from the last week, then one a day for three months, then one a month for 12 months, and then nothing. Saying “deleted everywhere, instantly” would be the easy sentence and it would not be true; saying “until it ages out” and never giving a number would be the easy evasion.

Two things about that number are worth saying plainly. It is not a shelf life somebody typed onto a page: it is the rule the pruning tool applies, imported into this page from the same module, and a test fails if the two ever disagree. And it has one exception: the tool never deletes the last backup in existence, whatever its date, because a policy that can leave a server with no backup at all is a worse failure than a copy kept too long. If backups had stopped a year ago, that single remaining copy would stay.

Statistics counted across everybody, and why there is no opt-out

One figure in this product is counted across all of its readers rather than for one of them: how often a given Japanese word is recalled at its first scheduled review, how often it lapses, how many repetitions it takes before it sticks. It is the one thing here that gets better the more people use it, and it is what will eventually let a new card start with a sensible interval instead of a guess. Only answers you gave deliberately feed it. A reading credit never does, and neither does an import.

There is no opt-out, and the reason is not that it would be inconvenient. A figure is published only when at least 50 distinct learners stand behind it; below that it is suppressed entirely rather than shown with a caveat, and the suppression is in the code that computes it, not in the screen that draws it. So no number that leaves this server can be one person’s answers, and a number that cannot be traced back to a person is not personal data. There is nothing there to opt out of, which is the same reason those figures survive your deletion, when everything else about you does not.

The honest half of that. The result identifies nobody; getting to it reads your reviews, with your account attached, because counting distinct learners is precisely what decides whether a figure may be shown at all. So the derivation runs over personal data even though what comes out of it does not. Deleting your account stops your rows feeding it: the next recomputation no longer contains you, and cells already computed stay because they hold no identifier and cannot be unmixed back into people.

The floor of 50 was not originally chosen for privacy. It was chosen because a recall rate computed from four people is not a fact anybody should act on, and one visibly wrong figure would cost more trust than the feature earns. It does both jobs, and it is worth knowing that it was set by the first argument.

One thing you are enrolled in without being asked

This is not a privacy question. It is here because answering it inside one would be the quiet way to tell you, and because you would otherwise find it out by being interrupted. Yomu’s central claim is that reading a word counts as reviewing it. Nobody knows whether that is true, including us, so the product runs an experiment on itself, and every account is in it.

What actually happens. Your account is put in one of two halves by a hash of its id: sampled, or not. It is decided the first time you start a sitting, never changes, is never chosen by you and is never accepted from your device. A client that could name its own half would be an opt-out with extra steps. If you are in the sampled half then you are sometimes asked to recall a few words cold (before a batch of reviews starts, at the end of a chapter, or as you put a book down), at most 6 in a sitting, and never in the middle of a page. The words asked are the ones reading has been carrying, paired against words you actually reviewed, because the whole point is the comparison between the two.

Why you cannot switch it off. An opt-out would not remove a random slice of readers. It would remove the readers most confident that reading alone is working for them, which is exactly the group whose self-assessment the check exists to test. What was left would be a calibration built from the people who least needed calibrating, and it would report that the mechanic works whether or not it does. The honest options were to ask nobody and claim nothing, or to ask everybody and say so on a page like this one. The rate can be turned down; the thing cannot be turned off.

What it costs you, plainly. Being quizzed is the opposite of immersive reading, and that tension is real rather than rhetorical. It is why the checks sit at pauses you had already taken, why there are never more than 6 of them in a sitting, and why the unsampled half exists at all: if being asked makes people read less, that shows up as a difference between the halves, and it is evidence about the feature rather than an argument about it. If it turns out that reading does not clear reviews, the product’s headline claim is the thing that changes.

What it stores. Which half you are in is recorded on each sitting, and your answers are ordinary events in your own history, both listed in the table above, and both deleted with your account.

Who is responsible, and how to complain

Fergus Leen runs this. Not a company and not a team: one person, who wrote the code this page describes and who holds the database it describes. In the language of the GDPR he is the data controller, which is not a title anybody awards. It follows from deciding what is collected and why, and that is what running it means. Anything about your data (a copy, a correction, a deletion, an objection to how it is used, or a complaint) goes to an address this deployment has not published, and a person reads it.

Three details are still missing from that paragraph and they are named at the bottom of this page rather than guessed at: an address to write to, which is the one gap on this list that stops you exercising any of the rights above; and a postal address, for anything that has to be served on paper; and the country this is established in, which is what decides the supervisory authority that supervises us. None of them changes how you complain, which is the next paragraph.

You may complain to your own country’s authority, wherever we are. This page used to say that how to complain followed from knowing where the controller sits. That was wrong. Article 77 of the GDPR gives you the right to complain to the supervisory authority of the country where you live, the country where you work, or the country where the thing you are complaining about happened. Whichever of those you prefer, and regardless of where we are established. So: write to the address above first, because most things are quicker to fix than to adjudicate, and if that does not satisfy you, take it to your own national authority. It is obliged to deal with your complaint and to tell you what has come of it.

And there is a court after that. If a supervisory authority does not handle your complaint, or does not tell you within three months what it is doing with it, you can take that authority to court (Article 78). You can also go to court against us directly, without complaining to anyone first (Article 79). Nothing on this page, and nothing you agreed to on the way in, asks you to give either of those up.

Sixteen

You need to be sixteen to have an account here. Sixteen is the age the GDPR takes as its default for someone deciding on their own behalf about their own data. Individual countries are allowed to set it lower, and some go to thirteen. This does not follow them down or vary by country, because one age that is right everywhere is worth more here than an age that is exactly right in one place.

Nothing asks your age and nothing checks it. There is no date-of-birth box, no estimate made from anything, and no document. That is a plain statement of what the software does rather than a policy: the requirement is real and it is unenforced, and a page that implied a check existed would be lying about the one thing on it that a parent would care about.

If we learn that an account belongs to someone under sixteen, it is deleted by the same deletion described above, which is a real one: the history, the books, the email address, all of it, in one transaction that refuses to report success while a row naming that person is still there. If you believe a child has an account here, write to the address named above.

What this page cannot yet tell you

These are real gaps, not oversights, and they are here rather than filled with something that sounds right. A stated gap can be closed; an invented promise cannot be withdrawn. This list used to have six things on it. Four are now answered above rather than quietly dropped, a fifth has shrunk to the two facts below, and the last one has not moved, because it is the price of a deliberate choice rather than something waiting on a decision.

If something on this page turns out not to match what the software does, the software is the thing that is wrong, and it is a bug worth reporting.

yomu · what we hold