Home

Data attributions

Kotonoto’s factual language data comes from open lexical resources. As their licenses require:

Japanese

  • JMdict — Japanese–English dictionary data (entries, readings, senses, glosses). © Electronic Dictionary Research and Development Group (EDRDG). Used under CC BY-SA 4.0.
  • KANJIDIC2 — kanji information (readings, meanings, stroke counts, grades, frequency). © EDRDG. Used under CC BY-SA 4.0.
  • JMnedict — Japanese proper-noun dictionary (names, places, organisations). © EDRDG. Used under CC BY-SA 4.0, via the jmdict-simplified JSON conversion (conversion released under CC0).

Any redistributed derivative of this data — including the vocabulary export — remains under CC BY-SA 4.0 and carries this attribution.

Chinese

  • CC-CEDICT — Chinese–English dictionary data (headwords in both scripts, pinyin, senses). Maintained by MDBG. Used under CC BY-SA 4.0. Redistributed derivatives — including the vocabulary export — remain under CC BY-SA 4.0 and carry this attribution.
  • Unihan Database — per-character facts (Mandarin readings, English definitions, stroke counts, Kangxi radicals). © Unicode, Inc. Used under the Unicode License.

Tools

  • kuromoji.js (Apache-2.0) with the IPADIC dictionary — Japanese tokenization.
  • wanakana (MIT) — kana / romaji transliteration.
  • jieba-rs (via @node-rs/jieba, MIT) — Chinese word segmentation.
  • complete-hsk-vocabulary (MIT) — HSK level lists, used to flag common Chinese words.