Kotonoto’s factual language data comes from open lexical resources. As their licenses require:
Japanese
- JMdict — Japanese–English dictionary data (entries, readings, senses, glosses). © Electronic Dictionary Research and Development Group (EDRDG). Used under CC BY-SA 4.0.
- KANJIDIC2 — kanji information (readings, meanings, stroke counts, grades, frequency). © EDRDG. Used under CC BY-SA 4.0.
- JMnedict — Japanese proper-noun dictionary (names, places, organisations). © EDRDG. Used under CC BY-SA 4.0, via the jmdict-simplified JSON conversion (conversion released under CC0).
Any redistributed derivative of this data — including the vocabulary export — remains under CC BY-SA 4.0 and carries this attribution.
Chinese
- CC-CEDICT — Chinese–English dictionary data (headwords in both scripts, pinyin, senses). Maintained by MDBG. Used under CC BY-SA 4.0. Redistributed derivatives — including the vocabulary export — remain under CC BY-SA 4.0 and carry this attribution.
- Unihan Database — per-character facts (Mandarin readings, English definitions, stroke counts, Kangxi radicals). © Unicode, Inc. Used under the Unicode License.
Tools
- kuromoji.js (Apache-2.0) with the IPADIC dictionary — Japanese tokenization.
- wanakana (MIT) — kana / romaji transliteration.
- jieba-rs (via
@node-rs/jieba, MIT) — Chinese word segmentation. - complete-hsk-vocabulary (MIT) — HSK level lists, used to flag common Chinese words.