HindiInput third-party language data notices ============================================ Dakshina Dataset 1.0 Copyright Google LLC and the respective dataset authors/contributors. Licensed under the Creative Commons Attribution-ShareAlike 4.0 International License: https://creativecommons.org/licenses/by-sa/4.0/ Source: https://github.com/google-research-datasets/dakshina HindiInput uses Roman/native Hindi pairs derived from the Dakshina Hindi training lexicon. The data is normalized, deduplicated, ranked, and compiled into the runtime lexicon and offline subword model. Development and test splits are reserved for evaluation and are not included in either model. Citation: Brian Roark, Lawrence Wolf-Sonkin, Christo Kirov, Sabrina J. Mielke, Cibu Johny, Isin Demirsahin, and Keith Hall. "Processing South Asian Languages Written in the Latin Script: the Dakshina Dataset." LREC 2020. StoryWeaver/AI4Bharat Hindi Transliteration Dataset Copyright StoryWeaver, AI4Bharat, and the respective contributors. Licensed under the Creative Commons Attribution 4.0 International License: https://creativecommons.org/licenses/by/4.0/ Source: https://github.com/AI4Bharat/IndicNLP-Transliteration/releases/tag/DATA HindiInput uses cleaned Roman/native Hindi pairs from the training split. Malformed edge-marked entries are rejected by the audited importer. Validation and test splits are excluded from the application binary.