hmar-heritage-org/paragraphs
Viewer • Updated • 33k • 13
None defined yet.
The Hmar Heritage Foundation is an independent community organization dedicated to building open-access digital infrastructure, linguistic datasets, and cultural archives for the Hmar language (hmr, ISO 639-3). Our mission is to ensure that indigenous and regional languages of Northeast India are preserved, standardized, and represented as first-class citizens in modern computational systems.
We curate, harmonize, and maintain open-access datasets for the Hmar language and the Zo language family on Hugging Face:
| Dataset | Scope / Records | DOI | Description | Link |
|---|---|---|---|---|
sentences |
101,867 train sentences (~2.5M words) | 10.57967/hf/10421 |
Multi-register sentence corpus and evaluation benchmark for language modeling. | sentences |
unigrams |
52,564 unique tokens | 10.57967/hf/10422 |
Frequency-weighted vocabulary compiled across 2.73M+ words. | unigrams |
wordlist |
43,509 lexical entries | 10.57967/hf/10423 |
Structured digital lexicon compiled from 5 lexicographical sources. | wordlist |
zo-bible |
30,974 verse anchors | 10.57967/hf/10424 |
Sentence-aligned parallel Bible corpus across 8 Zo speech varieties & English. | zo-bible |
numeral-words |
999,999 parallel rows | 10.57967/hf/10425 |
Sequential spelled-out numerals in Hmar and English (1 to 999,999). |
numeral-words |
hmingtluon |
10.19M synthetic names | 10.57967/hf/10426 |
Procedural anthroponyms dataset for NER, identity resolution, and clan genealogy. | hmingtluon |
zo-cognates |
Comparative matrix | — | Sound shifts and lexical correspondences across the Zo language family. | zo-cognates |
corpus-archive |
Multi-format vault | 10.57967/hf/10427 |
Scanned books (PDF), structured texts (JSON), and academic linguistic repositories. | corpus-archive |
culture-dump |
Community vault | — | Unstructured web publisher HTML dumps, news archives, and cultural assets. | culture-dump |
| Resource | Type | Author | DOI | Description | Link |
|---|---|---|---|---|---|
hmaraniam |
Engine / Space | Hmar Heritage Foundation | — | Fast heuristic language identification library for Hmar. | Hmaraniam Space |
| HmarBERT | Model | Donal Muolhoi | 10.57967/hf/10428 |
Foundational BERT-base model adapted from MizBERT via cognate vocabulary swaps (12.00 PPL). | azinamotoe/HmarBERT |
| Dolma | Space | azinamotoe | — | Interactive inference web interface powered by HmarBERT. | Dolma Space |
hmaraniam PyPI: pypi.org/project/hmaraniamhmr | Glottolog: hmar1241If you use datasets, language tools, or resources from the Hmar Heritage Foundation in your research or applications, please cite:
@misc{hmar_heritage_foundation_2026,
author = {{Hmar Heritage Foundation}},
title = {Open Computational Infrastructure, Corpora, and Linguistic Datasets for the Hmar Language},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/hmar-heritage-org},
keywords = {Hmar, hmr, hmar1241, Zo Languages, Low-Resource NLP, Language Preservation},
note = {Language: Hmar (ISO 639-3: hmr, Glottolog: hmar1241). Language family: Zo Languages.}
}