AI & ML interests

None defined yet.

Recent Activity

azinamotoe  updated a dataset 3 days ago
hmar-heritage-org/paragraphs
azinamotoe  published a dataset 3 days ago
hmar-heritage-org/paragraphs
hmar-heritage  updated a dataset 3 days ago
hmar-heritage-org/paragraphs
View all activity

Organization Card

Hmar Heritage Foundation

The Hmar Heritage Foundation is an independent community organization dedicated to building open-access digital infrastructure, linguistic datasets, and cultural archives for the Hmar language (hmr, ISO 639-3). Our mission is to ensure that indigenous and regional languages of Northeast India are preserved, standardized, and represented as first-class citizens in modern computational systems.


Linguistic Datasets & Corpora

We curate, harmonize, and maintain open-access datasets for the Hmar language and the Zo language family on Hugging Face:

Dataset Scope / Records DOI Description Link
sentences 101,867 train sentences (~2.5M words) 10.57967/hf/10421 Multi-register sentence corpus and evaluation benchmark for language modeling. sentences
unigrams 52,564 unique tokens 10.57967/hf/10422 Frequency-weighted vocabulary compiled across 2.73M+ words. unigrams
wordlist 43,509 lexical entries 10.57967/hf/10423 Structured digital lexicon compiled from 5 lexicographical sources. wordlist
zo-bible 30,974 verse anchors 10.57967/hf/10424 Sentence-aligned parallel Bible corpus across 8 Zo speech varieties & English. zo-bible
numeral-words 999,999 parallel rows 10.57967/hf/10425 Sequential spelled-out numerals in Hmar and English (1 to 999,999). numeral-words
hmingtluon 10.19M synthetic names 10.57967/hf/10426 Procedural anthroponyms dataset for NER, identity resolution, and clan genealogy. hmingtluon
zo-cognates Comparative matrix Sound shifts and lexical correspondences across the Zo language family. zo-cognates
corpus-archive Multi-format vault 10.57967/hf/10427 Scanned books (PDF), structured texts (JSON), and academic linguistic repositories. corpus-archive
culture-dump Community vault Unstructured web publisher HTML dumps, news archives, and cultural assets. culture-dump

Language Tools & Downstream Models

Resource Type Author DOI Description Link
hmaraniam Engine / Space Hmar Heritage Foundation Fast heuristic language identification library for Hmar. Hmaraniam Space
HmarBERT Model Donal Muolhoi 10.57967/hf/10428 Foundational BERT-base model adapted from MizBERT via cognate vocabulary swaps (12.00 PPL). azinamotoe/HmarBERT
Dolma Space azinamotoe Interactive inference web interface powered by HmarBERT. Dolma Space

Institutional Links & Standards


Citation & Attribution

If you use datasets, language tools, or resources from the Hmar Heritage Foundation in your research or applications, please cite:

@misc{hmar_heritage_foundation_2026,
  author    = {{Hmar Heritage Foundation}},
  title     = {Open Computational Infrastructure, Corpora, and Linguistic Datasets for the Hmar Language},
  year      = {2026},
  publisher = {Hugging Face},
  url       = {https://huggingface.co/hmar-heritage-org},
  keywords  = {Hmar, hmr, hmar1241, Zo Languages, Low-Resource NLP, Language Preservation},
  note      = {Language: Hmar (ISO 639-3: hmr, Glottolog: hmar1241). Language family: Zo Languages.}
}

models 0

None public yet