Drafting PhaseThis site is actively under development. Content and specifications are being incrementally refined. Please check back later for updates.
Official Policy & Standards Document
Doc Ref: HHF-POL-2026-01

Language Standards & Classification Policy

The Hmar Heritage Foundation formally establishes this policy to articulate our position on linguistic taxonomy, genealogical classification, and digital metadata standards across international language registries.

Section I · The Foundation's Position

Rejection of Colonial Exonyms and the Zo Autonym Candidate

The Hmar Heritage Foundation formally rejects legacy colonial umbrella terms, specifically the compound macro-labels "Kuki-Chin-Naga" (KCN), "Kuki-Chin" (KC), and their administrative derivatives, as obsolete exonyms that distort genetic language relationships and perpetuate colonial administrative conveniences.

In place of these colonial labels, the Foundation advocates for adopting the indigenous autonym Zo (in descriptive academic literature, South-Central Tibeto-Burman / South-Central Trans-Himalayan) as the overarching phylogenetic clade for the ~55 closely related speech varieties across Northeast India, Western Myanmar, and the Chittagong Hill Tracts. This autonym unites closely related speech varieties, including Hmar, Mizo (Lushai), Paite, Thadou, Vaiphei, Tedim, Mara, Gangte, Simte, Zou, Biate, Hrangkhol, Lai/Hakha, Khumi, and Cho/Daai, which share demonstrable, regular sound correspondences descending from a common ancestor.

To maintain full digital compatibility, the Foundation advocates preserving existing individual ISO 639-3 language codes (such as hmr, lus, pck, ted) while updating the overarching family classification from legacy exonyms to Zo across all software metadata, dataset tags, and international registries. Furthermore, we formally affirm the complete cladistic separation of Meitei (Manipuri), Naga, and Karbi branches out of the Zo node into their respective independent branches.

Section II · Historical Critique

Grierson's Stated Criteria vs. His Own Contradiction

The historical flaw began in the Linguistic Survey of India (LSI Vol. III, Part III, 1904), where G.A. Grierson acknowledged the outsider origin of the terminology:

"The name Kuki is an Assamese or Bengali word, applied to various hill tribes... Chin is a Burmese word used to denote the various hill tribes living in the country between Burma and the provinces of Assam and Bengal."G.A. Grierson, LSI Vol. III, Part III (1904, pp. 1–2)

However, Grierson broke his own stated criteria by placing Meitei (Manipuri), a historically valley-dwelling population of the Imphal Valley, under the same umbrella. He did this despite explicitly acknowledging that Meitei possesses structural, morphological, and lexical ties that align it more closely with Written Burmese and Classical Tibetan than with the surrounding hill languages:

"It will also be seen that Meithei in some respects agrees with written Burmese, as against the other languages of the group... Connection with Tibetan."G.A. Grierson, LSI Vol. III, Part III (1904, pp. 6, 20–24)

In the mid-20th century, tentative structural surveys by Robert Shafer (Kukish) and Paul K. Benedict (Kuki-Naga) merged Kuki-Chin, various heterogeneous Naga languages, Meitei, and Karbi under macro-rubrics like "Kuki-Chin-Naga". When Shafer published his classification in 1955, the undivided state of Assam had not yet been reorganized into modern states like Nagaland (1963) or Mizoram (1972/1987). Consequently, subsequent scholars easily confused legacy administrative hill groupings with true genetic language families, compounding colonial administrative shortcuts into permanent academic labels.

Section III · Cladistic Restructuring

Replacing the Parent Exonym Node with Zo

To replace colonial administrative shortcuts with scientific integrity, the Foundation advocates for replacing the top-level exonym node (Kuki-Chin-Naga / Kuki-Chin) with the authentic autonym Zo (55 speech varieties).

For immediate structural clarity, we illustrate how the Zo autonym maps onto existing subgroupings below. While sub-descriptors like "Central", "Northwestern", or "Peripheral" remain largely geographical, retaining them in this illustrative diagram ensures the tree remains immediately recognizable to researchers while demonstrating how the Zo autonym seamlessly replaces the colonial parent node.

Glottolog IDLegacy Glottolog LabelProposed Cladistic LabelScope & Core Speech Varieties
sino1245Sino-TibetanSino-Tibetan / Trans-HimalayanTop-level family root.
kuki1245Kuki-Chin-Naga (Legacy)[Node Dissolution]Reject as a non-monophyletic wastebasket node.
kuki1246Kuki-ChinZo Languages / South-Central55 Speech Varieties sharing Proto-Zo phonology.
├── cent2005Core Central Kuki-ChinCentral Zo (17 varieties)Hmar, Mizo (Lushai), Lai/Hakha, Maraic, Pangkhua.
├── oldk1252Northwestern Kuki-ChinNorthwestern Zo (16 varieties)Anal, Monsang, Moyon, Purum, Aimol, Lamkang, Tarao.
└── peri1260Peripheral Kuki-ChinPeripheral Zo (22 varieties)Tedim, Paite, Thadou, Vaiphei, Simte, Khomic, Ashö.
mani1292Manipuri / MeiteiMeitei Branch (Independent)Independent Sino-Tibetan node outside Zo.
karb1240KarbiKarbi Branch (Independent)Independent Sino-Tibetan node outside Zo.
anga1312Angami-Ao / NagaNaga Clades (Independent)Independent branches (Aoic, Angami-Pochuri, Zemeic, etc.).
Section IV · Indigenous Autonomy

Restoring Dignity and Respecting Indigenous Self-Determination

Crucially, advocating for the eviction of Naga, Karbi, and Meitei languages out of the Zo umbrella is an intentional act of restoring cultural dignity and linguistic autonomy.

Lumping these ancient, distinct groups under a colonial exonym was just as disrespectful to Naga, Karbi, and Meitei heritage as it was to Zo communities. However, the Foundation maintains clear institutional boundaries: while we strongly support evicting these groups from the colonial umbrella, where Naga, Karbi, or Meitei languages are ultimately classified within Sino-Tibetan is purely up to those respective communities, their scholars, and field linguists.

We do not presume to dictate the internal node structures of neighboring indigenous groups. Our objective is simply to remove colonial clutter so every community across the region has the freedom to define its own linguistic heritage.

Section V · Field Research

The Disconnect Between Comparative Linguistics and Legacy Registries

There is a persistent claim by institutional registries that legacy classifications are maintained to reflect "academic consensus." In reality, the primary comparative linguists actively conducting field research explicitly reject these colonial labels.

In his benchmark phonological reconstruction Proto-Kuki-Chin (STEDT Monograph 8, 2009), Dr. Kenneth Van Bik explicitly excludes all Naga languages (such as Ao, Tangkhul, Zeme) from the family, proving that Naga languages share no unique phonological innovations with Kuki-Chin.

Similarly, in "The Tibeto-Burman Languages of Northeastern India" (2003) and "The Sino-Tibetan Languages of Northeast India" (2017), Prof. Robbins Burling and Dr. Mark W. Post explicitly demonstrate that "Naga" and "Kuki-Chin-Naga" are not valid genetic units. They emphasize that compound colonial labels are areal catch-alls lacking comparative evidence of shared sound innovations.

Furthermore, Prof. Scott DeLancey (2013, 2015) treats the South-Central / Zo branch as a primary, distinct branch separate from surrounding Himalayan and Brahmaputran lineages, while Dr. Linda Konnerth (2018) explicitly recommends adopting South-Central Trans-Himalayan specifically because colonial exonyms carry negative connotations for native speakers. Citing "academic consensus" to justify 100-year-old administrative shortcuts while ignoring the explicit phonological reconstructions of modern field researchers misrepresents contemporary linguistics.

Section VI · Registry Infrastructure

Methodological Contradictions, Indexing Traps, and Institutional Neglect

This institutional resistance to change relies on two structural flaws: a fundamental methodological contradiction and a circular indexing trap.

First, Glottolog 5.3 (glottocode kuki1245) explicitly claims that its tree nodes are strictly genealogical, based on shared phonological innovations. Yet it continues to perpetuate top-level labels dating back to Grierson and Shafer, despite Grierson and Shafer themselves explicitly stating that terms like "Kuki-Chin-Naga" were merely geographical conveniences. Even SIL Ethnologue documentation admits that "Kuki-Chin-Naga" is a geographical clustering rather than a proven monophyletic node. Naming genealogical nodes with geographical labels is an unresolved methodological contradiction.

Second, registries create a self-reinforcing loop. Researchers are forced to include legacy exonyms in paper abstracts simply so their work remains discoverable within registry databases. A real-world example is seen on open platform archives like Hugging Face, where community-built speech corpora are named Zomi ASR by native developers, yet platform automated metadata tags still force legacy kuki-chin tags to ensure search indexing. Registries then cite those abstract keywords as proof that academics "prefer" the old terms, ignoring the fact that the papers themselves explicitly reject them.

Ultimately, maintaining this obsolete structure is nothing more than a sign of institutional neglect. Anyone who has actually studied these languages in the field recognizes these flaws immediately. That international registries have allowed a 100-year-old administrative shortcut to linger uncorrected demonstrates decades of neglect toward the region's linguistic heritage. As foundational infrastructure, it is the duty of institutions like ISO 639-3, Glottolog, and SIL Ethnologue to correct structural errors rather than hide behind circular database mechanics.

Section VII · Governance & Socio-Technical Impacts

The Cost of Misclassification and the Denial of Institutional Responsibility

Inaccurate language classification is not a sterile technical exercise; it has real-world consequences for governance, education, and cultural identity.

For decades, official government registries and census frameworks have lumped distinct speech communities together under broad colonial exonyms such as "Kuki", a label that Grierson himself acknowledged was an external exonym. While partial administrative updates have been made over time, the overarching official label still lingers in state policy.

This administrative lumping severely distorts history. For communities that transitioned from oral traditions into written literacy only within the last century, official state lumping corrupts how history is documented and passed down. When young generations read government documents that force a foreign administrative label onto their heritage, it rewrites their ancestral origin stories by bureaucratic fiat.

Furthermore, modern digital governance depends on precise language tags. Under educational mandates (such as India’s National Education Policy 2020), mother-tongue textbook development, curriculum funding, and teacher training depend on recognized language classifications. Flawed grouping leads to minority varieties being administratively subsumed, missing out on textbook printing, preservation grants, and localized digital software interfaces.

Pretending that these legacy labels are "just linguistic classifications", while ignoring their real-world impact on civil rights, policy, and human dignity, is a complete denial of institutional responsibility. Language is the vessel of cultural identity, and standards bodies must be held accountable for the administrative legacy they maintain.

Section VIII · Digital Implementation

Operationalizing the Zo Standard Across Open Datasets

To lead by example, the Hmar Heritage Foundation operationalizes this policy across all its open digital repositories, text corpora, speech recordings, and public language datasets.

We maintain individual ISO 639-3 codes (such as hmr for Hmar) for seamless POSIX locale, Unicode, and software compatibility, while replacing overarching exonym labels (KCN / KC) with Zo across all dataset metadata and documentation headers. All open-source text and audio datasets published by the Foundation on platforms like Hugging Face will pair stable ISO codes with updated taxonomic headers (e.g., language_iso639_3: "hmr", clade: "Zo", subgroup: "Central Zo (Hmaric)", academic_alternative: "South-Central Tibeto-Burman").

Finally, the Foundation will submit formal change proposals to the Glottolog Editorial Board to dissolve the non-monophyletic macro-node kuki1245 ("Kuki-Chin-Naga"), isolate Meitei (mani1292) and Naga branches, and rename kuki1246 to Zo / South-Central, while petitioning SIL International and ISO 639-5 to update overarching family classifications.

References & Selected Bibliography

Benedict, Paul K. (1972). Sino-Tibetan: A Conspectus. Contributing editor James A. Matisoff. Cambridge: Cambridge University Press.

Burling, Robbins. (2003). "The Tibeto-Burman Languages of Northeastern India." In Graham Thurgood and Randy J. LaPolla (eds.), The Sino-Tibetan Languages, pp. 169–191. London: Routledge.

DeLancey, Scott. (2013). "The History of Postverbal Agreement in Kuki-Chin." Journal of the Southeast Asian Linguistics Society (JSEALS), 6: 1–17.

DeLancey, Scott. (2015). "Morphological Evidence for Tani Subgrouping." In Linda Konnerth et al. (eds.), North East Indian Linguistics 7. Canberra: Pacific Linguistics.

Grierson, George A. (1904). Linguistic Survey of India: Vol. III, Part III, Tibeto-Burman Family: Specimens of the Kuki-Chin and Burma Groups. Calcutta: Office of the Superintendent of Government Printing.

Hammarström, Harald, Robert Forkel, Martin Haspelmath, & Sebastian Bank. (2024). Glottolog 5.0 / 5.3: Kuki-Chin-Naga (kuki1245). Leipzig: Max Planck Institute for Evolutionary Anthropology.

Konnerth, Linda. (2018). "The Historical Phonology of Monsang (Northwestern South-Central / 'Kuki-Chin')." Himalayan Linguistics, 17(1): 111–144.

Matisoff, James A. (2003). Handbook of Proto-Tibeto-Burman: System and Philosophy of Sino-Tibetan Reconstruction. Berkeley: University of California Press.

Post, Mark W., & Robbins Burling. (2017). "The Sino-Tibetan Languages of Northeast India." In Graham Thurgood and Randy J. LaPolla (eds.), The Sino-Tibetan Languages (2nd ed.), pp. 213–242. London: Routledge.

Shafer, Robert. (1955). "Classification of the Sino-Tibetan Languages." Word, 11(1): 94–111.

Van Bik, Kenneth. (2009). Proto-Kuki-Chin: A Reconstructed Ancestor of the Kuki-Chin Languages. STEDT Monograph Series No. 8. Berkeley: Center for Southeast Asia Studies, University of California.