Statutory Infrastructure & Project Briefs

Project Initiatives & Digital Infrastructure

The Hmar Heritage Foundation stewards five specialized project briefs to lay the open digital infrastructure, data schemas, and technical groundwork for the Hmar language. The Foundation's statutory role is opening the door for community participation, public archiving, and open research rather than guaranteeing third-party commercial integration or predicting community output.

The primary focus is the collection and digitization of texts and literature. Currently, this is focused on books, articles, and written content in the public domain, as well as collecting physical books until formal copyright permissions are secured before proceeding to scanning, ensuring all published data is free of copyright encumbrance. You can read more about this archival initiative under the hmar digital library and hmar corpus archival project.

Another core initiative is creating standardized datasets for UI and UX software terminology, enabling technology interaction in Hmar. While we recognize this is an ambitious endeavor, not only because we do not influence the major institutions and corporations capable of integrating the language into their operating systems, but also because it requires meticulous terminology design so the language does not create culture shock for native speakers who have used these interfaces in English, Assamese, or Hindi for decades. It remains at the sole discretion of AOSP, the Linux Foundation, Meta, Google, Microsoft, and other technology entities to integrate these open datasets. This initiative is currently in planning, and you can read more about our approach under the hmar locale project.

While we hope the Hmar Heritage Foundation can inspire natives to write more content in Hmar, be it a simple social media post, an entry in a personal blog, or a news article, we have no way of ensuring this happens. The only way we can guarantee a digital footprint is through the hmar wikipedia incubator initiative. This benefits the community in multiple ways and we hope it could become the cornerstone for an active digital Hmar community. It makes knowledge more accessible and relatable, gives students and learners a space to practice their writing and comprehension of the language, and at the same time creates valuable text datasets for machine learning models.

When translating technical, scientific, and modern concepts, writers and translators often encounter missing words or phrases that are nearly impossible to translate directly. Expressing these ideas frequently requires using unusual words, adapting terminology, or rewriting entire sections. This highlights the vital need for the hmar open lexicon. While we hope established institutions like the Hmar Literature Society will guide formal standards, language is ever-evolving and no single entity can permanently enforce vocabulary by decree. Standards can only ever serve as living recommendations. The Hmar Open Lexicon aims to be the open platform that provides standardization resources, term lookup databases, and a collaborative space where the community can build consensus on the evolution and use of words.

To see an index and list of the projects and related resources visit : resources