RS-010 /report /
IndicGPT.com: A Culturally-Grounded, Multilingual Foundational Model for the Indian Linguistic Landscape via a Dharmic Alignment Framework
Abhijeet Sarkar · Zenodo
Abstract
Abstract The advancement of Large Language Models (LLMs) has been transformative, yet it has inadvertently deepened the global digital linguistic divide. A significant portion of the world's population, particularly in linguistically diverse regions like India, remains underserved by models predominantly trained on Anglocentric data and worldviews. This paper introduces IndicGPT, a family of multilingual, trillion-parameter scale foundational models designed from the ground up to comprehend and generate content across the vast spectrum of Indian languages. We address the core challenges of Indic NLP through a series of architectural and methodological innovations. First, we detail the curation of the Bharat-Vani Corpus, a novel 15-trillion-token dataset meticulously sourced to represent India's civilizational and contemporary textual voice.
- #Foundational Models
- #Multilingual NLP
- #Indic Languages
- #AI Safety
- #Cultural AI
- #Mixture-of-Experts
- #Knowledge Graphs
- #Dharmic Ethics
Knowledge graph

