RS-010 /report /

IndicGPT.com: A Culturally-Grounded, Multilingual Foundational Model for the Indian Linguistic Landscape via a Dharmic Alignment Framework

Abhijeet Sarkar · Zenodo

DOI 10.5281/zenodo.17341822 Verified public record

Abstract

Abstract The advancement of Large Language Models (LLMs) has been transformative, yet it has inadvertently deepened the global digital linguistic divide. A significant portion of the world's population, particularly in linguistically diverse regions like India, remains underserved by models predominantly trained on Anglocentric data and worldviews. This paper introduces IndicGPT, a family of multilingual, trillion-parameter scale foundational models designed from the ground up to comprehend and generate content across the vast spectrum of Indian languages. We address the core challenges of Indic NLP through a series of architectural and methodological innovations. First, we detail the curation of the Bharat-Vani Corpus, a novel 15-trillion-token dataset meticulously sourced to represent India's civilizational and contemporary textual voice.

Knowledge graph

Connected in the archive