
અશોક કરણિયા
A young Spanish priest arrived in Gujarat in his twenties. He was a mathematician, sent to teach numbers. He did not know a word of Gujarati.
His name was Carlos G. Vallés — ફાધર વાલેસ. He stayed for the rest of his working life, wrote more than a hundred books in Gujarati, and became one of the language’s most loved writers. It took immersion, curiosity and, in some ways, a lifetime.
Last week I asked an AI system to write me a paragraph in Gujarati. It took four seconds. Every case ending was correct. Every vowel mark was in place. And it was completely dead.
Sixty years. Four seconds. That contrast holds both the extraordinary promise and the uncomfortable question facing languages like Gujarati. A machine can produce Gujarati. But does it know Gujarati?
The number that stopped me
Gujarati is spoken by around 55 million people — sixth most spoken in India, roughly twenty-sixth in the world. More speakers than Italian. More than Korean. Gujarati Wikipedia has about 30,800 articles. Marathi, a neighbouring language, has over 100,000.
But here is the comparison that should trouble anyone who loves this language. Gujarati Vishwakosh — twenty-five volumes, compiled by hand, on paper, completed in 2009 — contains 23,090 articles.
And then the number that ends the argument:
85,607 people have registered accounts on Gujarati Wikipedia. Eighty-four of them edited it last month.
Eighty-four. For fifty-five million people.
Not small. Low-resource.
This distinction is the one almost every conversation misses.
Gujarati has centuries of literature, poetry, philosophy, journalism and oral tradition. But AI does not learn a language by counting its speakers. It learns from what has been digitised and made available to machines.
Much of our wealth still lives in books, old journals, private collections, handwritten manuscripts, folk performance — and in people. What isn’t digital is effectively invisible to the machine.
So the model learns Gujarati largely from translations: news copy, government notices, subtitles, nearly all rendered from English. And it returns what it was given.
You can see the seams if you know where to look.
Respect collapses. Gujarati distinguishes તમે from તું. Machines routinely flatten the two and address an elder the way you would address a child. Every Gujarati reader feels it instantly. English has no equivalent to get wrong.
Idiom goes literal. ઊંટના મોંમાં જીરું becomes “cumin in a camel’s mouth.” Technically correct. But a Gujarati feels what it means before the sentence ends.
Dialect gets quietly corrected. Ask for Kathiawadi and you tend to get standard Gujarati with two words swapped — હાલો becomes the textbook ચાલો. Not because anyone decided to erase it, but because nobody wrote it down at scale.
Sometimes what emerges is not really Gujarati at all. It is Gujarati-shaped English.
Which leads to a far better question than whether AI will replace Gujarati.
How will Gujarati exist inside AI?
Preserving more than words
For generations, preservation meant dictionaries, books and archives. They remain invaluable — Bhagwadgomandal, compiled over twenty-six years by Bhagvatsinhji of Gondal, holds more than 280,000 words and stands among the great achievements of this language.
But a language is more than its vocabulary.
It is pronunciation. The rhythm of a sentence. The pause before a punchline. A lullaby. The working vocabulary of a potter or a salt-pan worker. The Gujarati of Kutch, of Saurashtra, of Surat, of a tribal community. A grandmother telling a story exactly as her grandmother told it.
AI lets us hold those things at a scale earlier generations could scarcely imagine.
Imagine a grandmother in Bhavnagar recording her memories, recipes and songs today. Twenty years on, her great-grandchild in Canada asks:
“બાની સૌથી પ્રિય વાર્તા કઈ હતી?” What was Ba’s favourite story?
And instead of a paragraph assembled from the internet, the child finds Ba’s actual story — in her vocabulary, her dialect, perhaps her own voice.That is not data preservation. That is cultural memory.
Our goal should not be to record Gujarati words. It should be to record the language.
Making Gujarati easier to use
Preservation alone saves nothing. Languages survive because people use them — and this may be where AI offers Gujarati its greatest opportunity.
Consider a child growing up in New Jersey or London. She understands Gujarati at home but cannot read it confidently, and there is no teacher nearby.
Now imagine an infinitely patient tutor available every day. It talks with her. Corrects her pronunciation. Builds a five-minute daily lesson. Turns vocabulary into a game. Explains a poem at her level. Writes stories about the things she already loves — and adjusts tomorrow based on what she found hard today.
The same capability lets a visually impaired reader hear Gujarati literature, and lets an elderly person who has never been comfortable with a keyboard simply speak to a digital service in Gujarati.
For most of history, preserving a language meant preserving its books. In the AI era, it also means reducing the friction required to use it. If Gujarati becomes effortless to speak, read, search and create in, more people will use it. If it stays inconvenient, another language wins by default.
Reopening our own literature
There is a third possibility that excites me most.
AI can make the literature we already have newly discoverable. Imagine asking:
- How has rain creatively appeared across a century of Gujarati poetry?
- What connects Kalapi’s treatment of longing to poets elsewhere in the world?
- Explain a Meghani story to a ten-year-old.
- Search fifty years of a Gujarati journal by theme rather than by title.
Questions that once demanded weeks in a library can now be opened in minutes. AI does not merely lower the cost of creating content. It lowers the cost of curiosity.
That alone could bring a generation back to Gujarati literature.
It has been done — by a radio station in New Zealand.
None of this is theoretical. Te Hiku Media, a small Māori community broadcaster, decided in 2018 that their language needed speech recognition and that nobody else would build it. They ran a competition. In ten days, their community recorded and annotated 300 hours of speech — enough to train working models. Those models now transcribe te reo Māori at around 92% accuracy.
Then they did the part that mattered most: they refused to hand the dataset to large technology companies, and wrote their own licence instead. The corpus is held in trust for the community; benefit flows back to the speakers who made it.
Four lessons transfer directly:
- Use an institution that already exists and is already trusted — not a new one.
- Ask your own community, with a deadline and a reason to take part.
- Native speakers must correct the transcripts. That correction is what turns audio into training data.
- Write the licence before anyone asks for the data.
Te reo Māori has a few hundred thousand speakers. Gujarati has fifty-five million.
The arithmetic is not subtle.
We are not starting from zero
India is building an extraordinary multilingual AI ecosystem. Bhashini provides national infrastructure for translation, speech and OCR — launched, as it happens, in Gandhinagar. BhashaDaan lets citizens contribute language data directly. AI4Bharat publishes open Indic models. BharatGen and Sarvam AI are building Indian foundation and voice models. Project Vaani is collecting spoken language across the country’s dialects.
And Gujarati already has remarkable digital foundations: Bhagwadgomandal. GujaratiLexicon. Gujarati Vishwakosh. Ekatra Foundation. Opinion Archives. Public-domain collections.
Generations before us did the hard work of collecting, classifying, publishing and digitising.
India is building the plumbing. Gujarat has already built much of the knowledge. Our generation’s job is to connect the two.
But AI is not automatically good for Gujarati
We should be equally clear about the risks.
AI can fabricate a Kalapi couplet and present it with total confidence. It can flatten Kathiawadi and Surti into standard Gujarati. It can absorb copyrighted literature without permission, strip attribution from folk material, or clone an elder’s voice without consent. Used badly in a classroom, it becomes a substitute for learning rather than a tutor.
None of these is a reason to stay away.
They are reasons for Gujarati writers, scholars, publishers and institutions to arrive early enough to help set the rules.
The Gujarati AI Mission
What if we set a collective ambition for 2035?
- Every Gujarati child with access to a patient Gujarati tutor
- Every out-of-copyright Gujarati book searchable
- Every dialect — Kathiawadi to Kutchi and beyond — recorded in a real voice
- Every folk song and lullaby preserved before its last singer goes
- Every proverb explained with context, not merely translated
- Every senior citizen able to use technology simply by speaking Gujarati
- Every researcher able to search centuries of Gujarati in seconds
- Every Gujarati school with Gujarati AI — not English AI translated into Gujarati
Look again at that list. Not one item requires a scientific breakthrough. Every one requires collection, correction, licensing and patience.
Which is to say: it requires people.
AI does not preserve a language. People do. AI only multiplies what they were already willing to do.
So it doesn’t begin with a billion-dollar model
It begins with us.
Record one elder’s story. Digitise one public-domain book. Preserve one folk song. Document ten words from a disappearing trade. Create one Gujarati lesson. Help one child find the pleasure of reading Gujarati.
And if you want something concrete to do tonight: give ten minutes to BhashaDaan. Record your voice. Validate someone else’s. Do it in your own dialect, not your school Gujarati — that is precisely what’s missing.
If you run an institution — a Parishad, a trust, a university department, a magazine — you may be Gujarat’s Te Hiku. You have the trust and the reach. Run a fortnight-long recording drive, and decide your licence before anyone asks for the data.
If you fund things, fund the boring middle. Everyone wants to fund a launch. Almost nobody funds annotation — and annotation is where the value actually is.
Thousands of people making thousands of small contributions can build something no single institution can.
The decision
Narmad spent eight years producing the first Gujarati dictionary — 25,268 words, one man, one lamp. Bhagvatsinhji spent twenty-six years and did not live to see his finished. Ratilal Chandaria, a school dropout who became an industrialist, gave the last twenty-five years of his life to Gujarati’s digital foundations: fonts, spellcheck, the lexicon online.
Not one of them had a machine. Every one of them decided this language was worth the work.
As the old line about the earth goes — we did not inherit it from our ancestors, we borrowed it from our children. The same is true of a language. And ours may be the first generation with a real chance of returning it richer than we received it.
નર્મદ પાસે યંત્ર નહોતું. ભગવતસિંહજી પાસે યંત્ર નહોતું. આપણી પાસે યંત્ર છે — નિર્ણય નથી.
Narmad had no machine. Bhagvatsinhji had no machine. We have the machine — what we don’t have is the decision.
If you work on any low-resource language — Gujarati or otherwise — I’d like to hear what is working where you are. The Te Hiku model transfers. Somebody should be running it everywhere.
This article is adapted from my talk at the third session of the 10th Bhasha-Sahitya Parishad, organised by the Gujarat Literary Academy — Saturday, 8 August 2026.
![]()

