Received: 10 August 2025; Revised: 27 December 2025; Accepted: 27 December 2025; Published Online: 29 December 2025.
J. Inf. Commun. Technol. Algorithms Syst. Appl., 2025, 1(3), 25316 | Volume 1 Issue 3 (December 2025) | DOI: https://doi.org/10.64189/ict.25316
© The Author(s) 2025
This article is licensed under Creative Commons Attribution NonCommercial 4.0 International (CC-BY-NC 4.0)
Spell Checker for Low-resource Konkani Language
Annie Rajan,
1,*
Nehal Kalita
2
and Ambuja Salgaonkar
3
1
Department of Computer Science, DCT’s Dhempe College of Arts and Science, Panaji, Goa, 403001, India
2
Independent Researcher, Navi Mumbai, Maharashtra, India
3
Department of Computer Science, University of Mumbai, Mumbai, Maharashtra, 400098, India
*Email: ann_raj_2000@yahoo.com (Annie Rajan)
Abstract
A spell checker is an application that identifies misspelled words by analyzing the sequence of characters in
each word. Spell checking applications exist for many of the Indian languages in the Eighth Schedule of the
Indian Constitution. However, there are not as many spell checkers for languages that were added later, some of
which are low-resource languages. Konkani is one such language. This is the first time a spell checker has been
developed for Konkani, in Devanagari script. Konkani is a macrolanguage, and developing a spell checker is
challenging. We have presented the design and implementation of the spell checker. The proposed approach
makes use of dictionary lookup to identify correct words and minimum edit distance to suggest correct words
for misspelled words. This spell checker also achieved a high F-score after being tested with a set of Konkani
words. It has 1,510,514 unique words in the dictionary.
Keywords: Natural language processing; Indian language; Diacritics; Minimum edit distance; Python.
1. Introduction
India is a country with 22 languages in its Eighth Schedule language list
[1]
to the constitution and 38 languages
that are not on the list. A developing country like India needs the digital footprint of these languages as various
Natural Language Processing (NLP) tools. Building NLP tools like Part-of-Speech (PoS) tagger,
[2]
Named Entity
Recognizer (NER),
[3]
morphological analyzer,
[4]
etc. is a challenging task since to build these tools there is a need
for an annotated corpus in the language in which the tool is built. Information processing in low-resource
languages involves technologies and methods to understand corpora and linguistic databases. By processing
this information, the resources of low-resource languages can be preserved, helping to bridge communication
gaps.
A spell checker is an application that flags words in a document that may not be spelled correctly.
[5,6]
A spell
checker is a basic need of a word processor in any language. The spell checker analyzes the written text in order
to identify any misspellings and gives the best correct suggestions for those misspellings. Spell checking
applications present valid
document. The user then selects from a list of suggestions or chooses to ignore the suggestions and accept the
current word as valid. Regardless of how often this is done, the spell-checking application will perform its task
independent of the types of misspelled words most commonly made by the user. The spell checkers are crucial
in making quality content without any mistakes or ambiguity. A misspelled word can change the meaning, focus,
and intention of a word and therefore its content, and also can lead to reading and attention discomfort.
The essence of digital applications is exponentially increasing day by day. It is difficult to imagine a regular day
without search engines, social media, online news, emails, and word processing. Further, there are other NLP
applications like speech-to-text and text-to-speech engines, Optical Character Recognition (OCR) systems,
speech synthesizers, and Machine Translation (MT) systems that are evolving. Spell checkers and correctors
play a crucial role in the development of all these applications, and they are deeply coupled with the NLP
ecosystem. Extensive work is reported for English spelling detection and a limited number of Indian languages,
whereas no work has been reported for Konkani, the state language of Goa, India. Konkani is a low-resource
macrolanguage that has multiple scripts and is classified in the linguistic database ISO 639-3.
[7]
The Konkani
language has 36 consonants and 12 vowels. The methods available for other languages cannot be directly
applied to Konkani. One such example is phonetic based spell checking which cannot be directly implemented
in a generalized Konkani spell checker. This is because the pronunciation differs among different variants of