Indian Institute Of Information Technology, Nagpur, Maharashtra
parulsahare2387@gmail.com
CO-Principal Investigator
Nil
Project Overview
Sanskrit is classical and ancient language of India and belongs to one of the oldest Indo-European languages. It contains historical texts of Jainism and Buddhism cultures, and is also the holy language of Hindu mythology and philosophy. The Vedas, Upanishads and Puranas are the Hindu scriptures primarily written in Sanskrit. These Sanskrit scriptures are now available digitally and can be also called as Sanskrit scanned documents. India is a multi-lingual country, where at present, 23 official languages including English exist, and 11 different scripts are used to write them. As on date, around 44% of the total Indian population uses Hindi language for speaking and writing, whereas being as sacred language, only 0.002% peoples of India use Sanskrit. Sanskrit is also a second official language of states like Himachal Pradesh and Uttarakhand. Continous efforts are being made by the Government of India to promote, preserve and rejuvenate this ancient language and further to indorse a sense of national identity and pride. This language is also seen as a storehouse of information and knowledge because of its philosophical, mythological and linguistic importance beyond its modern usage. However, still there is a gap exists between steps taken by the Government of India for its betterment and awareness of research community that works in the field of Document Image Processing. Thus, to overcome all such limitations, there is a need to provide an end-to-end and systematic solution which not only effectively pre-processes the Sanskrit documents for its digitization but also helps to understand the literature written in this holy language for general residents of India using artificial intelligence. In this project, our main focus will be on four aspects of Sanskrit document image analysis and recognition. These are namely text segmentation, text and noise classification, text recognition and then, finally we will convert Sanskrit text into Hindi and English texts. In addition to this, integration of these steps will also be done for easy accessibility of the system. More attention will be on proposing new and enhanced techniques for segmentation, identification, classification and recognition, which will then be tested on benchmark databases. This will ensure high performance during text-translation. There is a requirement to find proper architecture of artificial intelligence technique for training and testing purposes in each step. Further, the project also aims to develop a system in real-time, where the application will be built for Android operating system and can be easily deployed on handheld digital devices. This work can be used for the digitization and indexing in offices like The National Archives of India, National Mission on Libraries (Ministry of Culture) and National Mission for Manuscripts (Ministry of Culture). The proposed work is a step towards Government of India’s policies of Digital India and National Education Policy 2020.