• Home
  • Advanced Search
  • Directory of Libraries
  • About lib.ir
  • Contact Us
  • History

عنوان
Combining Image and Text Processing for the Computational Reading of Arabic Calligraphy

پدید آورنده
Alsalamah, Seetah

موضوع
Computer science,Linguistics

رده

کتابخانه
Center and Library of Islamic Studies in European Languages

محل استقرار
استان: Qom ـ شهر: Qom

Center and Library of Islamic Studies in European Languages

تماس با کتابخانه : 32910706-025

NATIONAL BIBLIOGRAPHY NUMBER

Number
TLpq2440376443

LANGUAGE OF THE ITEM

.Language of Text, Soundtrack etc
انگلیسی

TITLE AND STATEMENT OF RESPONSIBILITY

Title Proper
Combining Image and Text Processing for the Computational Reading of Arabic Calligraphy
General Material Designation
[Thesis]
First Statement of Responsibility
Alsalamah, Seetah
Subsequent Statement of Responsibility
Batista-Navarro, Riza Theresa

.PUBLICATION, DISTRIBUTION, ETC

Name of Publisher, Distributor, etc.
The University of Manchester (United Kingdom)
Date of Publication, Distribution, etc.
2020

PHYSICAL DESCRIPTION

Specific Material Designation and Extent of Item
221

DISSERTATION (THESIS) NOTE

Dissertation or thesis details and type of degree
Ph.D.
Body granting the degree
The University of Manchester (United Kingdom)
Text preceding or following the note
2020

SUMMARY OR ABSTRACT

Text of Note
The Arabic language originally made use of ancient calligraphy, which is found in historical documents and the Holy Quran. This calligraphy shows a different representation of Arabic texts using a more cursive style and a mixture of complex constructed word forms. These types of writing styles in Arabic texts give a degree of difficulty in segmenting the letters and reading the text. Since they originated long ago, they have mostly been reflected in Islamic culture and the use of quotations from the Holy Quran. This art form is still used today for various purposes in Arabic representation and Islamic calligraphy. The challenges of this type of text motivate the search for a way to simplify the reading and digitisation processes. To the best of our knowledge, this is the first attempt to investigate the recognition of Arabic calligraphy images and the reading of text drawn in such images. Due to the lack of resources in the calligraphy domain, different datasets were developed for this research. An Arabic calligraphy image dataset was collected, from which calligraphy letter image datasets were generated. Finally, a calligraphy quotations corpus was manually annotated based on the Arabic calligraphy image dataset. All of these datasets were used for training, testing and support in the different phases that were applied to achieve the primary goal of reading calligraphy. A new approach to the recognition of Arabic calligraphy was developed to manipulate a scanned image and extract a list of probable quotations. It consists of comparing two detection methods, namely maximally stable extremal regions (MSER) and sliding window (SW), to obtain the identity of intersecting letters from the image. The letters detected had their features extracted through a comparison between the histogram of oriented gradient (HOG) features and a bag of speeded-up robust features (SURF) used in training two different recognition models, support vector machines (SVM) and the convolutional neural network (CNN). In the investigation into which of these models and image feature descriptors most accurately fit the calligraphy letters, the results from the recognition process were placed in a bag of letters (BOL) feature. This feature was used to search the corpus according to two different methodologies to produce the list of probable quotations. The first method compares the target BOL with the corpus index BOL for each element, while the second method generates a list of related words from the BOL and then searches the corpus for any quotations that contain two or more of these words. The results from reading 388 calligraphy images showed that the MSER method outperforms the SW method in detecting letters. Moreover, BOL matching and searching the corpus predicts more accurate lists of quotations than the word generation process. The best methodology is based on a combination of the SVM recognition model and HOG feature extraction, correctly predicting more than 74% of the top ten quotations using the BOL matching process.

TOPICAL NAME USED AS SUBJECT

Computer science
Linguistics

PERSONAL NAME - PRIMARY RESPONSIBILITY

Alsalamah, Seetah
Batista-Navarro, Riza Theresa

ELECTRONIC LOCATION AND ACCESS

Electronic name
 مطالعه متن کتاب 

p

[Thesis]
276903

a
Y

Proposal/Bug Report

Warning! Enter The Information Carefully
Send Cancel
This website is managed by Dar Al-Hadith Scientific-Cultural Institute and Computer Research Center of Islamic Sciences (also known as Noor)
Libraries are responsible for the validity of information, and the spiritual rights of information are reserved for them
Best Searcher - The 5th Digital Media Festival