Thelast paragraph was actually about handwritten or as you mentioned Nastaligh fonts. Which there are a very few of them available right now that would create a good and natural form. These fonts have a lot of glyphs and hundreds lines of scripts. For example a few years ago we had a designer in Iran who created the IranNastaligh font funded by the government and it was something like 1,3 MB. And yet you could encounter a lot of dots and letters overlapping. I have to say he did a nice job, but it bothers me somehow that after all these years the only way to write with a Nastaligh or Tahriri or Persian handwritten style which based on Nastaligh and Tahriri is using some apps or software to correct the position of dots and letters.
If you are familiar with a font that does not have that problem I really appreciated if you introduce it to me. And could we use contextual variant to solve the many many ligatures that are needed for a handwritten font?
They don't like the idea of seeing one and one only non-cursive Persian font. The font that tries to make exactly similar letters for both Persian/Arabic and Latin. And when they don't make a distinction between a script and a font what can you say? I have never seen anything like this in a Latin type community.
Arabic is the one of the fifth spoken language across the globe (422 million speak Arabic). Accurate recognition of Arabic characters is a challenging work because of two reasons; firstly, major dissimilarities and minor similarities with other languages and secondly right to left writing style. Hence, good quality features are to be extracted and better classifiers are required to develop the optical character recognition of Arabic characters/font. This manuscript explores automatic and appropriate selection of classifiers and features sets for recognition of character in Arabic handwritten script. After analyzing the problem, this manuscript investigates a new model which is concentrated on fusion strategy. To achieve this objective, fusion of various classifiers and features has been done. The suggested work is on multi-font script recognition that contains fonts of SH Roqa, Naskh, Farsi, and Igaza. To perform experimentation, different classifiers and features which are appropriate to the script recognition for Arabic font nature have been considered and implemented. The suggested work has been executed on AHCD and IFN/ENIT datasets. The experimentation results indicate that fusion strategies bring results better than traditional methods.
Recognition of script in languages like English, Latin and Chinese have sufficient literature review however previous studies shows that recognition of Arabic text in Arabic document is still insufficient works. Thus, Arabic character recognition is considered as more complicated due to different reasons. First, different styles of writing which makes character and text line recognition very intricate. Second, huge number of characters with their different shapes in different location and large character sets with high similarity of characters all makes the task of the off-line scripts recognition is not easy. Thus, a good system of recognition should be dealing with various styles of handwriting, various font sizes, and connected characters. Document images require to be preprocessed to remove noise. Then line, word, and characters to be segmented. After segmentation, discriminate features need to be identified, and identifying individual classifiers with the use of classifiers.
Because cursive nature of Arabic script, the improvement of Arabic optical character recognition systems which denoted by OCR includes numerous technical issues, particularly in the phase of extraction of features and classification. Most Arabic characters have madda, zigzag(s), and dot(s) related with the characters and these could be inside, below, or on top of the character. Several characters have the same form only the dots number, secondary strokes number, or position makes them distinct. Though several researchers are discussing and examining solutions for take care of these issues. In this direction little advancement has been done.
There is no best pattern recognition method investigated for recognition of Arabic handwritten script so far. Every methodology has some kind of advantages, disadvantages and limitations. To overcome challenges addressed earlier, the method need to take benefit out of variety of strengths from existed methods to construct multiple data based frame work in addition to the novelty.
The hybrid methodologies are formed by systems for recognition which utilizes dataset at classification level by combining at least two paradigms of complementary classification and a combination function for decision which combine the classifier outcome to produce a single decision. Or the combination at feature extraction stage combining at least two primitives types to get better input character description, or at each stage of classification and feature extraction.
This work aims at evaluation of methods by considering set of four features and three classifiers is the first aim of this manuscript. The features which are used in this work are moments invariants (MI), run length matrix (RLM), statistical properties of intensity histogram (SFIH) and wavelet decomposition (WD). The classifiers which are used in this proposed work are modified quadratic discriminate functions (MQDF), support vector machine (SVM) and random forest (RF). To study the fusion effect of features, classifiers and accuracy rate of Arabic character recognition by choosing different combinations of classifiers and features is the second aim of this manuscript.
In any character recognition systems, segmentation phase is the significant task especially in Arabic handwritten recognition. A new text recognition method has been presented in [1], which depends on analysis of a subset of documents. There are three main sub-set domains. In order to avert failure in detection, the skeletonization method, this did the correction in the situation of false alarms into building the segmentation of vertically connected characters. Accuracy rate achieved is 98.9 and 97.4 on IFN/ENIT and IFN/ENIT datasets respectively which is considered high.
A new method has been suggested in [3] for locating the segmentation points. It used word image thinning to fetch the width of the stroke of one pixel. To detect the ligatures of Arabic characters, geometry and shape are utilized in the segmentation procedure. The approach of the proposed work is worked with touching characters, in both case of ligatures touching the segmentation existing between the characters of consecutive closed and the existence of ligatures within characters in the state of open characters.
Defined three basic approaches for word segmentation which are holistic, internal segmentation and external segmentation have been presented in [4]. In holistic segmentation approach also called no segmentation in which generic entire word features are utilized for the purpose of recognition. So, no characters are segmented in this method for recognition. In the internal segmentation, characters segmentation and recognition have been followed. The approach commonly used in the segmentation is the external segmentation which is prior to the recognition.
Another method which depended on mathematical morphology namely regularities and singularities has been presented in [5]. The segmentation candidates are the regularities. By applying an opening operation to the word image of the Arabic handwritten, authors specified singularities along with regularities by subtracting the singularities from the original document images. Segmentation works for overlapping of characters, the accuracy rate that achieved in the algorithm was 81.88%.
Methods are applied to achieve a good accuracy rate using Hough Transform with skeletonization for Arabic text recognition and modeling framework of Gaussian mixture method has been suggested in [7] that cover possible false correction to create proficiency in vertical connected characters segmentation. The authors used Euclidean distance metrics and convex fusion to calculate the distance among neighboring overlap components, which is classified as an intra-word distance or an inter-word distance.
A model of the Gaussian mixture has been suggested in [8] for identification at word level. An endeavor has been done to examine and compare various classifiers utilizing various statistical tests has been presented in [9]. Authors prescribed to utilize 39 particular features relying upon the convex hull and topological features. From eight different classifiers with multilayer perceptron (MLP) accomplished 99.87% accuracy rate. The directional discrete cosine transform (DDCT) has been exploited in [10] where KNN and LDA are utilized for classification.
A recognition method for handwritten documents has been proposed in [11]. In this method extracted five features of the connected components namely, number of holes, relative X centroid, aspect ratio, sphericity, and relative Y centroid. After that computed skew of all the features, standard deviation and mean were computed and utilized LDA for script recognition. Different techniques of feature extraction used for recognition of handwritten script have been reviewed in [12]. Detailed review on recognition of script from multi-script has been explained in [13].
This manuscript concentrates on combination of features and fusion of classifiers, proposed a system as mentioned in Fig. 1 to recognize handwritten characters of Arabic script with multi fonts. This system integrates numerous features and classifiers to increase the recognition rate of classifiers.
Basically, the suggested work focuses on the character level recognition. The suggested model mimics the data recognition done by human during analysis from various viewpoints. To automate the script recognition, the fusion strategies concept is utilized. Fusion technique utilized to determine Arabic script type followed by recognizing the character. The classifiers and features are fused to arrive at final decision.
3a8082e126