Structural image retrieval using automatic image annotation and region based inverted file
- Authors: Zhang, Dengsheng , Islam, Md , Lu, Guojun
- Date: 2013
- Type: Text , Journal article
- Relation: Journal of Visual Communication and Image Representation Vol. 24, no. 7 (2013), p. 1087-1098
- Full Text: false
- Reviewed:
- Description: Image retrieval has lagged far behind text retrieval despite more than two decades of intensive research effort. Most of the research on image retrieval in the last two decades are on content based image retrieval or image retrieval based on low level features. Recent research in this area focuses on semantic image retrieval using automatic image annotation. Most semantic image retrieval techniques in literature, however, treat an image as a bag of features/words while ignore the structural or spatial information in the image. In this paper, we propose a structural image retrieval method based on automatic image annotation and region based inverted file. In the proposed system, regions in an image are treated the same way as keywords in a structural text document, semantic concepts are learnt from image data to label image regions as keywords and weight is assigned to each keyword according to spatial position and relationship. As the result, images are indexed and retrieved in the same way as structural document retrieval. Specifically, images are broken down to regions which are represented using colour, texture and shape features. Region features are then quantized to create visual dictionaries which are similar to monolingual dictionaries like English or Chinese dictionaries. In the next step, a semantic dictionary similar to a bilingual dictionary like the English–Chinese dictionary is learnt to mapping image regions to semantic concepts. Finally, images are then indexed and retrieved using a novel region based inverted file data structure. Results show the proposed method has significant advantage over the widely used Bayesian annotation models.
Rotation invariant curvelet features for region based image retrieval
- Authors: Zhang, Dengsheng , Islam, Md , Lu, Guojun , Sumana, Ishrat
- Date: 2011
- Type: Text , Journal article
- Relation: International Journal of Computer Vision Vol. 98, no. 2 (2011), p. 187-201
- Full Text: false
- Reviewed:
- Description: There have been much interest and a large amount of research on content based image retrieval (CBIR) in recent years due to the ever increasing number of digital images. Texture features play a key role in CBIR. Many texture features exist in literature, however, most of them are neither rotation invariant nor robust to scale and other variations. Texture features based on Gabor filters have been shown with significant advantages over other methods, and they are adopted by MPEG-7 as one of the texture descriptors for image retrieval. In this paper, we propose a rotation invariant curvelet features for texture representation. With systematic analysis and rigorous experiments, we show that the proposed curvelet texture features significantly outperforms the widely used Gabor texture features. A novel region padding method is also proposed to apply curvelet transform to region based image retrieval. Retrieval results from standard image databases show that curvelet features are promising for both texture and region representation.
On low-rank regularized least squares for scalable nonlinear classification
- Authors: Fu, Zhouyu , Lu, Guojun , Ting, Kaiming , Zhang, Dengsheng
- Date: 2011
- Type: Text , Conference paper
- Relation: International Conference on Neural Information Processing p. 490-499
- Full Text: false
- Reviewed:
- Description: In this paper, we revisited the classical technique of Regularized Least Squares (RLS) for the classification of large-scale nonlinear data. Specifically, we focus on a low-rank formulation of RLS and show that it has linear time complexity in the data size only and does not rely on the number of labels and features for problems with moderate feature dimension. This makes low-rank RLS particularly suitable for classification with large data sets. Moreover, we have proposed a general theorem for the closed-form solutions to the Leave-One-Out Cross Validation (LOOCV) estimation problem in empirical risk minimization which encompasses all types of RLS classifiers as special cases. This eliminates the reliance on cross validation, a computationally expensive process for parameter selection, and greatly accelerate the training process of RLS classifiers. Experimental results on real and synthetic large-scale benchmark data sets have shown that low-rank RLS achieves comparable classification performance while being much more efficient than standard kernel SVM for nonlinear classification. The improvement in efficiency is more evident for data sets with higher dimensions.
Efficient nonlinear classification via low-rank regularised least squares
- Authors: Fu, Zhouyu , Lu, Guojun , Ting, Kaiming , Zhang, Dengsheng
- Date: 2013
- Type: Text , Journal article
- Relation: Neural Computing and Applications Vol. 22, no. 7-8(2013), p. 1279-1289
- Full Text: false
- Reviewed:
- Description: We revisit the classical technique of regularised least squares (RLS) for nonlinear classification in this paper. Specifically, we focus on a low-rank formulation of the RLS, which has linear time complexity in the size of data set only, independent of both the number of classes and number of features. This makes low-rank RLS particularly suitable for problems with large data and moderate feature dimensions. Moreover, we have proposed a general theorem for obtaining the closed-form estimation of prediction values on a holdout validation set given the low-rank RLS classifier trained on the whole training data. It is thus possible to obtain an error estimate for each parameter setting without retraining and greatly accelerate the process of cross-validation for parameter selection. Experimental results on several large-scale benchmark data sets have shown that low-rank RLS achieves comparable classification performance while being much more efficient than standard kernel SVM for nonlinear classification. The improvement in efficiency is more evident for data sets with higher dimensions.
Learning naive Bayes classifiers for music classification and retrieval
- Authors: Fu, Zhouyu , Lu, Guojun , Ting, Kaiming , Zhang, Dengsheng
- Date: 2010
- Type: Text , Conference paper
- Relation: Proceedings of the 20th International Conference on Pattern Recognition p. 4589-4592
- Full Text: false
- Reviewed:
- Description: In this paper, we explore the use of naive Bayes classifiers for music classification and retrieval. The motivation is to employ all audio features extracted from local windows for classification instead of just using a single song-level feature vector produced by compressing the local features. Two variants of naive Bayes classifiers are studied based on the extensions of standard nearest neighbor and support vector machine classifiers. Experimental results have demonstrated superior performance achieved by the proposed naive Bayes classifiers for both music classification and retrieval as compared to the alternative methods.
Semantic image retrieval using region based inverted file
- Authors: Zhang, Dengsheng , Islam, Md , Lu, Guojun , Hou, Jin
- Date: 2009
- Type: Text , Journal article
- Relation: Journal of Visual Communication and Image Representation Vol. 24, no. 7 (2009), p.242-249
- Full Text: false
- Reviewed:
- Description: Image retrieval has lagged far behind text retrieval despite more than two decades of intensive research effort. Most of the research on image retrieval in the last two decades are on content based image retrieval or image retrieval based on low level features. Recent research in this area focuses on semantic image retrieval using automatic image annotation. Most semantic image retrieval techniques in literature, however, treat an image as a bag of features/words while ignore the structural or spatial information in the image. In this paper, we propose a structural image retrieval method based on automatic image annotation and region based inverted file. In the proposed system, regions in an image are treated the same way as keywords in a structural text document, semantic concepts are learnt from image data to label image regions as keywords and weight is assigned to each keyword according to spatial position and relationship. As the result, images are indexed and retrieved in the same way as structural document retrieval. Specifically, images are broken down to regions which are represented using colour, texture and shape features. Region features are then quantized to create visual dictionaries which are similar to monolingual dictionaries like English or Chinese dictionaries. In the next step, a semantic dictionary similar to a bilingual dictionary like the English–Chinese dictionary is learnt to mapping image regions to semantic concepts. Finally, images are then indexed and retrieved using a novel region based inverted file data structure. Results show the proposed method has significant advantage over the widely used Bayesian annotation models.
Spherical harmonics and distance transform for image representation and retrieval
- Authors: Sajjanhar, Atul , Lu, Guojun , Zhang, Dengsheng , Hou, Jingyu , Chen, Yi-Ping Phoebe
- Date: 2009
- Type: Text , Conference paper
- Relation: Proceedings of the Intelligent Data Engineering and Automated Learning p. 309-316
- Full Text: false
- Reviewed:
- Description: In this paper, we have proposed a method for 2D image retrieval based on object shapes. The method relies on transforming the 2D images into 3D space based on distance transform. Spherical harmonics are obtained for the 3D data and used as descriptors for the underlying 2D images. The proposed method is compared against two existing methods which use spherical harmonics for shape based retrieval of images. MPEG-7 Still Images Content Set is used for performing experiments; this dataset consists of 3621 still images. Experimental results show that the performance of the proposed descriptors is significantly better than other methods in the same category.
Rotation invariant curvelet features for texture image retrieval
- Authors: Islam, Md , Zhang, Dengsheng , Lu, Guojun
- Date: 2009
- Type: Text , Conference paper
- Relation: Proceedings of the 2009 IEEE International Conference on Multimedia and Expo p. 562-565
- Full Text: false
- Reviewed:
- Description: Effective texture feature is an essential component in any content based image retrieval system. In the past, spectral features, like Gabor and wavelet, have shown superior retrieval performance than many other statistical and structural based features. Recent researches on multi-resolution analysis have found that curvelet captures texture properties, like curves, lines, and edges, more accurately than Gabor filters. However, the texture feature extracted using curvelet transform is not rotation invariant. This can degrade its retrieval performance significantly, especially in cases where there are many similar images with different orientations. This paper analyses the curvelet transform and derives a useful approach to extract rotation invariant curvelet features. Experimental results show that the new rotation invariant curvelet feature outperforms the curvelet feature without rotation invariance.
Connectivity-based shape descriptors
- Authors: Sajjanhar, Atul , Lu, Guojun , Zhang, Dengsheng , Zhou, Wanle
- Date: 2010
- Type: Text , Journal article
- Relation: International Journal of Computers and Applications Vol. 32, no. 1 (2010), p. 93-98
- Full Text: false
- Reviewed:
- Description: In this paper, we propose a method for indexing and retrieval of images based on shapes of objects. The concept of connectivity is introduced. 3D models are used to represent 2D images. 2D images are decomposed a priori using connectivity which is followed by 3D model construction. 3D model descriptors are obtained for 3D models and used to represent the underlying 2D shapes. We have used spherical harmonics descriptors as the 3D model descriptors. Difference between two images is computed as the Euclidean distance between their descriptors. Experiments are performed to test the effectiveness of spherical harmonics for retrieval of 2D images. The proposed method is compared with methods based on principal components analysis (PCA) and generic Fourier descriptors (GFD). It is found that the proposed method is effective. Item S8 within the MPEG-7 still images content set is used for performing experiments.
Image retrieval based on semantics of intra-region color properties
- Authors: Sajjanhar, Atul , Lu, Guojun , Zhang, Dengsheng , Zhou, Wanlei , Chen, Yi-Ping Phoebe
- Date: 2008
- Type: Text , Conference paper
- Relation: Proceedings of 2008 IEEE 8th International Conference on Computer and Information Technology p. 338-343
- Full Text: false
- Reviewed:
- Description: Traditional image retrieval systems are content based image retrieval systems which rely on low-level features for indexing and retrieval of images. CBIR systems fail to meet user expectations because of the gap between the low level features used by such systems and the high level perception of images by humans. Semantics based methods have been used to describe images according to their high level features. In this paper, we performed experiments to identify the failure of existing semantics-based methods to retrieve images in a particular semantic category. We have proposed a new semantic category to describe the intra-region color feature. The proposed semantic category complements the existing high level descriptions. Experimental results confirm the effectiveness of the proposed method
Automatic image search based on improved feature descriptors and decision tree
- Authors: Hou, Jin , Chen, Zeng , Qin, Xue , Zhang, Dengsheng
- Date: 2011
- Type: Text , Journal article
- Relation: Integrated Computer-Aided Engineering Vol. 18, no. 2 (2011), p. 167-180
- Full Text: false
- Reviewed:
- Description: There has been a growing interest in implementing image search engine at the semantic level. However, most existing practical systems including popular commercial image search engines like Google and Yahoo! are either text-based or a simple hybrid of texts and visual features. This paper proposes a novel image search system based on automatic image annotation. We develop a technology which learns semantic image concepts from image contents and transforms unstructured images into textual documents, so that images are indexed and retrieved in the same way as textual documents. Existing database management systems can be used to effectively manage image contents, and image search can be as efficient as text search by transforming images into textual documents through machine learning. Experiments in both the Corel dataset and real Web dataset are performed to validate our system and the results are promising. This system suggests a new combination of texts and visual features in order to achieve a semantic image search, and is expected to become a re-ranking system to the existing image search result via the Internet.
On feature combination for music classification
- Authors: Fu, Zhouyu , Lu, Guojun , Ting, Kaiming , Zhang, Dengsheng
- Date: 2010
- Type: Text , Conference proceedings
- Full Text: false
A geometric method to compute directionality features for texture images
- Authors: Islam, Md , Zhang, Dengsheng , Lu, Guojun
- Date: 2008
- Type: Text , Conference paper
- Relation: Proceedings of the 2008 IEEE International Conference on Multimedia and Expo p. 1521-1524
- Full Text: false
- Reviewed:
- Description: In content based image analysis and retrieval, texture feature is an essential component due to its strong discriminative power. Directionality is one of the most significant texture features which are well perceived by the human visual system. A new method to calculate the directionality of image is proposed in this paper. In contrast to Tamura method which uses the statistical property of the directional histogram of an image to calculate its directionality, the proposed method makes use of the geometric property of the directional histogram. Both subjective and objective analyses prove that the proposed method outperforms the conventional Tamura method. It has also been shown that the proposed directionality has better retrieval performance than the conventional Tamura directionality.
A review on automatic image annotation techniques
- Authors: Zhang, Dengsheng , Islam, Md , Lu, Guojun
- Date: 2012
- Type: Text , Journal article
- Relation: Pattern Recognition Letters Vol. 45, no. 1 (2012), p. 346-362
- Full Text: false
- Reviewed:
- Description: Nowadays, more and more images are available. However, to find a required image for an ordinary user is a challenging task. Large amount of researches on image retrieval have been carried out in the past two decades. Traditionally, research in this area focuses on content based image retrieval. However, recent research shows that there is a semantic gap between content based image retrieval and image semantics understandable by humans. As a result, research in this area has shifted to bridge the semantic gap between low level image features and high level semantics. The typical method of bridging the semantic gap is through the automatic image annotation (AIA) which extracts semantic features using machine learning techniques. In this paper, we focus on this latest development in image retrieval and provide a comprehensive survey on automatic image annotation. We analyse key aspects of the various AIA methods, including both feature extraction and semantic learning methods. Major methods are discussed and illustrated in details. We report our findings and provide future research directions in the AIA area in the conclusions
A survey of audio-based music classification and annotation
- Authors: Fu, Zhouyu , Lu, Guojun , Ting, Kaiming , Zhang, Dengsheng
- Date: 2011
- Type: Text , Journal article
- Relation: IEEE Transactions on Multimedia Vol. 13, no. 2 (2011), p. 303-319
- Full Text: false
- Reviewed:
- Description: Music information retrieval (MIR) is an emerging research area that receives growing attention from both the research community and music industry. It addresses the problem of querying and retrieving certain types of music from large music data set. Classification is a fundamental problem in MIR. Many tasks in MIR can be naturally cast in a classification setting, such as genre classification, mood classification, artist recognition, instrument recognition, etc. Music annotation, a new research area in MIR that has attracted much attention in recent years, is also a classification problem in the general sense. Due to the importance of music classification in MIR research, rapid development of new methods, and lack of review papers on recent progress of the field, we provide a comprehensive review on audio-based classification in this paper and systematically summarize the state-of-the-art techniques for music classification. Specifically, we have stressed the difference in the features and the types of classifiers used for different classification tasks. This survey emphasizes on recent development of the techniques and discusses several open issues for future research.
Region-based image retrieval with high-level semantics using decision tree learning
- Authors: Liu, Ying , Zhang, Dengsheng , Lu, Guojun
- Date: 2008
- Type: Text , Journal article
- Relation: Pattern Recognition Vol. 41, no. 8 (2008), p. 2554-2570
- Full Text: false
- Reviewed:
- Description: Semantic-based image retrieval has attracted great interest in recent years. This paper proposes a region-based image retrieval system with high-level semantic learning. The key features of the system are: (1) it supports both query by keyword and query by region of interest. The system segments an image into different regions and extracts low-level features of each region. From these features, high-level concepts are obtained using a proposed decision tree-based learning algorithm named DT-ST. During retrieval, a set of images whose semantic concept matches the query is returned. Experiments on a standard real-world image database confirm that the proposed system significantly improves the retrieval performance, compared with a conventional content-based image retrieval system. (2) The proposed decision tree induction method DT-ST for image semantic learning is different from other decision tree induction algorithms in that it makes use of the semantic templates to discretize continuous-valued region features and avoids the difficult image feature discretization problem. Furthermore, it introduces a hybrid tree simplification method to handle the noise and tree fragmentation problems, thereby improving the classification performance of the tree. Experimental results indicate that DT-ST outperforms two well-established decision tree induction algorithms ID3 and C4.5 in image semantic learning.
Building sparse support vector machines for multi-instance classification
- Authors: Fu, Zhouyu , Lu, Guojun , Ting, Kaiming , Zhang, Dengsheng
- Date: 2011
- Type: Text , Conference paper
- Relation: European Conference on Machine Learning Knowledge Discovery in Databases (ECML PKDD) p. 471-486
- Full Text: false
- Reviewed:
- Description: We propose a direct approach to learning sparse Support Vector Machine (SVM) prediction models for Multi-Instance (MI) classification. The proposed sparse SVM is based on a “label-mean” formulation of MI classification which takes the average of predictions of individual instances for bag-level prediction. This leads to a convex optimization problem, which is essential for the tractability of the optimization problem arising from the sparse SVM formulation we derived subsequently, as well as the validity of the optimization strategy we employed to solve it. Based on the “label-mean” formulation, we can build sparse SVM models for MI classification and explicitly control their sparsities by enforcing the maximum number of expansions allowed in the prediction function. An effective optimization strategy is adopted to solve the formulated sparse learning problem which involves the learning of both the classifier and the expansion vectors. Experimental results on benchmark data sets have demonstrated that the proposed approach is effective in building very sparse SVM models while achieving comparable performance to the state-of-the-art MI classifiers.
An annotation rule extraction algorithm for image retrieval
- Authors: Chen, Zeng , Hou, Jin , Zhang, Dengsheng , Qin, Xue
- Date: 2012
- Type: Text , Journal article
- Relation: Pattern Recognition Letters Vol. 33, no. 10 (2012), p.1257-1268
- Full Text: false
- Reviewed:
- Description: Automatic image annotation can be used to facilitate semantic search in large image databases. However, retrieval performance of the existing annotation schemes is far from the users’ expectation. In this paper, we propose a novel method to automatically annotate image through the rules generated by support vector machines and decision trees. In order to obtain the rules, we collect a set of training regions by image segmentation, feature extraction and discretization. We first employ a support vector machine as a preprocessing technique to refine the input training data and then use it to improve the rules generated by decision tree learning. The preprocessing can effectively deal with the similar regions in an image as well. Moreover, we integrate the original rules to the modified ones, so as to formulate the complete and effective annotation rules. We can translate an unknown image into text by this algorithm, and the proposed system can retrieve images queried by both images and keywords. Experiments are carried out in a standard Corel dataset and images collected from the Web to test the accuracy and robustness of the proposed system. Experimental results show the proposed algorithm can annotate and retrieve images more efficiently than traditional learning algorithms.
Automatic image annotation based on decision tree machine learning
- Authors: Jiang, Lixing , Hou, Jin , Zeng, Chen , Zhang, Dengsheng
- Date: 2009
- Type: Text , Conference paper
- Relation: Proceedings of the International Conference on Cyber-Enabled Distributed Computing and Knowledge Discovery p. 170-175
- Full Text: false
- Reviewed:
- Description: With the rapid development of digital imaging technology, image annotation is an important and challenging task in image retrieval. At present, many machine learning methods have been applied to solve the problem of automatic image annotation (AIA). However, there exists enormous semantic expressive gap between the low-level image features and high-level semantic concepts. Due to the problem, the annotation performance of existing methods is not satisfactory, and needs to be further improved. This paper proposes an automatic annotation framework via a novel decision tree-based Bayesian (DTB) machine learning algorithm. It is a hybrid approach that attempts to utilize the advantages of both DT and Naive-Bayesian (NB). We firstly segment an image into different regions and extract low-level features of each region. From these features, high-level semantic concepts are obtained using a DTB learning algorithm. Finally, experiments conducted on the Corel dataset demonstrate the effectiveness of DTB machine learning. The DTB can not only enhance the classification accuracy, but also associate low-level region features with high-level image concepts. This method presents the advantages of the Bayesian method and the DT. Moreover, this semantic interpretation capability is a natural simulation of human learning.
Composite feature modeling and retrieval
- Authors: Hou, Jin , Zhang, Dengsheng , Chen, Zeng , Xu, Xuerong , Nakamura, Takahiro
- Date: 2008
- Type: Text , Conference paper
- Relation: Proceedings of the 2008 10th International Conference on Control, Automation, Robotics & Vision p. 2176-2181
- Full Text: false
- Reviewed:
- Description: Feature-based intelligent design and manufacturing systems in the Internet environment are an evolution of traditional geometric and solid modeling systems. This paper presents some novel algorithms including a new face-base representation, composite feature modeling and retrieval technology, and efficient communication mechanism, to construct an interactive framework for composite feature modeling and retrieval. The proposed system consists of a feature modeler developed on Wolfram Research Mathematica, Java and Java 3D enabled GUI (graphical user interface), and DB (database). Experiments demonstrate that this system reflects designers' intent properly and is user-friendly to experts coming from various technical backgrounds. This paper provides some fundamental principles for composite feature modeling and retrieval in web-based distributed environment.