Most prior work on information extraction has focused on extracting information from text in digital documents. However, often, the most important information being reported in an...
Abstract— In this work, web-based metrics for semantic similarity computation between words or terms are presented and compared with the state-of-the-art. Starting from the funda...
Noun phrases in queries are identified and classified into four types: proper names, dictionary phrases, simple phrases and complex phrases. A document has a phrase if all content...
Automatically categorizing documents into pre-defined topic hierarchies or taxonomies is a crucial step in knowledge and content management. Standard machine learning techniques ...
To efficiently compress rasterized compound documents, an encoder must be content-adaptive. Content adaptivity may be achieved by using a layered approach. In such an approach, a ...
George Pavlidis, Sofia Tsekeridou, Christodoulos C...