+612 9045 4394
 
CHECKOUT
Explorations in Automatic Thesaurus Discovery : The Springer International Series in Engineering and Computer Science - Gregory Grefenstette

Explorations in Automatic Thesaurus Discovery

The Springer International Series in Engineering and Computer Science

Hardcover Published: 31st July 1994
ISBN: 9780792394686
Number Of Pages: 305

Share This Book:

Hardcover

RRP $562.99
$389.50
31%
OFF
or 4 easy payments of $97.38 with Learn more
Ships in 7 to 10 business days

Other Available Editions (Hide)

  • Paperback View Product Published: 21st November 2012
    $227.94

Explorations in Automatic Thesaurus Discovery presents an automated method for creating a first-draft thesaurus from raw text. It describes natural processing steps of tokenization, surface syntactic analysis, and syntactic attribute extraction. From these attributes, word and term similarity is calculated and a thesaurus is created showing important common terms and their relation to each other, common verb--noun pairings, common expressions, and word family members.
The techniques are tested on twenty different corpora ranging from baseball newsgroups, assassination archives, medical X-ray reports, abstracts on AIDS, to encyclopedia articles on animals, even on the text of the book itself. The corpora range from 40,000 to 6 million characters of text, and results are presented for each in the Appendix.
The methods described in the book have undergone extensive evaluation. Their time and space complexity are shown to be modest. The results are shown to converge to a stable state as the corpus grows. The similarities calculated are compared to those produced by psychological testing. A method of evaluation using Artificial Synonyms is tested. Gold Standards evaluation show that techniques significantly outperform non-linguistic-based techniques for the most important words in corpora.
Explorations in Automatic Thesaurus Discovery includes applications to the fields of information retrieval using established testbeds, existing thesaural enrichment, semantic analysis. Also included are applications showing how to create, implement, and test a first-draft thesaurus.

Preface
Introductionp. 1
Semantic Extractionp. 7
Sextantp. 33
Evaluationp. 69
Applicationsp. 101
Conclusionp. 137
Preprocessorsp. 149
Webster Stopword Listp. 151
Similarity Listp. 153
Semantic Clusteringp. 163
Automatic Thesaurus Generationp. 171
Corpora Treatedp. 181
Indexp. 303
Table of Contents provided by Blackwell. All Rights Reserved.

ISBN: 9780792394686
ISBN-10: 0792394682
Series: The Springer International Series in Engineering and Computer Science
Audience: Professional
Format: Hardcover
Language: English
Number Of Pages: 305
Published: 31st July 1994
Publisher: Springer
Country of Publication: NL
Dimensions (cm): 23.5 x 15.5  x 2.69
Weight (kg): 1.39