The Springer International Series in Engineering and Computer Science

Explorations in Automatic Thesaurus Discovery

Authors: Grefenstette, Gregory

Buy this book

eBook 107,09 €
price for Spain (gross)
  • ISBN 978-1-4615-2710-7
  • Digitally watermarked, DRM-free
  • Included format: PDF
  • ebooks can be used on all reading devices
  • Immediate eBook download after purchase
Hardcover 171,59 €
price for Spain (gross)
  • ISBN 978-0-7923-9468-6
  • Free shipping for individuals worldwide
  • Usually dispatched within 3 to 5 business days.
  • The final prices may differ from the prices shown due to specifics of VAT rules
Softcover 135,19 €
price for Spain (gross)
  • ISBN 978-1-4613-6167-1
  • Free shipping for individuals worldwide
  • Usually dispatched within 3 to 5 business days.
  • The final prices may differ from the prices shown due to specifics of VAT rules
About this book

Explorations in Automatic Thesaurus Discovery presents an automated method for creating a first-draft thesaurus from raw text. It describes natural processing steps of tokenization, surface syntactic analysis, and syntactic attribute extraction. From these attributes, word and term similarity is calculated and a thesaurus is created showing important common terms and their relation to each other, common verb--noun pairings, common expressions, and word family members.
The techniques are tested on twenty different corpora ranging from baseball newsgroups, assassination archives, medical X-ray reports, abstracts on AIDS, to encyclopedia articles on animals, even on the text of the book itself. The corpora range from 40,000 to 6 million characters of text, and results are presented for each in the Appendix.
The methods described in the book have undergone extensive evaluation. Their time and space complexity are shown to be modest. The results are shown to converge to a stable state as the corpus grows. The similarities calculated are compared to those produced by psychological testing. A method of evaluation using Artificial Synonyms is tested. Gold Standards evaluation show that techniques significantly outperform non-linguistic-based techniques for the most important words in corpora.
Explorations in Automatic Thesaurus Discovery includes applications to the fields of information retrieval using established testbeds, existing thesaural enrichment, semantic analysis. Also included are applications showing how to create, implement, and test a first-draft thesaurus.

Table of contents (6 chapters)

Buy this book

eBook 107,09 €
price for Spain (gross)
  • ISBN 978-1-4615-2710-7
  • Digitally watermarked, DRM-free
  • Included format: PDF
  • ebooks can be used on all reading devices
  • Immediate eBook download after purchase
Hardcover 171,59 €
price for Spain (gross)
  • ISBN 978-0-7923-9468-6
  • Free shipping for individuals worldwide
  • Usually dispatched within 3 to 5 business days.
  • The final prices may differ from the prices shown due to specifics of VAT rules
Softcover 135,19 €
price for Spain (gross)
  • ISBN 978-1-4613-6167-1
  • Free shipping for individuals worldwide
  • Usually dispatched within 3 to 5 business days.
  • The final prices may differ from the prices shown due to specifics of VAT rules
Loading...

Recommended for you

Loading...

Bibliographic Information

Bibliographic Information
Book Title
Explorations in Automatic Thesaurus Discovery
Authors
Series Title
The Springer International Series in Engineering and Computer Science
Series Volume
278
Copyright
1994
Publisher
Springer US
Copyright Holder
Springer Science+Business Media New York
eBook ISBN
978-1-4615-2710-7
DOI
10.1007/978-1-4615-2710-7
Hardcover ISBN
978-0-7923-9468-6
Softcover ISBN
978-1-4613-6167-1
Series ISSN
0893-3405
Edition Number
1
Number of Pages
XIII, 305
Topics