article

Free Access

A stop list for general text

Author:
Christopher Fox

View Profile

Authors Info & Claims

ACM SIGIR Forum Volume 24 Issue 1-2Fall 89/Winter 90pp 19–21https://doi.org/10.1145/378881.378888

Published:01 September 1989Publication History

ACM SIGIR Forum

Abstract

A stop list, or negative dictionary is a device used in automatic indexing to filter out words that would make poor index terms. Traditionally stop lists are supposed to have included only the most frequently occurring words. In practice, however, stop lists have tended to include infrequently occurring words, and have not included many frequently occurring words. Infrequently occurring words seem to have been included because stop list compilers have not, for whatever reason, consulted empirical studies of word frequencies. Frequently occurring words seem to have been left out for the same reason, and also because many of them might still be important as index terms.This paper reports an exercise in generating a stop list for general text based on the Brown corpus of 1,014,000 words drawn from a broad range of literature in English. We start with a list of tokens occurring more than 300 times in the Brown corpus. From this list of 278 words, 32 are culled on the grounds that they are too important as potential index terms. Twenty-six words are then added to the list in the belief that they may occur very frequently in certain kinds of literature. Finally, 149 words are added to the list because the finite state machine based filter in which this list is intended to be used is able to filter them at almost no cost. The final product is a list of 421 stop words that should be maximally efficient and effective in filtering the most frequently occurring and semantically neutral words in general literature in English.

References

van Rijsbergen, C. J., Information Retrieval, Butterworths, 1975. Google ScholarDigital Library
Luhn, H. P., "A Statistical Approach to Mechanized Encoding and Searching of Literary Information," IBM Journal of Research and Development 1(4), October, 1957.Google ScholarDigital Library
Francis, W. Nelson, and Henry, Kucera, Frequency Analysis of English Usage, Houghton Mifflin, 1982.Google Scholar
Aho, Alfred, Ravi Sethi, and Jeffrey Ullman, Compilers: Principles, Techniques, and Tools, Addison-Wesley, 1986. Google ScholarDigital Library

Recommendations

Automatic Construction of Generic Stop Words List for Hindi Text
Abstract
In this technological world, Hindi text is speedily increasing on web and attracting many users and researchers to retrieve useful information from this data. In Information Retrieval (IR) process, user comes across many words of least or no ...
Read More
Automatic construction of Chinese stop word list
ACOS'06: Proceedings of the 5th WSEAS international conference on Applied computer science

In modern information retrieval systems, effective indexing can be achieved by removal of stop words. Till now many stop word lists have been developed for English language. However, no standard stop word list has been constructed for Chinese language ...
Read More
Creation of a Russian Stop Word List
Abstract—
This article describes three identifying characteristics of stop words—statistical, semantic, and morphological—and postulates new principles for the creation of stop word lists based on these characteristics. The application of the principles is ...
Read More

Comments

Login options

Check if you have access through your login credentials or your institution to get full access on this article.

Full Access

Get this Article

Published in
ACM SIGIR Forum Volume 24, Issue 1-2
Fall 89/Winter 90
82 pages
ISSN:0163-5840
DOI:10.1145/378881
Editor:
Christopher Fox
Software Productivity Consortium, Herndon, VA
Issue’s Table of Contents
Copyright © 1989 Author
Sponsors
In-Cooperation
Publisher
Association for Computing Machinery
New York, NY, United States
Publication History
- Published: 1 September 1989
Check for updates
Qualifiers
- article
Conference
Funding Sources
Other Metrics
View Article Metrics

Article Metrics
- 201
  Total Citations
  View Citations
- 2,425
  Total Downloads
- Downloads (Last 12 months)221
- Downloads (Last 6 weeks)15
Other Metrics
View Author Metrics
Cited By
View all

PDF Format

View or Download as a PDF file.

PDF

eReader

View online with eReader.

eReader

A stop list for general text

ACM SIGIR Forum

Abstract

References

Cited By

Recommendations

Automatic Construction of Generic Stop Words List for Hindi Text

Automatic construction of Chinese stop word list

Creation of a Russian Stop Word List

Comments

Login options

Full Access

Published in

Sponsors

In-Cooperation

Publisher

Publication History

Check for updates

Qualifiers

Conference

Funding Sources

Other Metrics

Article Metrics

Other Metrics

Cited By

PDF Format

eReader

Digital Edition

Caption

A stop list for general text

ACM SIGIR Forum

Abstract

References

Cited By

Recommendations

Automatic Construction of Generic Stop Words List for Hindi Text

Automatic construction of Chinese stop word list

Creation of a Russian Stop Word List

Comments

Login options

Full Access

Published in

Sponsors

In-Cooperation

Publisher

Publication History

Check for updates

Qualifiers

Conference

Funding Sources

Article Metrics

Other Metrics

PDF Format

eReader

Digital Edition

Share this Publication link

Share on Social Media