JOURNAL ARTICLE

An asymptotically optimal data compression algorithm based on an inverted index

Abstract

Summary form only given. An alternate approach to representing a data sequence is to associate with each source letter, the list of locations at which it appears in the data sequence. We present a data compression algorithm based on a generalization of this idea. The algorithm parses the data with respect to a static dictionary of phrases and associates with each phrase in the dictionary a list of locations at which the phrase appears in the parsed data. Each list of locations is then run-length encoded. This collection of run-length encoded lists constitutes the compressed representation of the data. We refer to the collection of lists as an inverted index. While in information retrieval systems, the inverted index is an adjunct to the main database used to speed up searching, we regard it here as a self-contained representation of the database itself. Further, our inverted index does not necessarily list every occurrence of a phrase in the data, only every occurrence in the parsing. This allows us to be asymptotically optimal in terms of compression, though at the cost of a loss in searching efficiency. We discuss this trade-off between compression and searching efficiency. We prove that in terms of compression, this algorithm is asymptotically optimal universally over the class of discrete memoryless sources. We also show that pattern matching can be performed efficiently in the compressed domain. Compressing and storing data in this manner may be useful in applications which require frequent searching of a large but mostly static database.

Keywords:
Inverted index Computer science Data compression Parsing Algorithm Phrase Asymptotically optimal algorithm Generalization Sequence (biology) Index (typography) Compression (physics) Data structure Theoretical computer science Mathematics Search engine indexing Artificial intelligence

Metrics

1
Cited By
0.00
FWCI (Field Weighted Citation Impact)
2
Refs
0.12
Citation Normalized Percentile
Is in top 1%
Is in top 10%

Topics

Algorithms and Data Compression
Physical Sciences →  Computer Science →  Artificial Intelligence
semigroups and automata theory
Physical Sciences →  Computer Science →  Computational Theory and Mathematics
DNA and Biological Computing
Life Sciences →  Biochemistry, Genetics and Molecular Biology →  Molecular Biology

Related Documents

JOURNAL ARTICLE

Optimal data compression algorithm

I. Sadeh

Journal:   Computers & Mathematics with Applications Year: 1996 Vol: 32 (5)Pages: 57-72
BOOK-CHAPTER

Inverted Index Compression

Giulio Ermanno PibiriRossano Venturini

Encyclopedia of Big Data Technologies Year: 2018 Pages: 1-8
BOOK-CHAPTER

Inverted Index Compression

Giulio Ermanno PibiriRossano Venturini

Encyclopedia of Big Data Technologies Year: 2012 Pages: 1-9
BOOK-CHAPTER

Inverted Index Compression

Giulio Ermanno PibiriRossano Venturini

Encyclopedia of Big Data Technologies Year: 2019 Pages: 1051-1058
© 2026 ScienceGate Book Chapters — All rights reserved.