org.apache.lucene.index
Class IndexReader

java.lang.Object
  extended byorg.apache.lucene.index.IndexReader
Direct Known Subclasses:
FilterIndexReader, MultiReader

public abstract class IndexReader
extends java.lang.Object

IndexReader is an abstract class, providing an interface for accessing an index. Search of an index is done entirely through this abstract interface, so that any subclass which implements it is searchable.

Concrete subclasses of IndexReader are usually constructed with a call to the static method open(java.lang.String).

For efficiency, in this API documents are often referred to via document numbers, non-negative integers which each name a unique document in the index. These document numbers are ephemeral--they may change as documents are added to and deleted from an index. Clients should thus not rely on a given document having the same number between sessions.

Version:
$Id: IndexReader.java,v 1.32 2004/04/21 16:46:30 goller Exp $
Author:
Doug Cutting

Constructor Summary
protected IndexReader(Directory directory)
          Constructor used if IndexReader is not owner of its directory.
 
Method Summary
 void close()
          Closes files associated with this index.
protected  void commit()
          Commit changes resulting from delete, undeleteAll, or setNorm operations
 void delete(int docNum)
          Deletes the document numbered docNum.
 int delete(Term term)
          Deletes all documents containing term.
 Directory directory()
          Returns the directory this index resides in.
abstract  int docFreq(Term t)
          Returns the number of documents containing the term t.
protected abstract  void doClose()
          Implements close.
protected abstract  void doCommit()
          Implements commit.
abstract  Document document(int n)
          Returns the stored fields of the nth Document in this index.
protected abstract  void doDelete(int docNum)
          Implements deletion of the document numbered docNum.
protected abstract  void doSetNorm(int doc, java.lang.String field, byte value)
          Implements setNorm in subclass.
protected abstract  void doUndeleteAll()
          Implements actual undeleteAll() in subclass.
protected  void finalize()
          Release the write lock, if needed.
static long getCurrentVersion(Directory directory)
          Reads version number from segments files.
static long getCurrentVersion(java.io.File directory)
          Reads version number from segments files.
static long getCurrentVersion(java.lang.String directory)
          Reads version number from segments files.
abstract  java.util.Collection getFieldNames()
          Returns a list of all unique field names that exist in the index pointed to by this IndexReader.
abstract  java.util.Collection getFieldNames(boolean indexed)
          Returns a list of all unique field names that exist in the index pointed to by this IndexReader.
abstract  java.util.Collection getIndexedFieldNames(boolean storedTermVector)
           
abstract  TermFreqVector getTermFreqVector(int docNumber, java.lang.String field)
          Return a term frequency vector for the specified document and field.
abstract  TermFreqVector[] getTermFreqVectors(int docNumber)
          Return an array of term frequency vectors for the specified document.
abstract  boolean hasDeletions()
          Returns true if any documents have been deleted
static boolean indexExists(Directory directory)
          Returns true if an index exists at the specified directory.
static boolean indexExists(java.io.File directory)
          Returns true if an index exists at the specified directory.
static boolean indexExists(java.lang.String directory)
          Returns true if an index exists at the specified directory.
abstract  boolean isDeleted(int n)
          Returns true if document n has been deleted
static boolean isLocked(Directory directory)
          Returns true iff the index in the named directory is currently locked.
static boolean isLocked(java.lang.String directory)
          Returns true iff the index in the named directory is currently locked.
static long lastModified(Directory directory)
          Deprecated. Replaced by getCurrentVersion(Directory)
static long lastModified(java.io.File directory)
          Deprecated. Replaced by getCurrentVersion(File)
static long lastModified(java.lang.String directory)
          Deprecated. Replaced by getCurrentVersion(String)
abstract  int maxDoc()
          Returns one greater than the largest possible document number.
abstract  byte[] norms(java.lang.String field)
          Returns the byte-encoded normalization factor for the named field of every document.
abstract  void norms(java.lang.String field, byte[] bytes, int offset)
          Reads the byte-encoded normalization factor for the named field of every document.
abstract  int numDocs()
          Returns the number of documents in this index.
static IndexReader open(Directory directory)
          Returns an IndexReader reading the index in the given Directory.
static IndexReader open(java.io.File path)
          Returns an IndexReader reading the index in an FSDirectory in the named path.
static IndexReader open(java.lang.String path)
          Returns an IndexReader reading the index in an FSDirectory in the named path.
 void setNorm(int doc, java.lang.String field, byte value)
          Expert: Resets the normalization factor for the named field of the named document.
 void setNorm(int doc, java.lang.String field, float value)
          Expert: Resets the normalization factor for the named field of the named document.
abstract  TermDocs termDocs()
          Returns an unpositioned TermDocs enumerator.
 TermDocs termDocs(Term term)
          Returns an enumeration of all the documents which contain term.
abstract  TermPositions termPositions()
          Returns an unpositioned TermPositions enumerator.
 TermPositions termPositions(Term term)
          Returns an enumeration of all the documents which contain term.
abstract  TermEnum terms()
          Returns an enumeration of all the terms in the index.
abstract  TermEnum terms(Term t)
          Returns an enumeration of all terms after a given term.
 void undeleteAll()
          Undeletes all documents currently marked as deleted in this index.
static void unlock(Directory directory)
          Forcibly unlocks the index in the named directory.
 
Methods inherited from class java.lang.Object
clone, equals, getClass, hashCode, notify, notifyAll, toString, wait, wait, wait
 

Constructor Detail

IndexReader

protected IndexReader(Directory directory)
Constructor used if IndexReader is not owner of its directory. This is used for IndexReaders that are used within other IndexReaders that take care or locking directories.

Parameters:
directory - Directory where IndexReader files reside.
Method Detail

open

public static IndexReader open(java.lang.String path)
                        throws java.io.IOException
Returns an IndexReader reading the index in an FSDirectory in the named path.

Throws:
java.io.IOException

open

public static IndexReader open(java.io.File path)
                        throws java.io.IOException
Returns an IndexReader reading the index in an FSDirectory in the named path.

Throws:
java.io.IOException

open

public static IndexReader open(Directory directory)
                        throws java.io.IOException
Returns an IndexReader reading the index in the given Directory.

Throws:
java.io.IOException

directory

public Directory directory()
Returns the directory this index resides in.


lastModified

public static long lastModified(java.lang.String directory)
                         throws java.io.IOException
Deprecated. Replaced by getCurrentVersion(String)

Returns the time the index in the named directory was last modified.

Synchronization of IndexReader and IndexWriter instances is no longer done via time stamps of the segments file since the time resolution depends on the hardware platform. Instead, a version number is maintained within the segments file, which is incremented everytime when the index is changed.

Throws:
java.io.IOException

lastModified

public static long lastModified(java.io.File directory)
                         throws java.io.IOException
Deprecated. Replaced by getCurrentVersion(File)

Returns the time the index in the named directory was last modified.

Synchronization of IndexReader and IndexWriter instances is no longer done via time stamps of the segments file since the time resolution depends on the hardware platform. Instead, a version number is maintained within the segments file, which is incremented everytime when the index is changed.

Throws:
java.io.IOException

lastModified

public static long lastModified(Directory directory)
                         throws java.io.IOException
Deprecated. Replaced by getCurrentVersion(Directory)

Returns the time the index in the named directory was last modified.

Synchronization of IndexReader and IndexWriter instances is no longer done via time stamps of the segments file since the time resolution depends on the hardware platform. Instead, a version number is maintained within the segments file, which is incremented everytime when the index is changed.

Throws:
java.io.IOException

getCurrentVersion

public static long getCurrentVersion(java.lang.String directory)
                              throws java.io.IOException
Reads version number from segments files. The version number counts the number of changes of the index.

Parameters:
directory - where the index resides.
Returns:
version number.
Throws:
java.io.IOException - if segments file cannot be read

getCurrentVersion

public static long getCurrentVersion(java.io.File directory)
                              throws java.io.IOException
Reads version number from segments files. The version number counts the number of changes of the index.

Parameters:
directory - where the index resides.
Returns:
version number.
Throws:
java.io.IOException - if segments file cannot be read

getCurrentVersion

public static long getCurrentVersion(Directory directory)
                              throws java.io.IOException
Reads version number from segments files. The version number counts the number of changes of the index.

Parameters:
directory - where the index resides.
Returns:
version number.
Throws:
java.io.IOException - if segments file cannot be read.

getTermFreqVectors

public abstract TermFreqVector[] getTermFreqVectors(int docNumber)
                                             throws java.io.IOException
Return an array of term frequency vectors for the specified document. The array contains a vector for each vectorized field in the document. Each vector contains terms and frequencies for all terms in a given vectorized field. If no such fields existed, the method returns null.

Throws:
java.io.IOException
See Also:
Field.isTermVectorStored()

getTermFreqVector

public abstract TermFreqVector getTermFreqVector(int docNumber,
                                                 java.lang.String field)
                                          throws java.io.IOException
Return a term frequency vector for the specified document and field. The vector returned contains terms and frequencies for those terms in the specified field of this document, if the field had storeTermVector flag set. If the flag was not set, the method returns null.

Throws:
java.io.IOException
See Also:
Field.isTermVectorStored()

indexExists

public static boolean indexExists(java.lang.String directory)
Returns true if an index exists at the specified directory. If the directory does not exist or if there is no index in it. false is returned.

Parameters:
directory - the directory to check for an index
Returns:
true if an index exists; false otherwise

indexExists

public static boolean indexExists(java.io.File directory)
Returns true if an index exists at the specified directory. If the directory does not exist or if there is no index in it.

Parameters:
directory - the directory to check for an index
Returns:
true if an index exists; false otherwise

indexExists

public static boolean indexExists(Directory directory)
                           throws java.io.IOException
Returns true if an index exists at the specified directory. If the directory does not exist or if there is no index in it.

Parameters:
directory - the directory to check for an index
Returns:
true if an index exists; false otherwise
Throws:
java.io.IOException - if there is a problem with accessing the index

numDocs

public abstract int numDocs()
Returns the number of documents in this index.


maxDoc

public abstract int maxDoc()
Returns one greater than the largest possible document number. This may be used to, e.g., determine how big to allocate an array which will have an element for every document number in an index.


document

public abstract Document document(int n)
                           throws java.io.IOException
Returns the stored fields of the nth Document in this index.

Throws:
java.io.IOException

isDeleted

public abstract boolean isDeleted(int n)
Returns true if document n has been deleted


hasDeletions

public abstract boolean hasDeletions()
Returns true if any documents have been deleted


norms

public abstract byte[] norms(java.lang.String field)
                      throws java.io.IOException
Returns the byte-encoded normalization factor for the named field of every document. This is used by the search code to score documents.

Throws:
java.io.IOException
See Also:
Field.setBoost(float)

norms

public abstract void norms(java.lang.String field,
                           byte[] bytes,
                           int offset)
                    throws java.io.IOException
Reads the byte-encoded normalization factor for the named field of every document. This is used by the search code to score documents.

Throws:
java.io.IOException
See Also:
Field.setBoost(float)

setNorm

public final void setNorm(int doc,
                          java.lang.String field,
                          byte value)
                   throws java.io.IOException
Expert: Resets the normalization factor for the named field of the named document. The norm represents the product of the field's boost and its length normalization. Thus, to preserve the length normalization values when resetting this, one should base the new value upon the old.

Throws:
java.io.IOException
See Also:
norms(String), Similarity.decodeNorm(byte)

doSetNorm

protected abstract void doSetNorm(int doc,
                                  java.lang.String field,
                                  byte value)
                           throws java.io.IOException
Implements setNorm in subclass.

Throws:
java.io.IOException

setNorm

public void setNorm(int doc,
                    java.lang.String field,
                    float value)
             throws java.io.IOException
Expert: Resets the normalization factor for the named field of the named document.

Throws:
java.io.IOException
See Also:
norms(String), Similarity.decodeNorm(byte)

terms

public abstract TermEnum terms()
                        throws java.io.IOException
Returns an enumeration of all the terms in the index. The enumeration is ordered by Term.compareTo(). Each term is greater than all that precede it in the enumeration.

Throws:
java.io.IOException

terms

public abstract TermEnum terms(Term t)
                        throws java.io.IOException
Returns an enumeration of all terms after a given term. The enumeration is ordered by Term.compareTo(). Each term is greater than all that precede it in the enumeration.

Throws:
java.io.IOException

docFreq

public abstract int docFreq(Term t)
                     throws java.io.IOException
Returns the number of documents containing the term t.

Throws:
java.io.IOException

termDocs

public TermDocs termDocs(Term term)
                  throws java.io.IOException
Returns an enumeration of all the documents which contain term. For each document, the document number, the frequency of the term in that document is also provided, for use in search scoring. Thus, this method implements the mapping:

The enumeration is ordered by document number. Each document number is greater than all that precede it in the enumeration.

Throws:
java.io.IOException

termDocs

public abstract TermDocs termDocs()
                           throws java.io.IOException
Returns an unpositioned TermDocs enumerator.

Throws:
java.io.IOException

termPositions

public TermPositions termPositions(Term term)
                            throws java.io.IOException
Returns an enumeration of all the documents which contain term. For each document, in addition to the document number and frequency of the term in that document, a list of all of the ordinal positions of the term in the document is available. Thus, this method implements the mapping:

This positional information faciliates phrase and proximity searching.

The enumeration is ordered by document number. Each document number is greater than all that precede it in the enumeration.

Throws:
java.io.IOException

termPositions

public abstract TermPositions termPositions()
                                     throws java.io.IOException
Returns an unpositioned TermPositions enumerator.

Throws:
java.io.IOException

delete

public final void delete(int docNum)
                  throws java.io.IOException
Deletes the document numbered docNum. Once a document is deleted it will not appear in TermDocs or TermPostitions enumerations. Attempts to read its field with the document(int) method will result in an error. The presence of this document may still be reflected in the docFreq(org.apache.lucene.index.Term) statistic, though this will be corrected eventually as the index is further modified.

Throws:
java.io.IOException

doDelete

protected abstract void doDelete(int docNum)
                          throws java.io.IOException
Implements deletion of the document numbered docNum. Applications should call delete(int) or delete(Term).

Throws:
java.io.IOException

delete

public final int delete(Term term)
                 throws java.io.IOException
Deletes all documents containing term. This is useful if one uses a document field to hold a unique ID string for the document. Then to delete such a document, one merely constructs a term with the appropriate field and the unique ID string as its text and passes it to this method. Returns the number of documents deleted.

Throws:
java.io.IOException

undeleteAll

public final void undeleteAll()
                       throws java.io.IOException
Undeletes all documents currently marked as deleted in this index.

Throws:
java.io.IOException

doUndeleteAll

protected abstract void doUndeleteAll()
                               throws java.io.IOException
Implements actual undeleteAll() in subclass.

Throws:
java.io.IOException

commit

protected final void commit()
                     throws java.io.IOException
Commit changes resulting from delete, undeleteAll, or setNorm operations

Throws:
java.io.IOException

doCommit

protected abstract void doCommit()
                          throws java.io.IOException
Implements commit.

Throws:
java.io.IOException

close

public final void close()
                 throws java.io.IOException
Closes files associated with this index. Also saves any new deletions to disk. No other methods should be called after this has been called.

Throws:
java.io.IOException

doClose

protected abstract void doClose()
                         throws java.io.IOException
Implements close.

Throws:
java.io.IOException

finalize

protected final void finalize()
                       throws java.io.IOException
Release the write lock, if needed.

Throws:
java.io.IOException

getFieldNames

public abstract java.util.Collection getFieldNames()
                                            throws java.io.IOException
Returns a list of all unique field names that exist in the index pointed to by this IndexReader.

Returns:
Collection of Strings indicating the names of the fields
Throws:
java.io.IOException - if there is a problem with accessing the index

getFieldNames

public abstract java.util.Collection getFieldNames(boolean indexed)
                                            throws java.io.IOException
Returns a list of all unique field names that exist in the index pointed to by this IndexReader. The boolean argument specifies whether the fields returned are indexed or not.

Parameters:
indexed - true if only indexed fields should be returned; false if only unindexed fields should be returned.
Returns:
Collection of Strings indicating the names of the fields
Throws:
java.io.IOException - if there is a problem with accessing the index

getIndexedFieldNames

public abstract java.util.Collection getIndexedFieldNames(boolean storedTermVector)
Parameters:
storedTermVector - if true, returns only Indexed fields that have term vector info, else only indexed fields without term vector info
Returns:
Collection of Strings indicating the names of the fields

isLocked

public static boolean isLocked(Directory directory)
                        throws java.io.IOException
Returns true iff the index in the named directory is currently locked.

Parameters:
directory - the directory to check for a lock
Throws:
java.io.IOException - if there is a problem with accessing the index

isLocked

public static boolean isLocked(java.lang.String directory)
                        throws java.io.IOException
Returns true iff the index in the named directory is currently locked.

Parameters:
directory - the directory to check for a lock
Throws:
java.io.IOException - if there is a problem with accessing the index

unlock

public static void unlock(Directory directory)
                   throws java.io.IOException
Forcibly unlocks the index in the named directory.

Caution: this should only be used by failure recovery code, when it is known that no other process nor thread is in fact currently accessing this index.

Throws:
java.io.IOException


Copyright © 2000-2005 Apache Software Foundation. All Rights Reserved.