lucene-dev mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Mark Harwood (JIRA)" <>
Subject [jira] [Created] (LUCENE-4069) Segment-level Bloom filters for a 2 x speed up on rare term searches
Date Fri, 18 May 2012 15:45:13 GMT
Mark Harwood created LUCENE-4069:

             Summary: Segment-level Bloom filters for a 2 x speed up on rare term searches
                 Key: LUCENE-4069
             Project: Lucene - Java
          Issue Type: Improvement
          Components: core/index
    Affects Versions: 3.6
            Reporter: Mark Harwood
            Priority: Minor
             Fix For: 3.6.1

An addition to each segment which stores a Bloom filter for selected fields in order to give
fast-fail to term searches, helping avoid wasted disk access.

Best suited for low-frequency fields e.g. primary keys on big indexes with many segments but
also speeds up general searching in my tests.

Overview slideshow here:

Benchmarks based on Wikipedia content here:

Patch based on 3.6 codebase attached.
There are no API changes currently - to play just add a field with "_blm" on the end of the
name to invoke special indexing/querying capability. Clearly a new Field or schema declaration(!)
would need adding to APIs to configure the service properly.

This message is automatically generated by JIRA.
If you think it was sent incorrectly, please contact your JIRA administrators:!default.jspa
For more information on JIRA, see:


To unsubscribe, e-mail:
For additional commands, e-mail:

View raw message