lucene-dev mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Otis Gospodnetic (JIRA)" <>
Subject [jira] Commented: (LUCENE-1491) EdgeNGramTokenFilter stops on tokens smaller then minimum gram size.
Date Tue, 02 Jun 2009 20:27:07 GMT


Otis Gospodnetic commented on LUCENE-1491:

I'm not 100% sure - I'm not using ngrams at the moment, so I have no place to test this out,
but skipping a shorter than minimal ngrams seems like it would result in silent data loss.

Ah, here, example:
What would happen to "to be or not to be" if min=4 and we relied on ngrams to perform phrase

All of those terms would be dropped, so a search for "to be or not to be" would result in
0 hits.

If the above is correct, I think this sounds like a bad thing that one wouldn't expect...

> EdgeNGramTokenFilter stops on tokens smaller then minimum gram size.
> --------------------------------------------------------------------
>                 Key: LUCENE-1491
>                 URL:
>             Project: Lucene - Java
>          Issue Type: Bug
>          Components: Analysis
>    Affects Versions: 2.4, 2.4.1, 2.9, 3.0
>            Reporter: Todd Feak
>            Assignee: Otis Gospodnetic
>             Fix For: 2.9
>         Attachments: LUCENE-1491.patch
> If a token is encountered in the stream that is shorter in length than the min gram size,
the filter will stop processing the token stream.
> Working up a unit test now, but may be a few days before I can provide it. Wanted to
get it in the system.

This message is automatically generated by JIRA.
You can reply to this email to add a comment to the issue online.

To unsubscribe, e-mail:
For additional commands, e-mail:

View raw message