lucene-dev mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Robert Muir (JIRA)" <>
Subject [jira] Commented: (LUCENE-2503) light/minimal stemming for euro languages
Date Thu, 17 Jun 2010 21:31:23 GMT


Robert Muir commented on LUCENE-2503:

bq. Man are you fast!

not really, i've been working it for a while but since someone asked i figure i would create
the issue.
testing isnt done, but english, french, portuguese I think are ok.
the others need a lot of tests and probably have bugs.

bq. Does the English one deal with women/ woman and foci / focus type stuff?

Nope, the english one is the Harman "s-stemming" algorithm.

its very simple:
if final is '-ies' but not '-eies' or '-aies' then
replace '-ies' by '-y', return;
if final is '-es' but not '-aes', '-ees' or '-oes' then
replace '-es' by '-e', return;
if final is '-s' but not '-us' or '-ss' then
remove '-s';

For special cases like you mentioned (if you want them), i would recommend adding these customizations
as documented here:

just make a tab-separated file of words-stems and put a StemmerOverrideFilter(Factory) before
the stemmer in the stream.

I think this alone provides a lot of flexibility. if it isn't enough, then i think these stemmers
are much simpler to modify if you wanted to go that route also :)

> light/minimal stemming for euro languages
> -----------------------------------------
>                 Key: LUCENE-2503
>                 URL:
>             Project: Lucene - Java
>          Issue Type: New Feature
>          Components: contrib/analyzers
>    Affects Versions: 3.1, 4.0
>            Reporter: Robert Muir
>            Assignee: Robert Muir
>            Priority: Minor
>             Fix For: 3.1, 4.0
>         Attachments: LUCENE-2503.patch
> The snowball stemmers are very aggressive and it would be nice if there were lighter
> Some applications may want to perform less aggressive stemming, for example:
> Good, relevance tested algorithms exist and I think we should provide these alternatives.

This message is automatically generated by JIRA.
You can reply to this email to add a comment to the issue online.

To unsubscribe, e-mail:
For additional commands, e-mail:

View raw message