lucene-dev mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From Michał Dybizbański (JIRA) <>
Subject [jira] [Updated] (LUCENE-2341) explore morfologik integration
Date Mon, 20 Jun 2011 21:48:47 GMT


Michał Dybizbański updated LUCENE-2341:

    Attachment: morfologik-stemming-1.5.0.jar


This patch introduces stemming filter and analyzer, that use [Morfologik library|],
developed by Dawid Weiss and Marcin Miłkowski.
Tokens are stemmed by Morfologik with a dictionary, and current distribution provides a dictionary
for polish language.

The MorfologikFilter yields one or more terms for each token. Each of those terms is given
the same position in the index.

I'm attaching a binary distribution of the library (morfologik-stemming-1.5.0.jar), that needs
to be placed in modules/analysis/morfologik/lib/ subdirectory.
It is also available as a [Maven artifact|].

The library is BSD-licensed and a dictionary uses data from [Polish dictionary for aspell/ispell/myspell
(SJP.PL)|], which is licensed under GPL, LGPL, MPL and CC SA

This is my first contribution to the Lucene project, so please be forgiving :)
Thanks to Dawid for help.


> explore morfologik integration
> ------------------------------
>                 Key: LUCENE-2341
>                 URL:
>             Project: Lucene - Java
>          Issue Type: New Feature
>          Components: modules/analysis
>            Reporter: Robert Muir
>            Assignee: Dawid Weiss
>         Attachments: LUCENE-2341.diff, morfologik-stemming-1.5.0.jar
> Dawid Weiss mentioned on LUCENE-2298 that there is another Polish stemmer available:
> This works differently than LUCENE-2298, and ideally would be another option for users.

This message is automatically generated by JIRA.
For more information on JIRA, see:


To unsubscribe, e-mail:
For additional commands, e-mail:

View raw message