lucene-dev mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Dawid Weiss (JIRA)" <>
Subject [jira] Commented: (LUCENE-2341) explore morfologik integration
Date Wed, 24 Mar 2010 09:17:27 GMT


Dawid Weiss commented on LUCENE-2341:

Oh, I forgot about this -- yes, you're right, it can emit multiple lemmas (including their
morphology). An example from the test case:

		final IStemmer s = new PolishStemmer();

		final String word = "liga";
		final List<WordData> response = s.lookup(word);

		final HashSet<String> stems = new HashSet<String>();
		final HashSet<String> tags = new HashSet<String>();

		for (WordData wd : response) {
			assertSame(word, wd.getWord());


This raises the question how we want to handle multiple stems... should they be indexed on
overlapping positions?

> explore morfologik integration
> ------------------------------
>                 Key: LUCENE-2341
>                 URL:
>             Project: Lucene - Java
>          Issue Type: New Feature
>          Components: contrib/analyzers
>            Reporter: Robert Muir
> Dawid Weiss mentioned on LUCENE-2298 that there is another Polish stemmer available:
> This works differently than LUCENE-2298, and ideally would be another option for users.

This message is automatically generated by JIRA.
You can reply to this email to add a comment to the issue online.

To unsubscribe, e-mail:
For additional commands, e-mail:

View raw message