lucene-dev mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Christian Moen (Commented) (JIRA)" <j...@apache.org>
Subject [jira] [Commented] (LUCENE-3897) KuromojiTokenizer fails with large docs
Date Wed, 21 Mar 2012 16:11:52 GMT

    [ https://issues.apache.org/jira/browse/LUCENE-3897?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=13234456#comment-13234456
] 

Christian Moen commented on LUCENE-3897:
----------------------------------------

I've been trying to make an even more isolated case that reproduces this problem.  I'm new
to {{LuceneTestCaseRunner}} and {{checkRandomData}}, but I received very helpful advise on
how to follow up on this.  Thanks, Robert!

I'll look further into this tomorrow.  Mike, if you have any ideas on what the root cause
of this could be, please feel free to chime in.  Many thanks.
                
> KuromojiTokenizer fails with large docs
> ---------------------------------------
>
>                 Key: LUCENE-3897
>                 URL: https://issues.apache.org/jira/browse/LUCENE-3897
>             Project: Lucene - Java
>          Issue Type: Bug
>          Components: modules/analysis
>            Reporter: Robert Muir
>             Fix For: 3.6, 4.0
>
>
> just shoving largeish random docs triggers asserts like:
> {noformat}
>     [junit] Caused by: java.lang.AssertionError: backPos=4100 vs lastBackTracePos=5120
>     [junit] 	at org.apache.lucene.analysis.kuromoji.KuromojiTokenizer.backtrace(KuromojiTokenizer.java:907)
>     [junit] 	at org.apache.lucene.analysis.kuromoji.KuromojiTokenizer.parse(KuromojiTokenizer.java:756)
>     [junit] 	at org.apache.lucene.analysis.kuromoji.KuromojiTokenizer.incrementToken(KuromojiTokenizer.java:403)
>     [junit] 	at org.apache.lucene.analysis.BaseTokenStreamTestCase.checkRandomData(BaseTokenStreamTestCase.java:404)
> {noformat}
> But, you get no seed...
> I'll commit the test case and @Ignore it.

--
This message is automatically generated by JIRA.
If you think it was sent incorrectly, please contact your JIRA administrators: https://issues.apache.org/jira/secure/ContactAdministrators!default.jspa
For more information on JIRA, see: http://www.atlassian.com/software/jira

        

---------------------------------------------------------------------
To unsubscribe, e-mail: dev-unsubscribe@lucene.apache.org
For additional commands, e-mail: dev-help@lucene.apache.org


Mime
View raw message