nutch-dev mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "nutch.newbie (JIRA)" <j...@apache.org>
Subject [jira] Commented: (NUTCH-444) Possibly use a different library to parse RSS feed for improved performance and compatibility
Date Tue, 13 Feb 2007 14:04:06 GMT

    [ https://issues.apache.org/jira/browse/NUTCH-444?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#action_12472663
] 

nutch.newbie commented on NUTCH-444:
------------------------------------

Hi all:

I didn't realize that there was version 6 patch for NUTCH-443. After applying the patch all
seems to be working. Furthermore I like to thank Dogacan for helping me on the way. Fetching,
crawling and dedup/index works just fine.

I would like to use parse-feed and be of help in terms of writing/testing index-feed and query-feed
so it would be nice if commiters would be kind enough to test the patch NUTCH-443 and apply
it to trunk. So that work regarding index-feed and query-feed can begin. 

If there are any more test or anything else that you guys want me to perform to test NUTCH-443
or NUTCH-444 please tell me. 

Regards


> Possibly use a different library to parse RSS feed for improved performance and compatibility
> ---------------------------------------------------------------------------------------------
>
>                 Key: NUTCH-444
>                 URL: https://issues.apache.org/jira/browse/NUTCH-444
>             Project: Nutch
>          Issue Type: Improvement
>          Components: fetcher
>    Affects Versions: 0.9.0
>            Reporter: Renaud Richardet
>            Priority: Minor
>             Fix For: 0.9.0
>
>         Attachments: parse-feed-v2.tar.bz2, parse-feed.tar.bz2
>
>
> As discussed by Nutch Newbie, Gal, and Chris on NUTCH-443, the current library (feedparser)
has the following issues:
> - OutOfMemory when parsing > 100k feeds, since it has to convert the feed to jdom
first
> - no support for Atom 1.0
> - there has been no development in the last year
> Alternatives are:
> - Rome 
> - Informa
> - custom implementation based on Stax
> - ??

-- 
This message is automatically generated by JIRA.
-
You can reply to this email to add a comment to the issue online.


Mime
View raw message