nutch-dev mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From Dogacan Güney (JIRA) <j...@apache.org>
Subject [jira] Updated: (NUTCH-444) Possibly use a different library to parse RSS feed for improved performance and compatibility
Date Sun, 11 Feb 2007 15:42:05 GMT

     [ https://issues.apache.org/jira/browse/NUTCH-444?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]

Dogacan Güney updated NUTCH-444:
--------------------------------

    Attachment: parse-feed-v2.tar.bz2

Updated parse-feed plugin. Still not ready for any serious use, but I think I fixed the problems
with indexing and dedup. Use it with NUTCH-443's v5 patch.

nutch.newbie: I change parse-plugins.xml as you do. For this plugin to work, you also have
to change default signature to TextProfileSignature(because MD5Signature takes the hash of
content, which is the same for every element in a parseMap). This is done by adding:
<property>
  <name>db.signature.class</name>
  <value>org.apache.nutch.crawl.TextProfileSignature</value>
</property>

to your nutch-site.xml.


> Possibly use a different library to parse RSS feed for improved performance and compatibility
> ---------------------------------------------------------------------------------------------
>
>                 Key: NUTCH-444
>                 URL: https://issues.apache.org/jira/browse/NUTCH-444
>             Project: Nutch
>          Issue Type: Improvement
>          Components: fetcher
>    Affects Versions: 0.9.0
>            Reporter: Renaud Richardet
>            Priority: Minor
>             Fix For: 0.9.0
>
>         Attachments: parse-feed-v2.tar.bz2, parse-feed.tar.bz2
>
>
> As discussed by Nutch Newbie, Gal, and Chris on NUTCH-443, the current library (feedparser)
has the following issues:
> - OutOfMemory when parsing > 100k feeds, since it has to convert the feed to jdom
first
> - no support for Atom 1.0
> - there has been no development in the last year
> Alternatives are:
> - Rome 
> - Informa
> - custom implementation based on Stax
> - ??

-- 
This message is automatically generated by JIRA.
-
You can reply to this email to add a comment to the issue online.


Mime
View raw message