hadoop-common-dev mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Abdul Qadeer (JIRA)" <j...@apache.org>
Subject [jira] Created: (HADOOP-3646) Providing bzip2 as codec
Date Wed, 25 Jun 2008 22:42:46 GMT
Providing bzip2 as codec

                 Key: HADOOP-3646
                 URL: https://issues.apache.org/jira/browse/HADOOP-3646
             Project: Hadoop Core
          Issue Type: Improvement
          Components: conf, io
    Affects Versions: 0.19.0
            Reporter: Abdul Qadeer

Hadoop recognizes gzip compressed input and automatically decompresses the data before providing
it to the mapper. But Hadoop can not split a gzip stream due to the very nature of the gzip
compression. Consequently one gzip stream (e.g a whole file) can go to only one mapper.  On
the contrary Bzip2 compressed stream can be split across its block delimiters.

We are interested in extending Hadoop to support splittable bzip2 with a codec.  (https://issues.apache.org/jira/browse/HADOOP-1823
 uses input reader to split the bzip2 files, which must be provided by the user and can handle
FileInputFormat.  If a user wants to use some other input format or wants to do custom record
handling, he must write a new input reader!)

We have a patch now that provides a basic bzip2 codec equivalent to the current gzip codec.
 We are in the process of extending that to support splitting.

This message is automatically generated by JIRA.
You can reply to this email to add a comment to the issue online.

View raw message