hadoop-pig-dev mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Benjamin Reed (JIRA)" <j...@apache.org>
Subject [jira] Created: (PIG-42) Pig should be able to split Gzip files like it can split Bzip files
Date Sat, 01 Dec 2007 22:16:43 GMT
Pig should be able to split Gzip files like it can split Bzip files
-------------------------------------------------------------------

                 Key: PIG-42
                 URL: https://issues.apache.org/jira/browse/PIG-42
             Project: Pig
          Issue Type: Improvement
          Components: impl
            Reporter: Benjamin Reed


It would be nice to be able to split gzip files like we can split bzip files. Unfortunately,
we don't have a sync point for the split in the gzip format.

Gzip file format supports the notion of concatenate gzipped files. When gzipped files are
concatenated together they are treated as a single file. So to make a gzipped file splittable
we can used an empty compressed file with some salt in the headers as a sync signature. Then
we can make the gzip file splittable by using this sync signature between compressed segments
of the file.

-- 
This message is automatically generated by JIRA.
-
You can reply to this email to add a comment to the issue online.


Mime
View raw message