hive-dev mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Chengxiang Li (JIRA)" <>
Subject [jira] [Commented] (HIVE-7142) Hive multi serialization encoding support
Date Mon, 11 Aug 2014 09:30:12 GMT


Chengxiang Li commented on HIVE-7142:

Hi, [~navis], this jira is trying to support table level configurable encoding, i take a look
at HIVE-6329, do you mean you want to implement column level configurable encoding? If yes,
that should be a quite different implementation. But that would be valuable as well, and i'm
glad to see that.

> Hive multi serialization encoding support
> -----------------------------------------
>                 Key: HIVE-7142
>                 URL:
>             Project: Hive
>          Issue Type: Improvement
>          Components: Serializers/Deserializers
>            Reporter: Chengxiang Li
>            Assignee: Chengxiang Li
>         Attachments: HIVE-7142.1.patch.txt, HIVE-7142.2.patch, HIVE-7142.3.patch
> Currently Hive only support serialize data into UTF-8 charset bytes or deserialize from
UTF-8 bytes, real world users may want to load different kinds of encoded data into hive directly.
This jira is dedicated to support serialize/deserialize all kinds of encoded data in SerDe
> For user, only need to configure serialization encoding on table level by set serialization
encoding through serde parameter, for example:
> {code:sql}
> CREATE TABLE person(id INT, name STRING, desc STRING)ROW FORMAT SERDE 'org.apache.hadoop.hive.serde2.lazy.LazySimpleSerDe'
WITH SERDEPROPERTIES("serialization.encoding"='GBK');
> {code}
> or
> {code:sql}
> ALTER TABLE person SET SERDEPROPERTIES ('serialization.encoding'='GBK'); 
> {code}
> LIMITATIONS: Only LazySimpleSerDe support "serialization.encoding" property in this patch.

This message was sent by Atlassian JIRA

View raw message