spark-reviews mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From GitBox <...@apache.org>
Subject [GitHub] [spark] prakharjain09 opened a new pull request #26569: [SPARK-29938] [SQL] Add batching support in Alter table add partition flow
Date Mon, 18 Nov 2019 06:31:19 GMT
prakharjain09 opened a new pull request #26569: [SPARK-29938] [SQL] Add batching support in
Alter table add partition flow
URL: https://github.com/apache/spark/pull/26569
 
 
   ### What changes were proposed in this pull request?
   This PR split the the single createPartitions() call AlterTableAddPartition flow into smaller
batches, which could prevent
    - SocketTimeoutException: Adding thousand of partitions in Hive metastore itself takes
lot of time. Because of this hive client fails with SocketTimeoutException.
   
   - Hive metastore from OOM (caused by millions of partitions).
   
   It will also try to gather stats (total size of all files in all new partitions) faster
by parallely listing the new partition paths.
   
   ### How was this patch tested?
   Added UT. 
   Also tested on a cluster in HDI with 15000 partitions with remote metastore server. Without
batching - operation fails with SocketTimeoutException, With batching it finishes in 25 mins.
   

----------------------------------------------------------------
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
 
For queries about this service, please contact Infrastructure at:
users@infra.apache.org


With regards,
Apache Git Services

---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org


Mime
View raw message