drill-issues mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Aman Sinha (JIRA)" <j...@apache.org>
Subject [jira] [Created] (DRILL-4530) Improve metadata cache performance for queries with single partition
Date Wed, 23 Mar 2016 00:52:25 GMT
Aman Sinha created DRILL-4530:
---------------------------------

             Summary: Improve metadata cache performance for queries with single partition

                 Key: DRILL-4530
                 URL: https://issues.apache.org/jira/browse/DRILL-4530
             Project: Apache Drill
          Issue Type: Improvement
          Components: Query Planning & Optimization
    Affects Versions: 1.6.0
            Reporter: Aman Sinha
            Assignee: Aman Sinha


Consider two types of queries which are run with Parquet metadata caching: 
{noformat}
query 1:
SELECT col FROM  `A/B/C`;

query 2:
SELECT col FROM `A` WHERE dir0 = 'B' AND dir1 = 'C';
{noformat}

For a certain dataset, the query1 elapsed time is 1 sec whereas query1 elapsed time is 9 sec
even though both are accessing the same amount of data.  The user expectation is that they
should perform roughly the same.  The main difference comes from reading the bigger metadata
cache file at the root level 'A' for query2 and then applying the partitioning filter.  query1
reads a much smaller metadata cache file at the subdirectory level. 




--
This message was sent by Atlassian JIRA
(v6.3.4#6332)

Mime
View raw message