spark-reviews mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From ron8hu <...@git.apache.org>
Subject [GitHub] spark pull request #15363: [SPARK-17791][SQL] Join reordering using star sch...
Date Wed, 15 Mar 2017 21:18:04 GMT
Github user ron8hu commented on a diff in the pull request:

    https://github.com/apache/spark/pull/15363#discussion_r106285527
  
    --- Diff: sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/optimizer/CostBasedJoinReorder.scala
---
    @@ -51,6 +51,11 @@ case class CostBasedJoinReorder(conf: CatalystConf) extends Rule[LogicalPlan]
wi
     
       def reorder(plan: LogicalPlan, output: AttributeSet): LogicalPlan = {
         val (items, conditions) = extractInnerJoins(plan)
    +    // Find the star schema joins. Currently, it returns the star join with the largest
    +    // fact table. In the future, it can return more than one star join (e.g. F1-D1-D2
    +    // and F2-D3-D4).
    +    val starJoinPlans = StarSchemaDetection(conf).findStarJoins(items, conditions.toSeq)
    --- End diff --
    
    @ioana-delaney Thanks for the pointer.  We had a good exchange to clarify our points.
 We definitely need a joint community team effort to improve Spark's cost based optimizer
in the future.


---
If your project is set up for it, you can reply to this email and have your
reply appear on GitHub as well. If your project does not have this feature
enabled and wishes so, or if the feature is enabled but not working, please
contact infrastructure at infrastructure@apache.org or file a JIRA ticket
with INFRA.
---

---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org


Mime
View raw message