flink-issues mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Sachin Goel (JIRA)" <j...@apache.org>
Subject [jira] [Commented] (FLINK-2312) Random Splits
Date Tue, 21 Jul 2015 16:04:05 GMT

    [ https://issues.apache.org/jira/browse/FLINK-2312?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14635302#comment-14635302

Sachin Goel commented on FLINK-2312:

[~pietropinoli], It didn't work. I, instead, decided to add an API function to both java and
scala DataSetUtils instead. It's finished now.

> Random Splits
> -------------
>                 Key: FLINK-2312
>                 URL: https://issues.apache.org/jira/browse/FLINK-2312
>             Project: Flink
>          Issue Type: Wish
>          Components: Machine Learning Library
>            Reporter: Maximilian Alber
>            Assignee: pietro pinoli
>            Priority: Minor
> In machine learning applications it is common to split data sets into f.e. training and
testing set.
> To the best of my knowledge there is at the moment no nice way in Flink to split a data
set randomly into several partitions according to some ratio.
> The wished semantic would be the same as of Sparks RDD randomSplit.

This message was sent by Atlassian JIRA

View raw message