hadoop-common-issues mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Haibo Chen (JIRA)" <j...@apache.org>
Subject [jira] [Commented] (HADOOP-12436) GlobPattern regex library has performance issues with wildcard characters
Date Tue, 10 Oct 2017 21:57:00 GMT

    [ https://issues.apache.org/jira/browse/HADOOP-12436?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16199434#comment-16199434
] 

Haibo Chen commented on HADOOP-12436:
-------------------------------------

[~aw] [~mattpaduano] This seems an incompatible change given GlobFilter and RegexFilter are
Public Evolving. Hence, I have added an incompatible tag. Feel free to remove it if you disagree

> GlobPattern regex library has performance issues with wildcard characters
> -------------------------------------------------------------------------
>
>                 Key: HADOOP-12436
>                 URL: https://issues.apache.org/jira/browse/HADOOP-12436
>             Project: Hadoop Common
>          Issue Type: Improvement
>          Components: fs
>    Affects Versions: 2.2.0, 2.7.1
>            Reporter: Matthew Paduano
>            Assignee: Matthew Paduano
>             Fix For: 3.0.0-alpha1
>
>         Attachments: HADOOP-12436.01.patch, HADOOP-12436.02.patch, HADOOP-12436.03.patch,
HADOOP-12436.04.patch, HADOOP-12436.05.patch
>
>
> java.util.regex classes have performance problems with certain wildcard patterns.  Namely,
consecutive * characters in a file name (not properly escaped as literals) will cause commands
such as "hadoop fs -ls file******name" to consume 100% CPU and probably never return in a
reasonable time (time scales with number of *'s). 
> Here is an example:
> {noformat}
> hadoop fs -touchz /user/mattp/job_1429571161900_4222-1430338332599-tda%2D%2D\\\+\\\*\\\*\\\*\\\*\\\*\\\*\\\*\\\*\\\*\\\*\\\*\\\*\\\*\\\*\\\*\\\*\\\*\\\*\\\*\\\*\\\*\\\*\\\*\\\*\\\*\\\*\\\*\\\*\\\*\\\*\\\+\\\+\\\+...%270%27%28Stage-1430338580443-39-2000-SUCCEEDED-production%2Dhigh-1430338340360.jhist
> hadoop fs -ls /user/mattp/job_1429571161900_4222-1430338332599-tda%2D%2D+******************************+++...%270%27%28Stage-1430338580443-39-2000-SUCCEEDED-production%2Dhigh-1430338340360.jhist
> {noformat}
> causes:
> {noformat}
> PID    COMMAND      %CPU   TIME  
> 14526  java         100.0  01:18.85 
> {noformat}
> Not every string of *'s causes this, but the above filename reproduces this reliably.



--
This message was sent by Atlassian JIRA
(v6.4.14#64029)

---------------------------------------------------------------------
To unsubscribe, e-mail: common-issues-unsubscribe@hadoop.apache.org
For additional commands, e-mail: common-issues-help@hadoop.apache.org


Mime
View raw message