Return-Path: X-Original-To: apmail-commons-issues-archive@minotaur.apache.org Delivered-To: apmail-commons-issues-archive@minotaur.apache.org Received: from mail.apache.org (hermes.apache.org [140.211.11.3]) by minotaur.apache.org (Postfix) with SMTP id 25F6411199 for ; Sun, 29 Jun 2014 19:29:25 +0000 (UTC) Received: (qmail 41507 invoked by uid 500); 29 Jun 2014 19:29:24 -0000 Delivered-To: apmail-commons-issues-archive@commons.apache.org Received: (qmail 41410 invoked by uid 500); 29 Jun 2014 19:29:24 -0000 Mailing-List: contact issues-help@commons.apache.org; run by ezmlm Precedence: bulk List-Help: List-Unsubscribe: List-Post: List-Id: Reply-To: issues@commons.apache.org Delivered-To: mailing list issues@commons.apache.org Received: (qmail 41396 invoked by uid 99); 29 Jun 2014 19:29:24 -0000 Received: from arcas.apache.org (HELO arcas.apache.org) (140.211.11.28) by apache.org (qpsmtpd/0.29) with ESMTP; Sun, 29 Jun 2014 19:29:24 +0000 Date: Sun, 29 Jun 2014 19:29:24 +0000 (UTC) From: "Thomas Neidhart (JIRA)" To: issues@commons.apache.org Message-ID: In-Reply-To: References: Subject: [jira] [Commented] (CODEC-187) Beider Morse Phonetic Matching producing incorrect tokens MIME-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Transfer-Encoding: 7bit X-JIRA-FingerPrint: 30527f35849b9dde25b450d4833f0394 [ https://issues.apache.org/jira/browse/CODEC-187?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14047218#comment-14047218 ] Thomas Neidhart commented on CODEC-187: --------------------------------------- Yes, I had planned to compare all rules with the original version, but this will take some time as I want to create a script that does an automatic consistency check to see differences in the rules files. > Beider Morse Phonetic Matching producing incorrect tokens > --------------------------------------------------------- > > Key: CODEC-187 > URL: https://issues.apache.org/jira/browse/CODEC-187 > Project: Commons Codec > Issue Type: Bug > Affects Versions: 1.9 > Reporter: michael tobias > Priority: Minor > Fix For: 1.10 > > Attachments: CODEC-187.patch, CODEC-187_ashkenazi_approx_any.patch, CODEC-187_ashkenazi_approx_any_v2.patch > > > I believe the Beider Morse Phonetic Matching algorithm was added in Commons Codec 1.6 > The BMPM algorithm is an EVOLVING algorithm that is currently on version 3.02 though it had been static since version 3.01 dated 19 Dec 2011 (it was first available as opensource as version 1.00 on 6 May 2009). > I can see nothing in the Commons Codec Docs to say which version of BMPM was implemented so I am not sure if the problem with the algorithm as coded in the Codec is simply an old version or whether there are more basic problems with the implementation. > How do I determine the version of the algorithm that was implemented in the Commons Codec? > How do we ensure that the algorithm is updated if/when the BMPM algorithm changes? > How do we ensure that the algorithm as coded in the Commons Codec is accurate and working as expected? -- This message was sent by Atlassian JIRA (v6.2#6252)