首页 文章 精选 留言 我的

精选列表

搜索[AI意图识别],共10000篇文章
优秀的个人博客,低调大师

Hadoop WordCount改进实现正确识别单词以及词频降序排序

0.参考资料: http://radarradar.javaeye.com/blog/289257 http://blog.chinaunix.net/u3/99156/showart_2157576.html 1.思路: 1.1过滤 MapReduce的第一操作就是要读取文件,不过我们经常会发现一个文本中会有一些我们不需要的字符,比如特殊字符。一般需要进行词频统计的都是单词或者是数字,所以那些非0-9,a-z,A-Z的字符基本都是垃圾字符,我们需要进行统计,这是我们可以通过一个正则表达式来进行过滤,当每次多去一行文字的时候,我们将所有非0-9,a-z,A-Z的垃圾字符都替换为空格,这样就清楚了垃圾字符。在我们最后的词频统计结果中,就不会出现这些特殊字符了。 1.2降序 定义一个用户排序比较的静态内部类,通过这个类来控制词频统计最后的排序结果。我们这里所使用的静态内部类是IntWritableDecreasingComparator。需要注意的是必须在main函数中主动声明使用这个比较器。 2.代码实例 package org.apache.hadoop.examples; import java.io.IOException; import java.util.Random; import java.util.StringTokenizer; import org.apache.hadoop.conf.Configuration; import org.apache.hadoop.fs.FileSystem; import org.apache.hadoop.fs.Path; import org.apache.hadoop.io.IntWritable; import org.apache.hadoop.io.Text; import org.apache.hadoop.io.WritableComparable; import org.apache.hadoop.mapreduce.Job; import org.apache.hadoop.mapreduce.Mapper; import org.apache.hadoop.mapreduce.Reducer; import org.apache.hadoop.mapreduce.lib.input.FileInputFormat; import org.apache.hadoop.mapreduce.lib.input.SequenceFileInputFormat; import org.apache.hadoop.mapreduce.lib.map.InverseMapper; import org.apache.hadoop.mapreduce.lib.output.FileOutputFormat; import org.apache.hadoop.mapreduce.lib.output.SequenceFileOutputFormat; import org.apache.hadoop.util.GenericOptionsParser; public class WordCount2 { public static class TokenizerMapper extends Mapper<Object, Text, Text, IntWritable> { private final static IntWritable one = new IntWritable(1); private Text word = new Text(); private String pattern = "[^//w]"; // 正则表达式,代表不是0-9, a-z, A-Z的所有其它字符,其中还有下划线 public void map(Object key, Text value, Context context) throws IOException, InterruptedException { String line = value.toString().toLowerCase(); // 全部转为小写字母 line = line.replaceAll(pattern, " "); // 将非0-9, a-z, A-Z的字符替换为空格 StringTokenizer itr = new StringTokenizer(line); while (itr.hasMoreTokens()) { word.set(itr.nextToken()); context.write(word, one); } } } public static class IntSumReducer extends Reducer<Text, IntWritable, Text, IntWritable> { private IntWritable result = new IntWritable(); public void reduce(Text key, Iterable<IntWritable> values, Context context) throws IOException, InterruptedException { int sum = 0; for (IntWritable val : values) { sum += val.get(); } result.set(sum); context.write(key, result); } } private static class IntWritableDecreasingComparator extends IntWritable.Comparator { public int compare(WritableComparable a, WritableComparable b) { return -super.compare(a, b); } public int compare(byte[] b1, int s1, int l1, byte[] b2, int s2, int l2) { return -super.compare(b1, s1, l1, b2, s2, l2); } } public static void main(String[] args) throws Exception { Configuration conf = new Configuration(); String[] otherArgs = new GenericOptionsParser(conf, args) .getRemainingArgs(); if (otherArgs.length != 2) { System.err.println("Usage: wordcount <in> <out>"); System.exit(2); } Path tempDir = new Path("wordcount-temp-" + Integer.toString( new Random().nextInt(Integer.MAX_VALUE))); //定义一个临时目录 Job job = new Job(conf, "word count"); job.setJarByClass(WordCount2.class); try{ job.setMapperClass(TokenizerMapper.class); job.setCombinerClass(IntSumReducer.class); job.setReducerClass(IntSumReducer.class); job.setOutputKeyClass(Text.class); job.setOutputValueClass(IntWritable.class); FileInputFormat.addInputPath(job, new Path(otherArgs[0])); FileOutputFormat.setOutputPath(job, tempDir);//先将词频统计任务的输出结果写到临时目 //录中, 下一个排序任务以临时目录为输入目录。 job.setOutputFormatClass(SequenceFileOutputFormat.class); if(job.waitForCompletion(true)) { Job sortJob = new Job(conf, "sort"); sortJob.setJarByClass(WordCount2.class); FileInputFormat.addInputPath(sortJob, tempDir); sortJob.setInputFormatClass(SequenceFileInputFormat.class); /*InverseMapper由hadoop库提供,作用是实现map()之后的数据对的key和value交换*/ sortJob.setMapperClass(InverseMapper.class); /*将 Reducer 的个数限定为1, 最终输出的结果文件就是一个。*/ sortJob.setNumReduceTasks(1); FileOutputFormat.setOutputPath(sortJob, new Path(otherArgs[1])); sortJob.setOutputKeyClass(IntWritable.class); sortJob.setOutputValueClass(Text.class); /*Hadoop 默认对 IntWritable 按升序排序,而我们需要的是按降序排列。 * 因此我们实现了一个 IntWritableDecreasingComparator 类, * 并指定使用这个自定义的 Comparator 类对输出结果中的 key (词频)进行排序*/ sortJob.setSortComparatorClass(IntWritableDecreasingComparator.class); System.exit(sortJob.waitForCompletion(true) ? 0 : 1); } }finally{ FileSystem.get(conf).deleteOnExit(tempDir); } } } 本文转自xwdreamer博客园博客,原文链接:http://www.cnblogs.com/xwdreamer/archive/2011/01/07/2297044.html,如需转载请自行联系原作者

优秀的个人博客,低调大师

使用Hive UDF和GeoIP库为Hive加入IP识别功能

导读:Hive是基于Hadoop的数据管理系统,作为分析人员的即时分析工具和ETL等工作的执行引擎,对于如今的大数据管理与分析、处理有着非常大的意义。GeoIP是一套IP映射库系统,它定时更新,并且提供了各种语言的API,非常适合在做地域相关数据分析时的一个数据源。 Hive是基于Hadoop的数据管理系统,作为分析人员的即时分析工具和ETL等工作的执行引擎,对于如今的大数据管理与分析、处理有着非常大的意义。GeoIP是一套IP映射库系统,它定时更新,并且提供了各种语言的API,非常适合在做地域相关数据分析时的一个数据源。 UDF是Hive提供的用户自定义函数的接口,通过实现它可以扩展Hive目前已有的内置函数。而为Hive加入一个IP映射函数,我们只需要简单地在UDF中调用GeoIP的Java API即可。 GeoIP的数据文件可以从这里下载:http://www.maxmind.com/download/geoip/database/,由于需要国家和城市的信息,我这里下载的是http://www.maxmind.com/download/geoip/database/GeoLiteCity.dat.gz GeoIP的各种语言的API可以从这里下载:http://www.maxmind.com/download/geoip/api/ 查看文本 copy to clipboard 打印 ? importjava.io.IOException; importorg.apache.hadoop.hive.ql.exec.UDF; importcom.maxmind.geoip.Location; importcom.maxmind.geoip.LookupService; importjava.util.regex.*; publicclassIPToCCextendsUDF{ privatestaticLookupServicecl=null; privatestaticStringipPattern="\\d+\\.\\d+\\.\\d+\\.\\d+"; privatestaticStringipNumPattern="\\d+"; staticLookupServicegetLS()throwsIOException{ Stringdbfile="GeoLiteCity.dat"; if(cl==null) cl=newLookupService(dbfile,LookupService.GEOIP_MEMORY_CACHE); returncl; } /** *@paramstrlike"114.43.181.143" **/ publicStringevaluate(Stringstr){ try{ LocationAl=null; MatchermIP=Pattern.compile(ipPattern).matcher(str); MatchermIPNum=Pattern.compile(ipNumPattern).matcher(str); if(mIP.matches()) Al=getLS().getLocation(str); elseif(mIPNum.matches()) Al=getLS().getLocation(Long.parseLong(str)); returnString.format("%s\t%s",Al.countryName,Al.city); }catch(Exceptione){ e.printStackTrace(); if(cl!=null) cl.close(); returnnull; } } } import java.io.IOException; import org.apache.hadoop.hive.ql.exec.UDF; import com.maxmind.geoip.Location; import com.maxmind.geoip.LookupService; import java.util.regex.*; public class IPToCC extends UDF { private static LookupService cl = null; private static String ipPattern = "\\d+\\.\\d+\\.\\d+\\.\\d+"; private static String ipNumPattern = "\\d+"; static LookupService getLS() throws IOException{ String dbfile = "GeoLiteCity.dat"; if(cl == null) cl = new LookupService(dbfile, LookupService.GEOIP_MEMORY_CACHE); return cl; } /** * @param str like "114.43.181.143" * */ public String evaluate(String str) { try{ Location Al = null; Matcher mIP = Pattern.compile(ipPattern).matcher(str); Matcher mIPNum = Pattern.compile(ipNumPattern).matcher(str); if(mIP.matches()) Al = getLS().getLocation(str); else if(mIPNum.matches()) Al = getLS().getLocation(Long.parseLong(str)); return String.format("%s\t%s", Al.countryName, Al.city); }catch(Exception e){ e.printStackTrace(); if(cl != null) cl.close(); return null; } } } 使用上也非常简单,将以上程序和GeoIP的API程序,一起打成JAR包iptocc.jar,和数据文件(GeoLiteCity.dat)一起放到Hive所在的服务器的一个位置。然后打开Hive执行以下语句: 查看文本 copy to clipboard 打印 ? addfile/tje/path/to/GeoLiteCity.dat; addjar/the/path/to/iptocc.jar; createtemporaryfunctionip2ccas'your.company.udf.IPToCC'; add file /tje/path/to/GeoLiteCity.dat; add jar /the/path/to/iptocc.jar; create temporary function ip2cc as 'your.company.udf.IPToCC'; 然后就可以在Hive的CLI中使用这个函数了,这个函数接收标准的IPv4地址格式的字符串,返回国家和城市信息;同样这个函数也透明地支持长整形的IPv4地址表示格式。如果想在每次启动Hive CLI的时候都自动加载这个自定义函数,可以在hive命令同目录下建立.hiverc文件,在启动写入以上三条语句,重新启动Hive CLI即可;如果在这台服务器上启动Hive Server,使用JDBC连接,执行以上三条语句之后,也可以正常使用这个函数;但是唯一一点不足是,HUE的Beeswax不支持注册用户自定义函数。 虽然不尽完美,但是加入这样一个函数,对于以后做地域相关的即时分析总是提供了一些方便的,还是非常值得加入的。 本文转自茄子_2008博客园博客,原文链接:http://www.cnblogs.com/xd502djj/p/3253411.html,如需转载请自行联系原作者。

资源下载

更多资源
Mario

Mario

马里奥是站在游戏界顶峰的超人气多面角色。马里奥靠吃蘑菇成长,特征是大鼻子、头戴帽子、身穿背带裤,还留着胡子。与他的双胞胎兄弟路易基一起,长年担任任天堂的招牌角色。

Nacos

Nacos

Nacos /nɑ:kəʊs/ 是 Dynamic Naming and Configuration Service 的首字母简称,一个易于构建 AI Agent 应用的动态服务发现、配置管理和AI智能体管理平台。Nacos 致力于帮助您发现、配置和管理微服务及AI智能体应用。Nacos 提供了一组简单易用的特性集,帮助您快速实现动态服务发现、服务配置、服务元数据、流量管理。Nacos 帮助您更敏捷和容易地构建、交付和管理微服务平台。

Rocky Linux

Rocky Linux

Rocky Linux(中文名:洛基)是由Gregory Kurtzer于2020年12月发起的企业级Linux发行版,作为CentOS稳定版停止维护后与RHEL(Red Hat Enterprise Linux)完全兼容的开源替代方案,由社区拥有并管理,支持x86_64、aarch64等架构。其通过重新编译RHEL源代码提供长期稳定性,采用模块化包装和SELinux安全架构,默认包含GNOME桌面环境及XFS文件系统,支持十年生命周期更新。

WebStorm

WebStorm

WebStorm 是jetbrains公司旗下一款JavaScript 开发工具。目前已经被广大中国JS开发者誉为“Web前端开发神器”、“最强大的HTML5编辑器”、“最智能的JavaScript IDE”等。与IntelliJ IDEA同源,继承了IntelliJ IDEA强大的JS部分的功能。

用户登录
用户注册