首页 文章 精选 留言 我的

精选列表

搜索[文本预处理],共10000篇文章
优秀的个人博客,低调大师

java中利用hanlp比较两个文本相似度的步骤

使用 HanLP - 汉语言处理包 来处理,他能处理很多事情,如分词、调用分词器、命名实体识别、人名识别、地名识别、词性识别、篇章理解、关键词提取、简繁拼音转换、拼音转换、根据输入智能推荐、自定义分词器 使用很简单,只要引入hanlp.jar包,便可处理(新版本的hanlp安装包可以去github下载安装),下面是某位大神的操作截图:

优秀的个人博客,低调大师

sklearn.feature_extraction.text.CountVectorizer提取文本特征,将文档词块化

sklearn.feature_extraction.text. CountVectorizer ( input=u'content' , encoding=u'utf-8' , decode_error=u'strict' , strip_accents=None , lowercase=True , preprocessor=None , tokenizer=None , stop_words=None , token_pattern=u'(? u)\b\w\w+\b' , ngram_range=(1 , 1) , analyzer=u'word' , max_df=1.0 , min_df=1 , max_features=None , vocabulary=None , binary=False , dtype=<type 'numpy.int64'> ) 作用:Convert a collection of text documents to a matrix of token counts(计算词汇的数量,即tf);结果由scipy.sparse.coo_matrix进行稀疏表示。 看下参数就知道CountVectorizer在提取tf时都做了什么: strip_accents : {‘ascii’, ‘unicode’, None}:是否除去“音调”,不知道什么是“音调”?看:http://textmechanic.com/?reqp=1&reqr=nzcdYz9hqaSbYaOvrt== lowercase : boolean, True by default:计算tf前,先将所有字符转化为小写。 这个参数一般为True。 preprocessor : callable or None (default):复写thepreprocessing (string transformation) stage,但保留tokenizing and n-grams generation steps. 这个参数可以自己写。 tokenizer : callable or None (default):复写the string tokenization step,但保留preprocessingand n-grams generation steps. 这个参数可以自己写。 stop_words : string {‘english’}, list, or None (default):如果是‘english’, a built-in stop word list for English is used。如果是a list,那么最终的tokens中将去掉list中的所有的stop word。如果是None, 不处理停顿词;但 参数 max_df 可以设置为 [0.7, 1.0) 之间,进而根据 intra corpus document frequency(df) of terms自动detect and filter stop words。这个参数要根据自己的需求调整。 token_pattern : string:正则表达式,默认筛选长度大于等于2的字母和数字混合字符(select tokens of 2 or more alphanumeric characters),参数analyzer设置为word时才有效。 ngram_range : tuple (min_n, max_n):n-values值得上下界,默认是 ngram_range=(1 , 1), 该范围之内的n元feature都会被提取出来! 这个参数要根据自己的需求调整。 analyzer : string, {‘word’, ‘char’, ‘char_wb’} or callable:特征基于wordn-grams还是character n-grams。如果是callable是自己复写的从the raw, unprocessed input提取特征的函数。 max_df : float in range [0.0, 1.0] or int, default=1.0: min_df : float in range [0.0, 1.0] or int, default=1:按比例,或绝对数量删除df超过max_df或者df小于min_df的word tokens。有效的前提是参数vocabulary设置成Node。 max_features : int or None, default=None:选择tf最大的max_features个特征。有效的前提是参数vocabulary设置成Node。 vocabulary : Mapping or iterable, optional:自定义的特征word tokens,如果不是None,则只计算vocabulary中的词的tf。 还是设为None靠谱。 binary : boolean, default=False:如果是True,tf的值只有0和1,表示出现和不出现,useful for discrete probabilistic models that model binary events rather than integer counts.。 dtype : type, optional:Type of the matrix returned by fit_transform() or transform().。

优秀的个人博客,低调大师

Linux高级文本处理之gawk变量的操作符(三)

一、变量 Awk 变量以字母开头,后续字符可以是数字、字母、或下划线。关键字不能用作 awk 变量。awk 变量可以直接使用而不需事先声明。 如果要初始化变量,最好在BEGIN 区域内做,它只会执行一次。Awk 中没有数据类型的概念,一个 awk 变量是 number 还是 string 取决于该变量所处的上下文。 实例1:使用”total”便是用户建立的用来存储公司所有雇员工资总和的变量。 [root@localhost~]#catemp4 101,JohnDoe,CEO,10000 102,JasonSmith,ITManager,5000 103,RajReddy,Sysadmin,4500 104,AnandRam,Developer,4500 105,JaneMiller,SalesManager,3000 [root@localhost~]#catemp.awk BEGIN{ FS=","; total=0; } { print$2"'ssalaryis:"$4; total=total+$4; } END{ print"---\nTotalcompanysalary=$"total; } [root@localhost~]#awk-femp.awkemp4 JohnDoe'ssalaryis:10000 JasonSmith'ssalaryis:5000 RajReddy'ssalaryis:4500 AnandRam'ssalaryis:4500 JaneMiller'ssalaryis:3000 --- Totalcompanysalary=$27000 awk自定义变量的方法: 1.借助-v选项,可以将外部值(并非来自stdin)传递给awk 实例2: [root@localhost~]#awk-vvar="young"'BEGIN{printvar,"\n","---"}{printvar}'./num young --- young young young 2.通过VAR=value的方式定义 实例3: [root@localhost~]#awk'{printv1,v2}'v1="young"v2="geek"./num younggeek younggeek younggeek 二、一元操作符 1.取正取反 只接受单个操作数的操作符叫做一元操作符。 实例1:取反操作 [root@localhost~]#catemp4 101,JohnDoe,CEO,10000 102,JasonSmith,ITManager,5000 103,RajReddy,Sysadmin,4500 104,AnandRam,Developer,4500 105,JaneMiller,SalesManager,3000 [root@localhost~]#awk-F,'{print-$4}'emp4 -10000 -5000 -4500 -4500 -3000 注意:取反只对数值类数据生效,字符串取反结果全部为0. 实例2: [root@localhost~]#catnum -1 -2 -3 [root@localhost~]#awk'{print+$1}'num -1 -2 -3 [root@localhost~]#awk'{print-$1}'num 1 2 3 2.自增自减 VAR1=++VAR或者VAR1=--VAR,表示VAR先增加或者减去1再赋值给VAR1,VAR1=VAR++或VAR1=VAR--,表示先将VAR赋值给VAR1,VAR再增减或者减去1. 实例1:前自加子减 [root@localhost~]#awk-F,'{print++$4}'emp4#前自加 10001 5001 4501 4501 3001 [root@localhost~]#awk-F,'{print--$4}'emp4#前自减 9999 4999 4499 4499 实例2:后自加子减 [root@localhost~]#awk-F,'{$4--;print$4}'emp4#后自减 9999 4999 4499 4499 2999 [root@localhost~]#awk-F,'{$4++;print$4}'emp4#后自加 10001 5001 4501 4501 3001 实例3:打印所有可登陆 shell 的用户总数: [root@localhost~]#awk-F':' >'$NF~/\/bin\/bash/{n++} >END{printn}'/etc/passwd 10 [root@localhost~]#grep-c'/bin/bash$'/etc/passwd 10 实例4: [root@localhost~]#catnum 1 2 1 1 3 4 2 [root@localhost~]#awk'/1/{printNF}'num 1 1 1 [root@localhost~]#awk'/1/{n++}END{printn}'num 3 [root@localhost~]#awk'/2/{n++}END{printn}'num 2 [root@localhost~]#awk'/3/{n++}END{printn}'num 1 [root@localhost~]#awk'/4/{n++}END{printn}'num 1 三、算术运算符 需要两个操作数的操作符,成为二元操作符。 Awk 中有多种基本二元操作符(如算术操作符、 字符串操作符、赋值操作符,等等)。 实例1:将每件单独的商品价格减少 20% 并且将每件单独的商品的数量减少 1 [root@localhost~]#catitems.txt 101,HDCamcorder,Video,210,10 102,Refrigerator,Appliance,850,2 103,MP3Player,Audio,270,15 104,TennisRacket,Sports,190,20 105,LaserPrinter,Office,475,5 [root@localhost~]#catdis.awk BEGIN{ FS=","; OFS=","; discount=0 } { discount=$4*20/100; print$1,$2,$3,$4-discount,$5-1 } [root@localhost~]#awk-fdis.awkitems.txt 101,HDCamcorder,Video,168,9 102,Refrigerator,Appliance,680,1 103,MP3Player,Audio,216,14 104,TennisRacket,Sports,152,19 105,LaserPrinter,Office,380,4 实例2:只打印偶数行 [root@localhost~]#awk'NR%2==0'items.txt 102,Refrigerator,Appliance,850,2 104,TennisRacket,Sports,190,20 四、字符串操作符 (空格)是连接字符串的操作符。 实例1: [root@localhost~]#catstr.awk BEGIN{ FS=","; OFS=","; str1="Audio"; str2="Video"; nustr="100"; str3=str1str2; print"Concatenatestringis:"str3; nustr=nustr+1; print"Strtonu:"nustr; } [root@localhost~]#awk-fstr.awkitems.txt Concatenatestringis:AudioVideo Strtonu:101 四、赋值操作符 实例1: [root@localhost~]#catfz.awk BEGIN{ FS=","; OFS=","; total1=total2=total3=total4=total5=10; total1+=5;printtotal1; total2-=5;printtotal2; total3*=5;printtotal3; total4/=5;printtotal4; total5%=5;printtotal5; } [root@localhost~]#awk-ffz.awkitems.txt 15 5 50 2 0 实例2:打印商品清单 [root@localhost~]#catitems.txt 101,HDCamcorder,Video,210,10 102,Refrigerator,Appliance,850,2 103,MP3Player,Audio,270,15 104,TennisRacket,Sports,190,20 105,LaserPrinter,Office,475,5 [root@localhost~]#awk-F,' >BEGIN{total=0}{total+=$5} >END{print"TotalQutantity:"total}'items.txt TotalQutantity:52 五、比较操作符 实例1:打印数量小于等于临界值 5 的商品信息 [root@localhost~]#awk-F,'$5<=5'items.txt 102,Refrigerator,Appliance,850,2 105,LaserPrinter,Office,475,5 实例2:打印编号为 103 的商品信息 [root@localhost~]#awk-F,'$1==103'items.txt 103,MP3Player,Audio,270,15 实例3:打印除 Video 以外的所有商品 [root@localhost~]#awk-F,'$3!="Video"'items.txt 102,Refrigerator,Appliance,850,2 103,MP3Player,Audio,270,15 104,TennisRacket,Sports,190,20 105,LaserPrinter,Office,475,5 实例4:同实例3,但只打印描述信息 [root@localhost~]#awk-F,'$3!="Video"{print$2}'items.txt Refrigerator MP3Player TennisRacket LaserPrinter 实例5:打印价钱低于 900 或者数量小于等于临界值 5 的商品信息 [root@localhost~]#awk-F,'$5<=5||$4<900'items.txt 101,HDCamcorder,Video,210,10 102,Refrigerator,Appliance,850,2 103,MP3Player,Audio,270,15 104,TennisRacket,Sports,190,20 105,LaserPrinter,Office,475,5 实例6:打印/etc/password 中最大的 UID(以及其所在的整行)。 [root@localhost~]#awk-F':'' >$3>maxuid >{maxuid=$3;maxline=$0} >END{printmaxuid,maxline}'/etc/passwd 1009user3:x:1009:1010::/home/user3:/bin/bash 实例8:打印/etc/passwd 中 UID 和 GROUP ID 相同的用户信息 [root@localhost~]#awk-F':''$3==$4'/etc/passwd root:x:0:0:young,geek,010110110,0101101101:/root:/bin/bash bin:x:1:1:bin:/bin:/sbin/nologin daemon:x:2:2:daemon:/sbin:/sbin/nologin nobody:x:99:99:Nobody:/:/sbin/nologin dbus:x:81:81:Systemmessagebus:/:/sbin/nologin vcsa:x:69:69:virtualconsolememoryowner:/dev:/sbin/nologin abrt:x:173:173::/etc/abrt:/sbin/nologin 实例9:打印/etc/passwd 中 UID >= 100 并且用户的 shell 是/bin/sh 的用户 [root@localhost~]#awk-F:'$3>=100&&$7=="/bin/sh"'/etc/passwd user1:x:800:800:testuser:/none:/bin/sh 或者: [root@localhost~]#awk-F':''$3>=100&&$NF~/\/bin\/sh/'/etc/passwd user1:x:800:800:testuser:/none:/bin/sh#正则表达式模式匹配 实例10:打印/etc/passwd 中没有注释信息(第 5 个字段)的用户 [root@localhost~]#awk-F:'$5==""'/etc/passwd abrt:x:173:173::/etc/abrt:/sbin/nologin ntp:x:38:38::/etc/ntp:/sbin/nologin postfix:x:89:89::/var/spool/postfix:/sbin/nologin tcpdump:x:72:72::/:/sbin/nologin sys:x:498:1001::/home/sys:/bin/bash natasha:x:1006:1007::/home/natasha:/bin/bash harry:x:1007:1008::/home/harry:/bin/bash sarah:x:497:497::/home/sarah:/bin/nologin 实例11:使用取反(!)运算符打印奇数行 [root@localhost~]#seq10|awk'i=!i' 1 3 5 7 9 实例12:打印偶数行 [root@localhost~]#seq10|awk-vi=1'i=!i' 2 4 6 8 10 六、正则表达式操作符 实例1: [root@localhost~]#catitems.txt 101,HDCamcorder,Video,210,10 102,Refrigerator,Appliance,850,2 103,MP3Player,Audio,270,15 104,TennisRacket,Sports,190,20 105,LaserPrinter,Office,475,5 [root@localhost~]#awk-F,'$2=="Tennis"'items.txt#精确匹配 [root@localhost~]#awk-F,'$2~"Tennis"'items.txt#模糊匹配 104,TennisRacket,Sports,190,20 实例2: [root@localhost~]#awk-F,'$2!~"Tennis"'items.txt#不匹配 101,HDCamcorder,Video,210,10 102,Refrigerator,Appliance,850,2 103,MP3Player,Audio,270,15 105,LaserPrinter,Office,475,5 实例3:打印 shell 为/bin/bash 的用户的总数,如果最后一个字段包含”/bin/bash”,则变量n 增加 1 [root@localhost~]#grep-c'/bin/bash'/etc/passwd 9 [root@localhost~]#awk-F:'$NF~/\/bin\/bash/{n++}END{printn}'/etc/passwd 9 补充说明: awk PATTERN模式的其他形式: empy:空模式,匹配所有行 关系表达式,表达式结果非0为真,则执行后面body中语句;0则为假,不执行。例如 awk-F:'$3>=500{print$1,$3}'/etc/passwd 3.行范围,类似sed或vim中的地址定界: /startpattern/,/endpattern/ 注意:不支持直接给出数字格式 实例: [root@localhost~]#awk'/^root/,/^mail/{print$0}'/etc/passwd root:x:0:0:young,geek,010110110,0101101101:/root:/bin/bash bin:x:1:1:bin:/bin:/sbin/nologin daemon:x:2:2:daemon:/sbin:/sbin/nologin adm:x:3:4:adm:/var/adm:/sbin/nologin lp:x:4:7:lp:/var/spool/lpd:/sbin/nologin sync:x:5:0:sync:/sbin:/bin/sync shutdown:x:6:0:shutdown:/sbin:/sbin/shutdown halt:x:7:0:halt:/sbin:/sbin/halt mail:x:8:12:mail:/var/spool/mail:/sbin/nologin

资源下载

更多资源
Mario

Mario

马里奥是站在游戏界顶峰的超人气多面角色。马里奥靠吃蘑菇成长,特征是大鼻子、头戴帽子、身穿背带裤,还留着胡子。与他的双胞胎兄弟路易基一起,长年担任任天堂的招牌角色。

腾讯云软件源

腾讯云软件源

为解决软件依赖安装时官方源访问速度慢的问题,腾讯云为一些软件搭建了缓存服务。您可以通过使用腾讯云软件源站来提升依赖包的安装速度。为了方便用户自由搭建服务架构,目前腾讯云软件源站支持公网访问和内网访问。

Sublime Text

Sublime Text

Sublime Text具有漂亮的用户界面和强大的功能,例如代码缩略图,Python的插件,代码段等。还可自定义键绑定,菜单和工具栏。Sublime Text 的主要功能包括:拼写检查,书签,完整的 Python API , Goto 功能,即时项目切换,多选择,多窗口等等。Sublime Text 是一个跨平台的编辑器,同时支持Windows、Linux、Mac OS X等操作系统。

WebStorm

WebStorm

WebStorm 是jetbrains公司旗下一款JavaScript 开发工具。目前已经被广大中国JS开发者誉为“Web前端开发神器”、“最强大的HTML5编辑器”、“最智能的JavaScript IDE”等。与IntelliJ IDEA同源,继承了IntelliJ IDEA强大的JS部分的功能。

用户登录
用户注册