首页 文章 精选 留言 我的

精选列表

搜索[妙招生态],共10000篇文章
优秀的个人博客,低调大师

大数据分析技术生态圈一览

大数据领域让人晕头转向。为了帮助你,我们决定制作这份厂商图标和目录。它并不是全面列出了这个领域的每家厂商,而是深入探讨大数据分析技术领域。我们希望这份资料新颖、实用。 这是一款面向Hadoop的自助服务式、无数据库模式的大数据分析应用软件。 Platfora 这是一款大数据发现和分析平台。 Qlikview 这是一款引导分析平台。 Sisense 这是一款商业智能软件,专门处理复杂数据的商业智能解决方案。 Sqream 这是一款快速、可扩展的大数据分析SQL数据库。 Splunk 这是一款运维智能平台。 Sumologic 这是一项安全的、专门定制的、基于云的机器数据分析服务。 Actian 这是一款大数据分析平台。 亚马逊Redshift 这是一项PB级云端数据仓库服务。 CitusData 可扩展PostgreSQL。 Exasol 这是一种用于分析数据的大规模并行处理(MPP)内存数据库。 惠普Vertica 这是一款SQL on Hadoop大数据分析平台。 Mammothdb 这是一款与SQL兼容的MPP分析数据库。 微软SQL Server 这是一款关系数据库管理系统。 甲骨文Exadata 这是一款计算和存储综合系统,针对甲骨文数据库软件进行了优化。 SAP HANA 这是一款内存计算平台。 Snowflake 这是一款云数据仓库。 Teradata 这是企业级大数据分析和服务。 数据探查 Apache Drill 这是一款无数据库模式的SQL查询引擎,面向Hadoop、NoSQL和云存储。 Cloudera Impala 这是一款开源大规模并行处理SQL查询引擎。 谷歌BigQuery 这是一项全面托管的NoOps数据分析服务。 Presto 这是一款面向大数据的分布式SQL查询引擎。 Spark 这是一款用于处理大数据的快速通用引擎。 平台/基础设施 亚马逊网络服务(AWS) 提供云计算服务 思科云 提供基础设施即服务 Heroku 为云端应用程序提供平台即服务 Infochimps 提供云服务的大数据解决方案 微软Azure 这是一款企业级云计算平台。 Rackspace 托管专业服务和云计算服务 Softlayer(IBM) 提供云基础设施即服务 数据基础设施 Cask 这是一款面向Hadoop解决方案的开源应用程序平台。 Cloudera 提供基于Hadoop的软件、支持和服务。 Hortonworks 管理HDP――这是一款开源企业Apache Hadoop数据平台。 MAPR 这是面向大数据部署环境的Apache Hadoop技术。 垂直领域应用/数据挖掘 Alpine Data Labs 这是一种高级分析平台,可处理Apache Hadoop和大数据。 R 这是一种免费软件环境,可处理统计计算和图形。 Rapidminer 这是一款开源预测分析平台 SAS 这是一款软件套件,可以挖掘、改动、管理和检索来自众多数据源的数据。 提取、转换和加载(ETL) IBM Datastage 使用一种高性能并行框架,整合多个系统上的数据。 Informatica 这是一款企业数据整合和管理软件。 Kettle-Pentaho Data Integration 提供了强大的提取、转换和加载(ETL)功能。 微软SSIS 这是一款用于构建企业级数据整合和数据转换解决方案的平台。 甲骨文Data Integrator 这是一款全面的数据整合平台。 SAP NetWeaver为整合来自各个数据源的数据提供了灵活方式。 Talend 提供了开源整合软件产品 Cassandra 这是键值数据库和列式数据库的混合解决方案。 CouchBase 这是一款开源分布式NoSQL文档型数据库。 Databricks 这是使用Spark的基于云的大数据处理解决方案。 Datastax 为企业版的Cassandra数据库提供商业支持。 IBM DB2 这是一款可扩展的企业数据库服务器软件。 MemSQL 这是一款分布式内存数据库。 MongoDB 这是一款跨平台的文档型数据库。 MySQL 这是一款流行的开源数据库。 甲骨文 这是一款企业数据库软件套件。 PostgresSQL 这是一款对象关系数据库管理系统。 Riak 这是一款分布式NoSQL数据库。 Splice Machine 这是一款Hadoop关系数据库管理系统。 VoltDB 这是一款内存NewSQL数据库。 Actuate 这是一款嵌入式分析和报表解决方案。 BiBoard 这是一款交互式商业智能仪表板和可视化工具。 Chart.IO 这是面向数据库的企业级分析工具。 IBM Cognos 这是一款商业智能和绩效管理软件。 D3.JS 这是一种使用HTML、SVG和CSS可视化显示数据的JavaScript库。 Highcharts 这是面向互联网的交互式JavaScirpt图表。 Logi Analytics 这是自助服务式、基于Web的商业智能和分析应用软件。 微软Power BI 这是交互式数据探查、可视化和演示工具。 Microstrategy 这是一款企业商业智能和分析软件。 甲骨文Hyperion 这是企业绩效管理和商业智能系统。 Pentaho 这是大数据整合和分析解决方案。 SAP Business Objects 这是商业智能解决方案。 Tableau 这是专注于商业智能的交互式数据可视化产品系列。 Tibco Jaspersoft 这是商业智能套件。 本文作者:佚名 来源:51CTO

优秀的个人博客,低调大师

探秘Hadoop生态12:分布式日志收集系统Flume

这位大侠,这是我的公众号:程序员江湖。分享程序员面试与技术的那些事。 干货满满,关注就送。 在具体介绍本文内容之前,先给大家看一下Hadoop业务的整体开发流程: 从Hadoop的业务开发流程图中可以看出,在大数据的业务处理过程中,对于数据的采集是十分重要的一步,也是不可避免的一步,从而引出我们本文的主角—Flume。本文将围绕Flume的架构、Flume的应用(日志采集)进行详细的介绍。 (一)Flume架构介绍 1、Flume的概念 flume是分布式的日志收集系统,它将各个服务器中的数据收集起来并送到指定的地方去,比如说送到图中的HDFS,简单来说flume就是收集日志的。 2、Event的概念 在这里有必要先介绍一下flume中event的相关概念:flume的核心是把数据从数据源(source)收集过来,在将收集到的数据送到指定的目的地(sink)。为了保证输送的过程一定成功,在送到目的地(sink)之前,会先缓存数据(channel),待数据真正到达目的地(sink)后,flume在删除自己缓存的数据。 在整个数据的传输的过程中,流动的是event,即事务保证是在event级别进行的。那么什么是event呢?—–event将传输的数据进行封装,是flume传输数据的基本单位,如果是文本文件,通常是一行记录,event也是事务的基本单位。event从source,流向channel,再到sink,本身为一个字节数组,并可携带headers(头信息)信息。event代表着一个数据的最小完整单元,从外部数据源来,向外部的目的地去。 为了方便大家理解,给出一张event的数据流向图: 一个完整的event包括:event headers、event body、event信息(即文本文件中的单行记录),如下所以: 其中event信息就是flume收集到的日记记录。 3、flume架构介绍 flume之所以这么神奇,是源于它自身的一个设计,这个设计就是agent,agent本身是一个java进程,运行在日志收集节点—所谓日志收集节点就是服务器节点。 agent里面包含3个核心的组件:source—->channel—–>sink,类似生产者、仓库、消费者的架构。 source:source组件是专门用来收集数据的,可以处理各种类型、各种格式的日志数据,包括avro、thrift、exec、jms、spooling directory、netcat、sequence generator、syslog、http、legacy、自定义。 channel:source组件把数据收集来以后,临时存放在channel中,即channel组件在agent中是专门用来存放临时数据的——对采集到的数据进行简单的缓存,可以存放在memory、jdbc、file等等。 sink:sink组件是用于把数据发送到目的地的组件,目的地包括hdfs、logger、avro、thrift、ipc、file、null、hbase、solr、自定义。 4、flume的运行机制 flume的核心就是一个agent,这个agent对外有两个进行交互的地方,一个是接受数据的输入——source,一个是数据的输出sink,sink负责将数据发送到外部指定的目的地。source接收到数据之后,将数据发送给channel,chanel作为一个数据缓冲区会临时存放这些数据,随后sink会将channel中的数据发送到指定的地方—-例如HDFS等,注意:只有在sink将channel中的数据成功发送出去之后,channel才会将临时数据进行删除,这种机制保证了数据传输的可靠性与安全性。 5、flume的广义用法 flume之所以这么神奇—-其原因也在于flume可以支持多级flume的agent,即flume可以前后相继,例如sink可以将数据写到下一个agent的source中,这样的话就可以连成串了,可以整体处理了。flume还支持扇入(fan-in)、扇出(fan-out)。所谓扇入就是source可以接受多个输入,所谓扇出就是sink可以将数据输出多个目的地destination中。 (二)flume应用—日志采集 对于flume的原理其实很容易理解,我们更应该掌握flume的具体使用方法,flume提供了大量内置的Source、Channel和Sink类型。而且不同类型的Source、Channel和Sink可以自由组合—–组合方式基于用户设置的配置文件,非常灵活。比如:Channel可以把事件暂存在内存里,也可以持久化到本地硬盘上。Sink可以把日志写入HDFS, HBase,甚至是另外一个Source等等。下面我将用具体的案例详述flume的具体用法。 其实flume的用法很简单—-书写一个配置文件,在配置文件当中描述source、channel与sink的具体实现,而后运行一个agent实例,在运行agent实例的过程中会读取配置文件的内容,这样flume就会采集到数据。 配置文件的编写原则: 1>从整体上描述代理agent中sources、sinks、channels所涉及到的组件 # Name the components on this agent a1.sources = r1 a1.sinks = k1 a1.channels = c1 1 2 3 4 2>详细描述agent中每一个source、sink与channel的具体实现:即在描述source的时候,需要 指定source到底是什么类型的,即这个source是接受文件的、还是接受http的、还是接受thrift 的;对于sink也是同理,需要指定结果是输出到HDFS中,还是Hbase中啊等等;对于channel 需要指定是内存啊,还是数据库啊,还是文件啊等等。 # Describe/configure the source a1.sources.r1.type = netcat a1.sources.r1.bind = localhost a1.sources.r1.port = 44444 # Describe the sink a1.sinks.k1.type = logger # Use a channel which buffers events in memory a1.channels.c1.type = memory a1.channels.c1.capacity = 1000 a1.channels.c1.transactionCapacity = 100 1 2 3 4 5 6 7 8 9 10 11 12 3>通过channel将source与sink连接起来 # Bind the source and sink to the channel a1.sources.r1.channels = c1 a1.sinks.k1.channel = c1 1 2 3 启动agent的shell操作: flume-ng agent -n a1 -c ../conf -f ../conf/example.file -Dflume.root.logger=DEBUG,console 1 2 参数说明: -n 指定agent名称(与配置文件中代理的名字相同) -c 指定flume中配置文件的目录 -f 指定配置文件 -Dflume.root.logger=DEBUG,console 设置日志等级 具体案例: 案例1: NetCat Source:监听一个指定的网络端口,即只要应用程序向这个端口里面写数据,这个source组件就可以获取到信息。 其中 Sink:logger Channel:memory flume官网中NetCat Source描述: Property Name Default Description channels – type – The component type name, needs to be netcat bind – 日志需要发送到的主机名或者Ip地址,该主机运行着netcat类型的source在监听 port – 日志需要发送到的端口号,该端口号要有netcat类型的source在监听 1 2 3 4 5 a) 编写配置文件: # Name the components on this agent a1.sources = r1 a1.sinks = k1 a1.channels = c1 # Describe/configure the source a1.sources.r1.type = netcat a1.sources.r1.bind = 192.168.80.80 a1.sources.r1.port = 44444 # Describe the sink a1.sinks.k1.type = logger # Use a channel which buffers events in memory a1.channels.c1.type = memory a1.channels.c1.capacity = 1000 a1.channels.c1.transactionCapacity = 100 # Bind the source and sink to the channel a1.sources.r1.channels = c1 a1.sinks.k1.channel = c1 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 b) 启动flume agent a1 服务端 flume-ng agent -n a1 -c ../conf -f ../conf/netcat.conf -Dflume.root.logger=DEBUG,console 1 c) 使用telnet发送数据 telnet 192.168.80.80 44444 big data world!(windows中运行的) 1 d) 在控制台上查看flume收集到的日志数据: 案例2:NetCat Source:监听一个指定的网络端口,即只要应用程序向这个端口里面写数据,这个source组件就可以获取到信息。 其中 Sink:hdfs Channel:file (相比于案例1的两个变化) flume官网中HDFS Sink的描述: a) 编写配置文件: # Name the components on this agent a1.sources = r1 a1.sinks = k1 a1.channels = c1 # Describe/configure the source a1.sources.r1.type = netcat a1.sources.r1.bind = 192.168.80.80 a1.sources.r1.port = 44444 # Describe the sink a1.sinks.k1.type = hdfs a1.sinks.k1.hdfs.path = hdfs://hadoop80:9000/dataoutput a1.sinks.k1.hdfs.writeFormat = Text a1.sinks.k1.hdfs.fileType = DataStream a1.sinks.k1.hdfs.rollInterval = 10 a1.sinks.k1.hdfs.rollSize = 0 a1.sinks.k1.hdfs.rollCount = 0 a1.sinks.k1.hdfs.filePrefix = %Y-%m-%d-%H-%M-%S a1.sinks.k1.hdfs.useLocalTimeStamp = true # Use a channel which buffers events in file a1.channels.c1.type = file a1.channels.c1.checkpointDir = /usr/flume/checkpoint a1.channels.c1.dataDirs = /usr/flume/data # Bind the source and sink to the channel a1.sources.r1.channels = c1 a1.sinks.k1.channel = c1 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 b) 启动flume agent a1 服务端 flume-ng agent -n a1 -c ../conf -f ../conf/netcat.conf -Dflume.root.logger=DEBUG,console 1 c) 使用telnet发送数据 telnet 192.168.80.80 44444 big data world!(windows中运行的) 1 d) 在HDFS中查看flume收集到的日志数据: 案例3:Spooling Directory Source:监听一个指定的目录,即只要应用程序向这个指定的目录中添加新的文件,source组件就可以获取到该信息,并解析该文件的内容,然后写入到channle。写入完成后,标记该文件已完成或者删除该文件。其中 Sink:logger Channel:memory flume官网中Spooling Directory Source描述: Property Name Default Description channels – type – The component type name, needs to be spooldir. spoolDir – Spooling Directory Source监听的目录 fileSuffix .COMPLETED 文件内容写入到channel之后,标记该文件 deletePolicy never 文件内容写入到channel之后的删除策略: never or immediate fileHeader false Whether to add a header storing the absolute path filename. ignorePattern ^$ Regular expression specifying which files to ignore (skip) interceptors – 指定传输中event的head(头信息),常用timestamp 1 2 3 4 5 6 7 8 9 Spooling Directory Source的两个注意事项: ①If a file is written to after being placed into the spooling directory, Flume will print an error to its log file and stop processing. 即:拷贝到spool目录下的文件不可以再打开编辑 ②If a file name is reused at a later time, Flume will print an error to its log file and stop processing. 即:不能将具有相同文件名字的文件拷贝到这个目录下 1 2 3 4 a) 编写配置文件: # Name the components on this agent a1.sources = r1 a1.sinks = k1 a1.channels = c1 # Describe/configure the source a1.sources.r1.type = spooldir a1.sources.r1.spoolDir = /usr/local/datainput a1.sources.r1.fileHeader = true a1.sources.r1.interceptors = i1 a1.sources.r1.interceptors.i1.type = timestamp # Describe the sink a1.sinks.k1.type = logger # Use a channel which buffers events in memory a1.channels.c1.type = memory a1.channels.c1.capacity = 1000 a1.channels.c1.transactionCapacity = 100 # Bind the source and sink to the channel a1.sources.r1.channels = c1 a1.sinks.k1.channel = c1 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 b) 启动flume agent a1 服务端 flume-ng agent -n a1 -c ../conf -f ../conf/spool.conf -Dflume.root.logger=DEBUG,console 1 c) 使用cp命令向Spooling Directory 中发送数据 cp datafile /usr/local/datainput (注:datafile中的内容为:big data world!) 1 d) 在控制台上查看flume收集到的日志数据: 从控制台显示的结果可以看出event的头信息中包含了时间戳信息。 同时我们查看一下Spooling Directory中的datafile信息—-文件内容写入到channel之后,该文件被标记了: [root@hadoop80 datainput]# ls datafile.COMPLETED 1 2 案例4:Spooling Directory Source:监听一个指定的目录,即只要应用程序向这个指定的目录中添加新的文件,source组件就可以获取到该信息,并解析该文件的内容,然后写入到channle。写入完成后,标记该文件已完成或者删除该文件。 其中 Sink:hdfs Channel:file (相比于案例3的两个变化) a) 编写配置文件: # Name the components on this agent a1.sources = r1 a1.sinks = k1 a1.channels = c1 # Describe/configure the source a1.sources.r1.type = spooldir a1.sources.r1.spoolDir = /usr/local/datainput a1.sources.r1.fileHeader = true a1.sources.r1.interceptors = i1 a1.sources.r1.interceptors.i1.type = timestamp # Describe the sink # Describe the sink a1.sinks.k1.type = hdfs a1.sinks.k1.hdfs.path = hdfs://hadoop80:9000/dataoutput a1.sinks.k1.hdfs.writeFormat = Text a1.sinks.k1.hdfs.fileType = DataStream a1.sinks.k1.hdfs.rollInterval = 10 a1.sinks.k1.hdfs.rollSize = 0 a1.sinks.k1.hdfs.rollCount = 0 a1.sinks.k1.hdfs.filePrefix = %Y-%m-%d-%H-%M-%S a1.sinks.k1.hdfs.useLocalTimeStamp = true # Use a channel which buffers events in file a1.channels.c1.type = file a1.channels.c1.checkpointDir = /usr/flume/checkpoint a1.channels.c1.dataDirs = /usr/flume/data # Bind the source and sink to the channel a1.sources.r1.channels = c1 a1.sinks.k1.channel = c1 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 b) 启动flume agent a1 服务端 flume-ng agent -n a1 -c ../conf -f ../conf/spool.conf -Dflume.root.logger=DEBUG,console 1 c) 使用cp命令向Spooling Directory 中发送数据 cp datafile /usr/local/datainput (注:datafile中的内容为:big data world!) 1 d) 在控制台上可以参看sink的运行进度日志: d) 在HDFS中查看flume收集到的日志数据: 从案例1与案例2、案例3与案例4的对比中我们可以发现:flume的配置文件在编写的过程中是非常灵活的。 案例5:Exec Source:监听一个指定的命令,获取一条命令的结果作为它的数据源 常用的是tail -F file指令,即只要应用程序向日志(文件)里面写数据,source组件就可以获取到日志(文件)中最新的内容 。 其中 Sink:hdfs Channel:file 这个案列为了方便显示Exec Source的运行效果,结合Hive中的external table进行来说明。 a) 编写配置文件: # Name the components on this agent a1.sources = r1 a1.sinks = k1 a1.channels = c1 # Describe/configure the source a1.sources.r1.type = exec a1.sources.r1.command = tail -F /usr/local/log.file # Describe the sink a1.sinks.k1.type = hdfs a1.sinks.k1.hdfs.path = hdfs://hadoop80:9000/dataoutput a1.sinks.k1.hdfs.writeFormat = Text a1.sinks.k1.hdfs.fileType = DataStream a1.sinks.k1.hdfs.rollInterval = 10 a1.sinks.k1.hdfs.rollSize = 0 a1.sinks.k1.hdfs.rollCount = 0 a1.sinks.k1.hdfs.filePrefix = %Y-%m-%d-%H-%M-%S a1.sinks.k1.hdfs.useLocalTimeStamp = true # Use a channel which buffers events in file a1.channels.c1.type = file a1.channels.c1.checkpointDir = /usr/flume/checkpoint a1.channels.c1.dataDirs = /usr/flume/data # Bind the source and sink to the channel a1.sources.r1.channels = c1 a1.sinks.k1.channel = c1 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 b)在hive中建立外部表—–hdfs://hadoop80:9000/dataoutput的目录,方便查看日志捕获内容 hive> create external table t1(infor string) > row format delimited > fields terminated by '\t' > location '/dataoutput/'; OK Time taken: 0.284 seconds 1 2 3 4 5 6 c) 启动flume agent a1 服务端 flume-ng agent -n a1 -c ../conf -f ../conf/exec.conf -Dflume.root.logger=DEBUG,console 1 d) 使用echo命令向/usr/local/datainput 中发送数据 echo big data > log.file 1 d) 在HDFS和Hive分别中查看flume收集到的日志数据: hive> select * from t1; OK big data Time taken: 0.086 seconds 1 2 3 4 e)使用echo命令向/usr/local/datainput 中在追加一条数据 echo big data world! >> log.file 1 d) 在HDFS和Hive再次分别中查看flume收集到的日志数据: hive> select * from t1; OK big data big data world! Time taken: 0.511 seconds 1 2 3 4 5 总结Exec source:Exec source和Spooling Directory Source是两种常用的日志采集的方式,其中Exec source可以实现对日志的实时采集,Spooling Directory Source在对日志的实时采集上稍有欠缺,尽管Exec source可以实现对日志的实时采集,但是当Flume不运行或者指令执行出错时,Exec source将无法收集到日志数据,日志会出现丢失,从而无法保证收集日志的完整性。 案例6:Avro Source:监听一个指定的Avro 端口,通过Avro 端口可以获取到Avro client发送过来的文件 。即只要应用程序通过Avro 端口发送文件,source组件就可以获取到该文件中的内容。 其中 Sink:hdfs Channel:file (注:Avro和Thrift都是一些序列化的网络端口–通过这些网络端口可以接受或者发送信息,Avro可以发送一个给定的文件给Flume,Avro 源使用AVRO RPC机制) Avro Source运行原理如下图: flume官网中Avro Source的描述: Property Name Default Description channels – type – The component type name, needs to be avro bind – 日志需要发送到的主机名或者ip,该主机运行着ARVO类型的source port – 日志需要发送到的端口号,该端口要有ARVO类型的source在监听 1 2 3 4 5 1)编写配置文件 # Name the components on this agent a1.sources = r1 a1.sinks = k1 a1.channels = c1 # Describe/configure the source a1.sources.r1.type = avro a1.sources.r1.bind = 192.168.80.80 a1.sources.r1.port = 4141 # Describe the sink a1.sinks.k1.type = hdfs a1.sinks.k1.hdfs.path = hdfs://hadoop80:9000/dataoutput a1.sinks.k1.hdfs.writeFormat = Text a1.sinks.k1.hdfs.fileType = DataStream a1.sinks.k1.hdfs.rollInterval = 10 a1.sinks.k1.hdfs.rollSize = 0 a1.sinks.k1.hdfs.rollCount = 0 a1.sinks.k1.hdfs.filePrefix = %Y-%m-%d-%H-%M-%S a1.sinks.k1.hdfs.useLocalTimeStamp = true # Use a channel which buffers events in file a1.channels.c1.type = file a1.channels.c1.checkpointDir = /usr/flume/checkpoint a1.channels.c1.dataDirs = /usr/flume/data # Bind the source and sink to the channel a1.sources.r1.channels = c1 a1.sinks.k1.channel = c1 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 b) 启动flume agent a1 服务端 flume-ng agent -n a1 -c ../conf -f ../conf/avro.conf -Dflume.root.logger=DEBUG,console 1 c)使用avro-client发送文件 flume-ng avro-client -c ../conf -H 192.168.80.80 -p 4141 -F /usr/local/log.file 1 注:log.file文件中的内容为: [root@hadoop80 local]# more log.file big data big data world! 1 2 3 d) 在HDFS中查看flume收集到的日志数据: 通过上面的几个案例,我们可以发现:flume配置文件的书写是相当灵活的—-不同类型的Source、Channel和Sink可以自由组合! 最后对上面用的几个flume source进行适当总结: ① NetCat Source:监听一个指定的网络端口,即只要应用程序向这个端口里面写数据,这个source组件 就可以获取到信息。 ②Spooling Directory Source:监听一个指定的目录,即只要应用程序向这个指定的目录中添加新的文 件,source组件就可以获取到该信息,并解析该文件的内容,然后写入到channle。写入完成后,标记 该文件已完成或者删除该文件。 ③Exec Source:监听一个指定的命令,获取一条命令的结果作为它的数据源 常用的是tail -F file指令,即只要应用程序向日志(文件)里面写数据,source组件就可以获取到日志(文件)中最新的内容 。 ④Avro Source:监听一个指定的Avro 端口,通过Avro 端口可以获取到Avro client发送过来的文件 。即只要应用程序通过Avro 端口发送文件,source组件就可以获取到该文件中的内容。 如有问题,欢迎留言指正!

资源下载

更多资源
Mario

Mario

马里奥是站在游戏界顶峰的超人气多面角色。马里奥靠吃蘑菇成长,特征是大鼻子、头戴帽子、身穿背带裤,还留着胡子。与他的双胞胎兄弟路易基一起,长年担任任天堂的招牌角色。

腾讯云软件源

腾讯云软件源

为解决软件依赖安装时官方源访问速度慢的问题,腾讯云为一些软件搭建了缓存服务。您可以通过使用腾讯云软件源站来提升依赖包的安装速度。为了方便用户自由搭建服务架构,目前腾讯云软件源站支持公网访问和内网访问。

Nacos

Nacos

Nacos /nɑ:kəʊs/ 是 Dynamic Naming and Configuration Service 的首字母简称,一个易于构建 AI Agent 应用的动态服务发现、配置管理和AI智能体管理平台。Nacos 致力于帮助您发现、配置和管理微服务及AI智能体应用。Nacos 提供了一组简单易用的特性集,帮助您快速实现动态服务发现、服务配置、服务元数据、流量管理。Nacos 帮助您更敏捷和容易地构建、交付和管理微服务平台。

WebStorm

WebStorm

WebStorm 是jetbrains公司旗下一款JavaScript 开发工具。目前已经被广大中国JS开发者誉为“Web前端开发神器”、“最强大的HTML5编辑器”、“最智能的JavaScript IDE”等。与IntelliJ IDEA同源,继承了IntelliJ IDEA强大的JS部分的功能。

用户登录
用户注册