首页 文章 精选 留言 我的

精选列表

搜索[es],共3901篇文章
优秀的个人博客,低调大师

ES doc_values的来源,field data——就是doc->terms的正向索引啊,不过它是在查询阶段通过读取倒排索引loadi...

Support in the Wild: My Biggest Elasticsearch Problem at Scale Java Heap Pressure Elasticsearch has so many wildly different use cases that I could not write a reasonably short blog post describing what can and cannot consume memory. However, there is one thing that constantly stands out above all of the other concerns that you might have while running an Elasticsearch cluster at scale. For the users that I help,fielddatais the problem that is the most likely to cause their cluster's instability. Fielddata is the bane of my existence and it's the most frequent cause of the highest severity issues that I handle with our customers. Understanding Fielddata The inverted index is the magic that makes Elasticsearch queries so fast. This data structure holds a sorted list of all the unique terms that appear in a field, and each term points to the list of documents that contain that term: Term: Docs: 1 2 3 4 5 ---------------------------- brown X X X fox X X quick X X ---------------------------- Search asks the question: What documents contain termbrownin thefoofield? The inverted index is the perfect structure to answer this question: look up the term of interest in the sorted list and you immediately know which documents match your query. Sorting or aggregations, however, need to be able to answer this question: What terms does Document 1 contain in thefoofield? To answer this, we need a data structure that is the opposite of the inverted index: Docs: Terms: ---------------------------- 1 [ brown ] 2 [ quick ] 3 [ brown ] 4 [ brown, fox, quick ] 5 [ fox ] ---------------------------- This is the purpose of fielddata.Fielddata can be generated at query time by reading the inverted index, inverting the term <-> doc data structure, and storing the results in memory.(就是doc->terms的正向索引啊,不过它是在查询阶段通过读取倒排索引得到的?如果真是这样,那么如何能够比doc values更快?)The two major downsides of this approach should be obvious: Loading fielddata can be slow, especially with big segments. It consumes a lot of valuable heap space. Because loading fielddata is costly, we try to do it as seldom as possible. Once loaded, we keep it in memory for as long as possible. By default, fielddata is loaded on demand, which means that you will not see it until you are using it. Also, by being loaded per segment, it means that new segments that get created will slowly add to your overall memory usage until the field's fielddata is evicted from memory. Eviction happens in only a few ways: Deleting the index or indices that contains it. Closing the index or indices that contains it. Segment fielddata is removed when segments are removed (e.g., background merging). This usually just means that the problem is moving rather than going away. Restarting the node containing the fielddata. Clearing the relevant fielddata cache. Automatically evicting the fielddata to make room for other fielddata. This will not happen with default settings. While the first two ways will cause the memory to be evicted, they're not useful in terms of solving the problem because they make the index unusable. Segment merging is happening in the background and it is not a way to clear fielddata. The fourth and fifth ways are unlikely to be a long term solution because they do not prevent fielddata from being reloaded. The sixth option, evicting fielddata when the cache is full, leads to different issues: one request triggers fielddata loading for one field and the next request triggers loading for another, causing the first field to be evicted. This causes memory thrashing and slow garbage collections, and your users suffer from very slow queries while they wait for their fielddata to be loaded. Simply put, once fielddata becomes a problem, then it stays a problem. Why Fielddata is Bad At small scales, you can generally get away with fielddata usage without even realizing that you are using it. In highly controlled environments, you may even enjoy that specific fields are being loaded into memory for theoretically faster access. However, almost without fail, you are bound to run into a problem with it eventually. Whether it's because someone ran a test request on the production system without thinking that it would be a problem (it's just one query, right?), your queries changed to match new data, or you just finally reached a scale where it no longer works: you will eventually run into memory pressure that does not go away. As I noted earlier, fielddata does not go away on its own. In Elasticsearch 1.3 and later, we allow up to 60% of your Java heap's memory to be consumed by fielddata per node. We control this via theFielddata Circuit Breaker, which checks incoming requests for potential fielddata usage and thenblocksthem if they require more memory than is currently available. Any circuit breaker's purpose is to prevent any bad requests, which means that it never gets the chance to cause a problem (e.g., allocate even more fielddata), but it's important to note that it will not clear any existing fielddata. For example, if a node has 10 GBs of Java heap, then 60% of that is going to be 6 GBs. If a new request requires 1 GB of fielddata to be loaded for that node that is already using 4 GBs of the heap for fielddata, then it will allow it because 4 GB, plus 1 GB, is less than 6 GB. However, if the next request needed 2 GB for yet another field's fielddata, then the entire request would be rejected because the fielddata is exhausted (5 GB + 2 GB = 7 GB, which is clearly greater than 6 GB). Note: for versions prior to Elasticsearch 1.3, we allowed an unlimited amount of your Java heap to be consumed by fielddata. Finding Your Fielddata Fortunately, it's not all bad news. Not only do we have a solution to the problem, but we also provide a way to find and understandyourproblem with it. $ curl -XGET 'localhost:9200/_cat/fielddata?v&fields=*' This will provide a list ofeach node with its fielddata usage.For instance, at startup, my local node is using absolutely no fielddata: id host ip node total iExRFXn1Qw23iRzhwor-Wg Chriss-MBP.home 192.168.1.2 WallE 0b To see it change, it's as simple as sorting, scripting, or aggregating any field. So let's do all three! $ curl -XGET localhost:9200/test-index/_search -d '{ "query": { "filtered": { "filter": { "script": { "script": "doc[\"percentage\"].value > 0.5" } } } }, "aggs": { "max_number": { "max": { "field": "number" } } }, "sort": [ { "@timestamp": { "order": "desc" } } ] }' Although order is irrelevant for this, the first field that will be impacted will be thepercentagefield that is accessed inside of the scripted filter. The second field used will be thenumberfield from the aggregation. Finally, the last field is the@timestampfield used to sort the filtered results. Taking another look at the_cat/fielddatacommand above confirms this: id host ip node total number @timestamp percentage iExRFXn1Qw23iRzhwor-Wg Chriss-MBP.home 192.168.1.2 WallE 49.9kb 24.8kb 24.8kb 192b Use Doc Values The solution to this fielddata problem is to avoid it altogether. Fortunately, you canavoid the use of fielddata bymanuallymapping all of your fields to usedoc values.Without repeating too much from the guide, doc values offload this burden by writing the fielddata to disk at index time, thereby allowing Elasticsearch to load the values outside of your Java heap as they are needed. By taking the burden out of your heap, you get fast access to the on-disk fielddata through the file system cache, which gives in-memory performance without the cost of garbage collections coming into play. This also frees up a lot of headroom for the Elasticsearch heap so that more operations (e.g., bulk indexing and concurrent searches) can use the heap without placing the node under memory pressure, which leads to garbage collection that will slow it down. 本文转自张昺华-sky博客园博客,原文链接:http://www.cnblogs.com/bonelee/p/6401686.html,如需转载请自行联系原作者

优秀的个人博客,低调大师

ES shrink ——一般是结合rollover一起使用的,一开始没有看懂官方shrink文档,当看了这个之后就明白了

rollover Elasticsearch 从 5.0 开始,为日志场景的用户提供了一个很不错的接口,叫 rollover。其作用是:当某个别名指向的实际索引过大的时候,自动将别名指向下一个实际索引。 因为这个接口是操作的别名,所以我们依然需要首先自己创建一个开始滚动的起始索引: # curl -XPUT 'http://localhost:9200/logstash-2016.11.25-1' -d '{ "aliases":{ "logstash":{} } }' 然后就可以尝试发起 rollover 请求了: # curl -XPOST 'http://localhost:9200/logstash/_rollover' -d '{ "conditions":{ "max_age":"1d", "max_docs":10000000 } }' 上面的定义意思就是:当索引超过 1 天,或者索引内的数据量超过一千万条的时候,自动创建并指向下一个索引。 这时候有几种可能性: 条件都没满足,直接返回一个 false,索引和别名都不发生实际变化; { "old_index":"logstash-2016.11.25-1", "new_index":"logstash-2016.11.25-1", "rolled_over":false, "dry_run":false, "acknowledged":false, "shards_acknowledged":false, "conditions":{ "[max_docs: 10000000]":false, "[max_age: 1d]":false } } 还没满一天,满了一千万条,那么下一个索引名会是:logstash-2016.11.25-000002; 还没满一千万条,满了一天,那么下一个索引名会是:logstash-2016.11.26-000002。 shrink Elasticsearch 一直以来都是固定分片数的。这个策略极大的简化了分布式系统的复杂度,但是在一些场景,比如存储 metric 的 TSDB、小数据量的日志存储,人们会期望在多分片快速写入数据以后,把老数据合并存储,节约过多的 cluster state 容量。从 5.0 版本开始,Elasticsearch 新提供了 shrink 接口,可以成倍数的合并分片数。 注:所谓成倍数的,就是原来有 15 个分片,可以合并缩减成 5 个或者 3 个或者 1 个分片。 整个合并缩减的操作流程,大概如下: 先把所有主分片都转移到一台主机上; 在这台主机上创建一个新索引,分片数较小,其他设置和原索引一致; 把原索引的所有分片,复制(或硬链接)到新索引的目录下; 对新索引进行打开操作恢复分片数据。 (可选)重新把新索引的分片均衡到其他节点上。 准备工作 因为这个操作流程需要把所有分片都转移到一台主机上,所以作为 shrink 主机,它的磁盘要足够大,至少要能放得下一整个索引。 最好是一整块磁盘,因为硬链接是不能跨磁盘的。靠复制太慢了。 开始迁移: # curl -XPUT 'http://localhost:9200/metric-2016.11.25/_settings' -d ' { "settings":{ "index.routing.allocation.require._name":"shrink_node_name", "index.blocks.write":true } }' shrink 操作 curl-XPOST'http://localhost:9200/metric-2016.11.25/_shrink/oldmetric-2016.11.25'-d' { "settings": { "index.number_of_replicas": 1, "index.number_of_shards": 3 }, "aliases": { "metric-tsdb": {} } }' 这个命令执行完会立刻返回,但是 Elasticsearch 会一直等到 shrink 操作完成的时候,才会真的开始做 replica 分片的分配和重均衡,此前分片都处于 initializing 状态。 注意:Elasticsearch 有一个硬编码限制,单个分片内的文档总数不得超过 2147483519 个。一般来说这个限制在日志场景下是不太会触发的,但是如果做 TSDB 用,则需要多加注意! 本文转自张昺华-sky博客园博客,原文链接:http://www.cnblogs.com/bonelee/p/8136708.html,如需转载请自行联系原作者

优秀的个人博客,低调大师

ES 断路器——本质上保护OOM提前抛出异常而已监控fielddata使用了多少内存以及是否有数据被驱逐是非常重要的。大量的数据被驱逐会导致...

监控fielddata使用了多少内存以及是否有数据被驱逐是非常重要的。大量的数据被驱逐会导致严重的资源问题以及不好的性能。 Fielddata使用可以通过下面的方式来监控: 对于单个索引使用 {ref}indices-stats.html[indices-statsAPI]: GET /_stats/fielddata?fields=* 对于单个节点使用 {ref}cluster-nodes-stats.html[nodes-statsAPI]: GET /_nodes/stats/indices/fielddata?fields=* 或者甚至单个节点单个索引 GET /_nodes/stats/indices/fielddata?level=indices&fields=* 通过设置?fields=*内存使用按照每个字段分解了. 断路器(breaker) 聪明的读者可能已经注意到fielddata大小设置的一个问题。fielddata的大小是在数据被加载之后才校验的。如果一个查询尝试加载到fielddata的数据比可用的内存大会发生什么情况?答案是不客观的:你将会获得一个OutOfMemory异常。 Elasticsearch包含了一个fielddata断路器,这个就是设计来处理这种情况的。断路器通过检查涉及的字段(它们的类型,基数,大小等等)来估计查询需要的内存。然后检查加 载需要的fielddata会不会导致总的fielddata大小超过设置的堆的百分比。 如果估计的查询大小超过限制,断路器就会触发并且查询会被抛弃返回一个异常。这个发生在数据被加载之前,这就意味着你不会遇到OutOfMemory异常。 Elasticsearch拥有一系列的断路器,所有的这些都是用来保证内存限制不会被突破: indices.breaker.fielddata.limit 这个fielddata断路器限制fielddata的大小为堆大小的60%,默认情况下。 indices.breaker.request.limit 这个request断路器估算完成查询的其他部分要求的结构的大小,比如创建一个聚集通, 以及限制它们到堆大小的40%,默认情况下。 indices.breaker.total.limit 这个total断路器封装了request和fielddata断路器去确保默认情况下这2个 使用的总内存不超过堆大小的70%。 断路器限制可以通过文件config/elasticsearch.yml指定,也可以在集群上动态更新: PUT /_cluster/settings { "persistent" : { "indices.breaker.fielddata.limit" : 40% (1) } } 这个限制设置的是堆的百分比。 最好把断路器设置成一个相对保守的值。记住fielddata需要和堆共享request断路器, 索引内存缓冲区,过滤器缓存,打开的索引的Lucene数据结构,以及各种各样别的临时数据 结构。所以默认为相对保守的60%。过分乐观的设置可能会导致潜在的OOM异常,从而导致整 个节点挂掉。 从另一方面来说,一个过分保守的值将会简单的返回一个查询异常,这个异常会被应用处理。 异常总比挂掉好。这些异常也会促使你重新评估你的查询:为什么单个的查询需要超过60%的 堆空间。 断路器和Fielddata大小 在Fielddata大小部分我们谈到了要给fielddata大小增加一个限制去保证老的不使用 的fielddata被驱逐出去。indices.fielddata.cache.size和indices.breaker.fielddata.limit的关系是非常重要的。如果断路器限制比缓冲区大小要小,就会没有数据会被驱逐。为了能够 让它正确的工作,断路器限制必须比缓冲区大小要大。 我们注意到断路器是和总共的堆大小对比查询大小,而不是和真正已经使用的堆内存区比较。 这样做是有一系列技术原因的(比如,堆可能看起来是满的,但是实际上可能正在等待垃圾 回收,这个很难准确的估算)。但是作为终端用户,这意味着设置必须是保守的,因为它是 和整个堆大小比较,而不是空闲的堆比较。 参考:Elasticsearch权威指南笔记 官网:https://www.elastic.co/guide/en/elasticsearch/guide/current/_limiting_memory_usage.html 本文转自张昺华-sky博客园博客,原文链接: http://www.cnblogs.com/bonelee/p/8202878.html ,如需转载请自行联系原作者

资源下载

更多资源
腾讯云软件源

腾讯云软件源

为解决软件依赖安装时官方源访问速度慢的问题,腾讯云为一些软件搭建了缓存服务。您可以通过使用腾讯云软件源站来提升依赖包的安装速度。为了方便用户自由搭建服务架构,目前腾讯云软件源站支持公网访问和内网访问。

Spring

Spring

Spring框架(Spring Framework)是由Rod Johnson于2002年提出的开源Java企业级应用框架,旨在通过使用JavaBean替代传统EJB实现方式降低企业级编程开发的复杂性。该框架基于简单性、可测试性和松耦合性设计理念,提供核心容器、应用上下文、数据访问集成等模块,支持整合Hibernate、Struts等第三方框架,其适用范围不仅限于服务器端开发,绝大多数Java应用均可从中受益。

Sublime Text

Sublime Text

Sublime Text具有漂亮的用户界面和强大的功能,例如代码缩略图,Python的插件,代码段等。还可自定义键绑定,菜单和工具栏。Sublime Text 的主要功能包括:拼写检查,书签,完整的 Python API , Goto 功能,即时项目切换,多选择,多窗口等等。Sublime Text 是一个跨平台的编辑器,同时支持Windows、Linux、Mac OS X等操作系统。

WebStorm

WebStorm

WebStorm 是jetbrains公司旗下一款JavaScript 开发工具。目前已经被广大中国JS开发者誉为“Web前端开发神器”、“最强大的HTML5编辑器”、“最智能的JavaScript IDE”等。与IntelliJ IDEA同源,继承了IntelliJ IDEA强大的JS部分的功能。

用户登录
用户注册