数据节点磁盘使用量达到危险水平时,基于磁盘的分片分配水位设置会触发,以保护节点磁盘。默认百分比阈值、对应行为及日志如下。
85%,低水位:Elasticsearch 停止向受影响节点分配副本分片和主分片,但新建索引的主分片除外。
low disk watermark [85%] exceeded on [NODE_ID][NODE_NAME] free: Xgb[X%], replicas will not be assigned to this node
90%,高水位:Elasticsearch 将分片从受影响节点重新平衡到其他节点。
high disk watermark [90%] exceeded on [NODE_ID][NODE_NAME] free: Xgb[X%], shards will be relocated away from this node
95%,洪泛水位:Elasticsearch 将受影响节点上的全部相关索引设为只读。当该节点磁盘使用量降到高水位以下,写入阻止会自动解除。
flood-stage watermark [95%] exceeded on [NODE_ID][NODE_NAME], all indices on this node will be marked read-only
注意:磁盘使用率达到 75% 时,Elastic Cloud 控制台会为节点显示红色磁盘指示器,提示使用率升高。这只是视觉提示,与 Elasticsearch 水位或磁盘限制行为无关,此时不会施加分片分配或写入限制。
为防止磁盘写满,节点达到洪泛水位后,Elasticsearch 会阻止向任何在该节点上拥有分片的索引写入。如果涉及系统索引,Kibana 或其他 Elastic Stack 功能可能不可用。例如,Kibana 显示 Kibana Server is not Ready yet,摄取 API 返回 HTTP 429,正文类似:
{
"reason": "index [INDEX_NAME] blocked by: [TOO_MANY_REQUESTS/12/disk usage exceeded flood-stage watermark, index has read-only-allow-delete block];",
"type": "cluster_block_exception"
}
常见原因包括:
- 突然摄取大量数据,磁盘消耗超过峰值负载测试预期。参阅索引性能注意事项。
- 索引设置低效、保存不必要字段、文档结构不理想,增加磁盘消耗。参阅磁盘使用优化指南。
- 副本数量过多,每个副本消耗与主分片相同的磁盘空间,存储需求迅速倍增。
- 分片过大,容易出现磁盘使用峰值,也会拖慢恢复与重新平衡。
更多原因和处理方法,见 Elastic 关于集群触发磁盘水位的博客文章。
监控磁盘使用量
AutoOps 可以实时检测集群问题,并建议解决方法与性能改进措施。
要跟踪磁盘使用趋势,请根据部署类型启用监控:
- 对于 Elastic Cloud 托管部署,推荐启用 AutoOps;也可以启用日志和指标,在 Kibana 的 Stack Monitoring 页面查看,并启用磁盘使用阈值告警;还可以在部署菜单的 Performance 页面查看磁盘图表。
- 对于自管理等其他适用部署,同样推荐 AutoOps;或启用 Elasticsearch 监控,通过 Kibana Stack Monitoring 查看信息,并配置磁盘阈值告警。
监控重新平衡
确认分片正在迁出受影响节点,直到磁盘使用量低于高水位,可以使用以下 API。
集群健康 API,查看 relocating_shards:
GET _cluster/health
CAT recovery API,查看恢复中的分片数量,以及已迁移字节百分比 bp 和总字节数 tb:
GET _cat/recovery?v=true&expand_wildcards=all&active_only=true&h=time,tb,bp,top,ty,st,snode,tnode,idx,sh&s=time:desc
如果分片仍停留在该节点,导致磁盘一直高于高水位,使用 CAT shards API 确认节点承载哪些分片:
GET _cat/shards?v=true
使用集群分配解释 API,获取所选分片的分配状态原因:
GET _cluster/allocation/explain
{
"index": "my-index-000001",
"shard": 0,
"primary": false
}
输出解读方式见集群分配 API 排查指南。
通常应等待 Elasticsearch 自行平衡。高级用户如果根据预计摄取速度或当前磁盘使用量,确认某个分片需要更快迁移,可以考虑使用集群 reroute API,立即将选定分片重新平衡到指定目标节点。
临时缓解
要立即恢复写入,可以暂时提高磁盘水位并解除写入阻止:
PUT _cluster/settings
{
"persistent": {
"cluster.routing.allocation.disk.watermark.low": "90%",
"cluster.routing.allocation.disk.watermark.low.max_headroom": "100GB",
"cluster.routing.allocation.disk.watermark.high": "95%",
"cluster.routing.allocation.disk.watermark.high.max_headroom": "20GB",
"cluster.routing.allocation.disk.watermark.flood_stage": "97%",
"cluster.routing.allocation.disk.watermark.flood_stage.max_headroom": "5GB",
"cluster.routing.allocation.disk.watermark.flood_stage.frozen": "97%",
"cluster.routing.allocation.disk.watermark.flood_stage.frozen.max_headroom": "5GB"
}
}
PUT */_settings?expand_wildcards=all
{
"index.blocks.read_only_allow_delete": null
}
长期解决方案到位后,重置或重新配置水位:
PUT _cluster/settings
{
"persistent": {
"cluster.routing.allocation.disk.watermark.low": null,
"cluster.routing.allocation.disk.watermark.low.max_headroom": null,
"cluster.routing.allocation.disk.watermark.high": null,
"cluster.routing.allocation.disk.watermark.high.max_headroom": null,
"cluster.routing.allocation.disk.watermark.flood_stage": null,
"cluster.routing.allocation.disk.watermark.flood_stage.max_headroom": null,
"cluster.routing.allocation.disk.watermark.flood_stage.frozen": null,
"cluster.routing.allocation.disk.watermark.flood_stage.frozen.max_headroom": null
}
}
注意:Elasticsearch 建议使用默认水位。高级用户可以覆盖阈值与预留空间,但可能无法给 force merge 等后台操作留下足够空间,无法匹配数据摄取速率与索引生命周期策略,甚至在磁盘达到 100% 时出现磁盘已满错误。
根本解决
要长期解决水位错误,可以采取以下措施之一:
- 横向扩展受影响数据层的节点数量。
- 纵向扩展现有节点,增加磁盘。数据层中的节点应使用匹配的硬件规格,避免热点。
- 使用删除索引 API 删除索引。不再需要的可永久删除;需要保留的可暂时删除,之后从快照恢复。
提示:Elastic Cloud Hosted 与 Elastic Cloud Enterprise 中,可能需要先通过 Elasticsearch API Console 暂时删除索引,以解除阻止部署变更的红色集群健康状态。问题解决后再从快照恢复。若此流程遇到困难,可联系 Elastic Support。
预防
- 启用自动扩缩容,根据存储与性能需求调整资源。
- 使用更严格的索引生命周期管理策略,让数据更早迁往后续数据层,控制较高数据层的磁盘占用。
- 避免过大和过小索引混杂,以免集群失衡。参阅分片大小规划指南。
原文:Watermark errors。作者/维护方:Elastic 文档维护者。本文为中文翻译,代码及命令保留原文。











暂无评论内容