
Elasticsearch 实战指南Elasticsearch 是搜索、日志分析领域的事实标准也是 CSDN 后端与大数据板块的高流量话题。本文从核心概念讲起覆盖索引与文档、分词器、查询 DSL、聚合分析、Spring Boot 集成、日志平台ELK、性能优化与常见坑全程带命令和代码。一、Elasticsearch 是什么1.1 定位基于 Lucene 的分布式搜索引擎——本质是一个存文档 全文检索 聚合分析的 NoSQL 数据库。和 MySQL 的区别维度MySQLElasticsearch查询精确匹配、关系型全文检索、模糊搜索性能千万级还行亿级毫秒响应扩展分库分表麻烦天然分布式一致性强一致事务近实时秒级可见典型场景业务数据存储搜索、日志、分析1.2 典型场景站内搜索电商商品搜索、文档搜索、论坛日志分析ELKElasticsearch Logstash Kibana指标监控APM、业务指标推荐/补全搜索建议、纠错。1.3 核心概念Elasticsearch MySQL 的实例/集群 Index索引 数据库database Type类型已废弃 表7.x 移除 Document文档 一行记录JSON Field字段 一列 Mapping映射 表结构字段类型定义一条文档示例{_index:products,_id:1001,name:iPhone 16 Pro,price:7999.00,tags:[手机,苹果],created_at:2026-09-30T10:00:00}二、安装与快速上手2.1 安装# Docker 最快dockerrun-d--namees\-p9200:9200-p9300:9300\-ediscovery.typesingle-node\-eES_JAVA_OPTS-Xms512m -Xmx512m\docker.elastic.co/elasticsearch/elasticsearch:8.13.0验证curlhttp://localhost:92002.2 索引操作# 创建索引PUT /products{settings:{number_of_shards:3,number_of_replicas:1}}# 查看索引GET /products# 删除索引慎用DELETE /products2.3 文档 CRUD# 新增/覆盖文档PUT /products/_doc/1001{name:iPhone 16 Pro,price:7999,tags:[手机,苹果]}# 查询GET /products/_doc/1001# 更新局部POST /products/_update/1001{doc:{price:7499}}# 删除DELETE /products/_doc/1001三、Mapping字段类型设计3.1 核心类型类型用途text全文检索分词、模糊匹配keyword精确匹配等值、排序、聚合integer / long整数float / double浮点boolean布尔date日期object / nested嵌套对象3.2 为什么要设计 MappingES 会自动推断类型但默认把字符串设为text keyword——text 不能排序和聚合keyword 不能全文检索。生产必须显式定义PUT/products{mappings:{properties:{name:{type:text,fields:{keyword:{type:keyword}// name.keyword 可精确匹配/排序}},price:{type:double},tags:{type:keyword},description:{type:text},created_at:{type:date,format:yyyy-MM-dd HH:mm:ss}}}}关键点加字段容易改字段类型要重建索引reindex线上先想清楚字段类型别依赖自动映射。3.3 重建索引reindex# 建新索引新 mappingPUT /products_v2{...}# 数据搬迁POST /_reindex{source:{index:products},dest:{index:products_v2}}# 改别名POST /_aliases{actions:[{remove:{index:products,alias:products_alias}},{add:{index:products_v2,alias:products_alias}}]}生产规范业务永远用别名访问重建索引无感切换。四、分词器Analyzer4.1 为什么需要分词搜索苹果手机如果文档是iPhone不分词就匹配不上。分词器把文本拆成词元token苹果手机很好用 → [苹果, 手机, 很好, 用]4.2 中文分词IKES 自带 standard 分词器对中文是逐字拆效果差——中文场景必须装IK 分词器./bin/elasticsearch-plugininstallanalysis-ik测试POST /_analyze{analyzer:ik_max_word,text:中华人民共和国成立了}# 输出: [中华人民共和国, 中华人民, 中华, 华人, 人民共和国, 人民, 共和国, 共和, 国, 成立, 了]IK 两种模式ik_max_word最细粒度索引用召回全ik_smart最粗粒度搜索用精确。4.3 自定义分词{settings:{analysis:{analyzer:{my_analyzer:{type:custom,tokenizer:ik_max_word,filter:[lowercase]}}}}}五、查询 DSL重点5.1 基本查询结构GET/products/_search{query:{...},// 查询条件from:0,size:10,// 分页sort:[{price:desc}],// 排序_source:[name,price]// 只返回指定字段}5.2 match全文检索{query:{match:{name:苹果手机}}}match 与 term 的区别必考match先分词再匹配适合 text 字段term精确匹配不分词适合 keyword 字段{query:{term:{tags:苹果}}}坑term查 text 字段基本查不到text 被分词了。5.3 组合查询 bool{query:{bool:{must:[{match:{name:手机}}],// ANDfilter:[{term:{tags:苹果}}],// 过滤不算分快should:[{match:{description:拍照}}],// OR加分must_not:[{term:{status:下架}}]// 排除}}}filter vs mustfilter 不参与相关度评分可缓存性能更好——纯过滤条件用 filter。5.4 范围与多值{query:{range:{price:{gte:1000,lte:8000}}}}{query:{terms:{tags:[苹果,华为]}}}5.5 分页与深分页{from:0,size:20}深分页问题同 MySQLfromsize超过 1 万条会被拒绝默认 max_result_window10000深度翻页用 search_after游标POST/products/_search{size:20,sort:[{price:asc},{_id:asc}],search_after:[4999,1001]// 上一页最后一条的值}5.6 高亮{query:{match:{name:苹果}},highlight:{fields:{name:{}},pre_tags:[em],post_tags:[/em]}}六、聚合分析Aggregation6.1 三种聚合桶bucket分组类似 GROUP BY指标metric统计类似 MAX/AVG/SUM管道pipeline聚合结果再聚合。6.2 实战电商统计POST/orders/_search{size:0,// 只要聚合结果aggs:{by_category:{// 按分类分组terms:{field:category,size:10},aggs:{avg_price:{avg:{field:price}},// 每组均价total_sales:{sum:{field:amount}},// 每组总额by_status:{// 每组内再按状态分terms:{field:status}}}}}}6.3 日期直方图按时间统计{size:0,aggs:{sales_per_day:{date_histogram:{field:created_at,calendar_interval:day}}}}七、Spring Boot 集成7.1 依赖与配置dependencygroupIdorg.springframework.boot/groupIdartifactIdspring-boot-starter-data-elasticsearch/artifactId/dependencyspring:elasticsearch:uris:http://localhost:92007.2 实体与 RepositoryDocument(indexNameproducts)publicclassProduct{IdprivateStringid;privateStringname;privateDoubleprice;privateListStringtags;// getter/setter}publicinterfaceProductRepositoryextendsElasticsearchRepositoryProduct,String{// 方法名自动生成查询ListProductfindByNameContaining(Stringkeyword);// 复杂查询用 QueryQuery DSL 字符串Query({\bool\: {\must\: [{\match\: {\name\: \?0\}}]}})ListProductsearchByName(Stringkeyword);}7.3 自定义查询RestControllerpublicclassProductSearchController{AutowiredprivateElasticsearchOperationsesOps;publicListProductsearch(Stringkeyword,doubleminPrice){QueryqueryNativeQuery.builder().withQuery(q-q.bool(b-b.must(m-m.match(t-t.field(name).query(keyword))).filter(f-f.range(r-r.field(price).gte(JsonData.of(minPrice)))))).build();returnesOps.search(query,Product.class).stream().map(hit-hit.getContent()).toList();}}八、ELK 日志平台8.1 架构应用日志 → Filebeat采集→ Logstash处理→ Elasticsearch存储索引 ↓ Kibana可视化Filebeat轻量采集器读日志文件发到 ES/LogstashLogstash解析、过滤、转换grok 解析日志格式Kibana搜索、图表、仪表盘。8.2 Logstash 配置示例input{beats{port5044}}filter{grok{match{message%{TIMESTAMP_ISO8601:log_time} %{LOGLEVEL:level} %{GREEDYDATA:msg}}}date{match[log_time,yyyy-MM-dd HH:mm:ss,SSS]}}output{elasticsearch{hosts[http://es:9200]indexapp-logs-%{YYYY.MM.dd}}}8.3 日志索引生命周期日志量大按天建索引配合索引生命周期管理ILM3 天内热节点读写快30 天内温节点压缩之后删除。九、性能优化9.1 写入优化批量写入_bulk一次几千条别一条一条写调大 refresh_interval默认 1s 刷新可设 30s日志场景关闭副本再写导入数据时 replicas0导完恢复。PUT/logs/_settings{index:{refresh_interval:30s}}9.2 查询优化用 filter 不用 must纯过滤可缓存不要深分页用 search_after只返回需要的字段_source 裁剪避免 wildcard 模糊查询性能差大聚合限制 size。9.3 分片设计单分片 30-50GB 为宜分片数 节点数 × 1~2每节点 3-5 个分片合适分片太多资源浪费、查询慢。9.4 内存设置# jvm.options-Xms4g-Xmx4g红线堆不要超过物理内存一半留一半给操作系统页缓存ES 大量用 OS cache。十、常见坑速查现象原因解决中文搜不到没装 IK 分词器装 IK字段用 ik 分词term 查 text 查不到text 被分词用 keyword 字段或 match聚合报错text 字段不能聚合用 .keyword 字段深分页报错超过 max_result_windowsearch_after索引变黄/红副本分配失败/磁盘满看集群健康、加磁盘写入慢单条写入/refresh 频繁批量 调 refresh_interval数据不实时近实时1s 刷新接受或调 refresh字段类型错了依赖自动映射显式 mappingreindex本章小结Elasticsearch 的实战核心Mapping 设计text/keyword、查询 DSLmatch/term/bool/range、聚合分析、分页策略search_after、性能优化。先学会用 curl 玩转 DSL再套框架Spring Data/ELK。下一篇讲设计模式——Java 面试与代码质量的常青话题。全文完