ARTICLE DETAIL

资讯详情

深耕郑州网站建设与运营推广的一线实战洞察。

Spring Boot 实现高并发分布式 ID 生成:雪花算法原理、时钟回拨处理与性能优化实战

Spring Boot 实现高并发分布式 ID 生成:雪花算法原理、时钟回拨处理与性能优化实战 最近做订单系统ID 生成这事儿看着简单真要整好坑不少。尤其是高并发场景一个破 UUID 都能把数据库索引干趴下。所以打算把雪花算法从原理到落地捋一遍顺便聊聊 Spring Boot 里怎么封装、怎么处理时钟回拨、怎么压榨性能。这篇文章是我实际写代码的经验总结不是那种教科书式的贴代码想到哪儿写到哪儿但技术点都是真刀真枪的。1. 业务场景与方案选择需要全局唯一 ID 的地方太多了订单号、消息 ID、链路追踪的 TraceId还有日志 ID、支付流水号。这些场景要求高并发、全局唯一、最好趋势递增而且生成过程不能成为瓶颈。常见的生成方案UUID最大的问题是无序、太长。字符串乱序数据库索引性能差排序也难受。优点是本地生成没网络开销。数据库自增 ID数字有序但依赖数据库分布式环境下有单点问题并发一高就是瓶颈。号段模式每次从数据库拿一批号性能不错但依赖数据库重启可能丢号段。雪花算法不依赖外部组件趋势递增性能极高缺点也明显——时钟回拨会产生重复 ID还需要维护 workerId。综合来看雪花算法是主流但得针对时钟回拨做增强。下面从原理开始。2. 雪花算法结构解析雪花算法最早是 Twitter 搞出来的核心思想很简洁把一个 64 位整数按位拆开分别表示时间戳、机器 ID 和序列号。0 1 - 41 42 - 51 52 - 63 ┌─┬─────────────┬─────────────┬─────────────┐ │0│ 时间戳(ms) │ 机器ID(10bit)│ 序列号(12bit)│ └─┴─────────────┴─────────────┴─────────────┘符号位永远是 0保证 ID 是正数。时间戳41 位毫秒时间戳。2^41 毫秒差不多 69 年实际用的时候要设一个起始时间比如从 2024-01-01 开始这样可以用到 2093 年。机器 ID10 位可以拆成 5 位机房 5 位机器支持 32 个机房每个机房 32 台机器总共 1024 个节点。序列号12 位同一毫秒内最多生成 4096 个 ID。序列号用完了就等下一毫秒。这个设计的好处很明显时间戳 机器 ID 序列号组合起来只要时钟不回拨、机器 ID 不冲突全局唯一整体趋势递增因为毫秒递增同一毫秒内序列号递增纯内存运算没网络开销。3. Spring Boot 实现一个雪花生成器先建一个基础版把位运算核心逻辑写出来。publicclassSnowflakeIdGenerator{// 起始时间戳2024-01-01 00:00:00privatestaticfinallongSTART_TIMESTAMP1704067200000L;// 各部分位数privatestaticfinallongSEQUENCE_BITS12L;privatestaticfinallongWORKER_ID_BITS10L;// 最大值privatestaticfinallongSEQUENCE_MASK~(-1LSEQUENCE_BITS);privatestaticfinallongWORKER_ID_MASK~(-1LWORKER_ID_BITS);// 位移量privatestaticfinallongWORKER_ID_SHIFTSEQUENCE_BITS;privatestaticfinallongTIMESTAMP_SHIFTSEQUENCE_BITSWORKER_ID_BITS;privatefinallongworkerId;privatelonglastTimestamp-1L;privatelongsequence0L;publicSnowflakeIdGenerator(longworkerId){if(workerId0||workerIdWORKER_ID_MASK){thrownewIllegalArgumentException(workerId must be between 0 and WORKER_ID_MASK);}this.workerIdworkerId;}publicsynchronizedlongnextId(){longcurrentTimestampSystem.currentTimeMillis();if(currentTimestamplastTimestamp){// 先留空后面处理时钟回拨thrownewIllegalStateException(Clock moved backwards);}if(currentTimestamplastTimestamp){sequence(sequence1)SEQUENCE_MASK;if(sequence0){currentTimestampwaitNextMillis(currentTimestamp);}}else{sequence0L;}lastTimestampcurrentTimestamp;return((currentTimestamp-START_TIMESTAMP)TIMESTAMP_SHIFT)|(workerIdWORKER_ID_SHIFT)|sequence;}privatelongwaitNextMillis(longcurrentTimestamp){while(currentTimestamplastTimestamp){currentTimestampSystem.currentTimeMillis();}returncurrentTimestamp;}}配置里加数据中心和机器 ID。我习惯把 10 位拆成 5 位机房 5 位机器最终workerId (dataCenterId 5) | machineId。ComponentConfigurationProperties(prefixsnowflake)publicclassSnowflakeProperties{privatelongdataCenterId1;privatelongmachineId1;// getter/setter 省略publiclonggetWorkerId(){return(dataCenterId5)|machineId;}}配置类里生成 BeanConfigurationpublicclassIdGeneratorConfig{BeanpublicSnowflakeIdGeneratorsnowflakeIdGenerator(SnowflakePropertiesproperties){returnnewSnowflakeIdGenerator(properties.getWorkerId());}}这样就能直接Autowired使用了。但还没处理时钟回拨接下来是关键。4. 时钟回拨问题处理雪花算法最怕系统时间往回跳一回到过去同一毫秒内就可能生成重复 ID。业界常见的招数有三招等待、抛错、历史时间补偿。4.1 等待如果回拨时间不长比如 5 秒内就让线程睡一会儿等系统时间追上来。if(currentTimestamplastTimestamp){longoffsetlastTimestamp-currentTimestamp;if(offset5000){try{Thread.sleep(offset*2);}catch(InterruptedExceptione){Thread.currentThread().interrupt();}currentTimestampSystem.currentTimeMillis();while(currentTimestamplastTimestamp){currentTimestampSystem.currentTimeMillis();}}else{thrownewIllegalStateException(Clock moved backwards too much);}}思路很直白但回拨频繁的话线程会被反复阻塞请求堆积。4.2 直接抛错回拨超过阈值直接抛异常让上层决定是降级还是拒绝。if(currentTimestamplastTimestamp){longoffsetlastTimestamp-currentTimestamp;if(offset100){thrownewIllegalStateException(Clock moved backwards. Refusing to generate ID for offset ms);}// 小回拨继续处理}这种策略容易导致服务不可用但能保证不生成重复 ID。4.3 历史时间补偿这个思路比较巧妙。既然当前时钟回拨了那就假装时间还在上次生成 ID 的时候序列号继续往上加直到加满 4096 个为止。这样避免等待也不抛错。publicsynchronizedlongnextId(){longcurrentTimestampSystem.currentTimeMillis();if(currentTimestamplastTimestamp){sequence(sequence1)SEQUENCE_MASK;if(sequence0){thrownewIllegalStateException(Sequence exhausted while clock moved backwards);}// 继续用 lastTimestamp序列号递增return((lastTimestamp-START_TIMESTAMP)TIMESTAMP_SHIFT)|(workerIdWORKER_ID_SHIFT)|sequence;}// 正常流程...}但注意这个方案只能补偿短时间回拨。如果回拨太久序列号用尽就只能抛错了。所以生产上我更喜欢“等待 补偿”的组合publicsynchronizedlongnextId(){longcurrentTimestampSystem.currentTimeMillis();longoffsetlastTimestamp-currentTimestamp;if(offset0){if(offset5000){thrownewIllegalStateException(Clock moved backwards, offsetoffsetms);}try{Thread.sleep(offset1);}catch(InterruptedExceptione){Thread.currentThread().interrupt();}currentTimestampSystem.currentTimeMillis();}if(currentTimestamplastTimestamp){sequence(sequence1)SEQUENCE_MASK;if(sequence0){currentTimestampwaitNextMillis(currentTimestamp);}}else{sequence0L;}lastTimestampcurrentTimestamp;return((currentTimestamp-START_TIMESTAMP)TIMESTAMP_SHIFT)|(workerIdWORKER_ID_SHIFT)|sequence;}另外如果你们有 Redis 或者 ZooKeeper可以把最近生成 ID 的时间戳存一份本机时间小于全局最大时间戳时直接拿全局时间戳用。不过这会引入网络开销我一般不用。5. 预生成 ID 池与双缓冲雪花算法本身延迟已经很低了但System.currentTimeMillis()是有系统调用开销的再加上synchronized锁高并发下还是会被拖死。有一个土办法先把 ID 批量生成好放在本地队列里用的时候直接从队列取一步到位。publicclassIdPool{privatefinalSnowflakeIdGeneratorgenerator;privatefinalBlockingQueueLongqueuenewLinkedBlockingQueue(10000);privatefinalExecutorServiceexecutorExecutors.newSingleThreadExecutor();publicIdPool(SnowflakeIdGeneratorgenerator){this.generatorgenerator;fill();}privatevoidfill(){for(inti0;i10000;i){queue.offer(generator.nextId());}}publicLongnextId()throwsInterruptedException{if(queue.size()5000){executor.submit(this::fill);// 异步补充注意防重}returnqueue.take();}}生产环境别这么裸至少加个AtomicBoolean防止重复提交填充任务。也可以用双缓冲两个队列一个消费一个后台填充消费到一半就交换。publicclassDoubleBufferIdPool{privatefinalSnowflakeIdGeneratorgenerator;privatefinalintbufferSize;privatevolatileQueueLongcurrentBuffer;privateQueueLongnextBuffer;privatefinalExecutorServiceexecutorExecutors.newSingleThreadExecutor();publicDoubleBufferIdPool(SnowflakeIdGeneratorgenerator,intbufferSize){this.generatorgenerator;this.bufferSizebufferSize;this.currentBuffernewArrayDeque();this.nextBuffernewArrayDeque();fillBuffer(currentBuffer);}privatevoidfillBuffer(QueueLongbuffer){for(inti0;ibufferSize;i){buffer.offer(generator.nextId());}}publicLongnextId(){LongidcurrentBuffer.poll();if(idnull){synchronized(this){if(currentBuffer.isEmpty()){QueueLongtmpcurrentBuffer;currentBuffernextBuffer;nextBuffertmp;fillBuffer(nextBuffer);}idcurrentBuffer.poll();}}else{if(currentBuffer.size()bufferSize/2nextBuffer.size()bufferSize){executor.submit(()-fillBuffer(nextBuffer));}}returnid;}}预生成会浪费一些 ID但相比性能提升这点浪费值得。另外批量接口可以一次返回多个 ID进一步减少调用次数。6. 集成 MyBatis-PlusMyBatis-Plus 默认的 ASSIGN_ID 就是雪花算法但我们可以换成自己的生成器做到统一管理。只需要实现IdentifierGenerator接口。ComponentpublicclassMyBatisPlusIdGeneratorimplementsIdentifierGenerator{AutowiredprivateIdPoolidPool;// 或者直接用 SnowflakeIdGeneratorOverridepublicNumbernextId(Objectentity){returnidPool.nextId();}}实体类主键用TableId(type IdType.ASSIGN_ID)就行。如果想生成字符串订单号可以单独写个服务ServicepublicclassOrderIdService{AutowiredprivateSnowflakeIdGeneratorgenerator;publicStringgenerateOrderId(){returnORDERnewSimpleDateFormat(yyyyMMdd).format(newDate())generator.nextId();}}7. 美团 Leaf 的设计思路美团 Leaf 是业界很经典的方案两种模式都值得学习。号段模式从数据库拿一个号段比如 1~1000进程内按顺序分配。数据库表里有biz_tag、max_id、step这些字段。当号段快用完时后台异步去数据库加载下一个号段。优点是对数据库压力小ID 趋势递增缺点是要维护号段表数据库挂了也玩不转。雪花模式Leaf 做雪花模式的时候解决了两个痛点。一个是 workerId 动态分配通过 ZooKeeper 注册节点启动时获取递增序号作为 workerId关闭时释放。另一个是时钟回拨Leaf 有checkTimestamp机制检测到回拨就抛异常然后告警人工处理。参考 Leaf我们在生产环境可以考虑用 Redis 动态分配 workerId但大多数团队规模用配置文件就行只要不搞混。8. 压测与监控我拿 8C16G 的机器压了一次JMeter 100 线程60 秒基础雪花算法 QPS 大概 3.5 万平均响应 2.8msP99 8.1ms。加了 ID 池双缓冲之后QPS 能到 12 万平均响应降到 0.8msP99 也只有 2.3ms。差距很明显。生产环境建议重点监控这几个指标生成耗时平均耗时、TP99。池水位剩余 ID 数量低于阈值就告警。时钟回拨次数记下每次回拨的时间差和频率。QPS评估容量和扩缩容。用 Micrometer Prometheus Grafana 就能搞定代码里埋点记录一下Timer和Counter就行。9. 总结雪花算法不是什么银弹但掌握了原理和变体大部分场景都能应对。我的建议是起始时间戳设得远一点保证 ID 用几十年。workerId 别重复运维要统一登记。开启 NTP 同步加-x参数避免时间跳变。日志里记下时钟回拨详情出了问题好排查。ID 生成服务要做限流降级别让上游突起压垮。差不多就是这些有问题评论区聊。官网www.farerboy.com
返回列表