
Go panic 与 recover从内核 stack unwind 到生产级容错模式每个 Go 工程师都怕线上 panic。掌握 panic 的本质、recover 的边界才能从容应对突发故障。一、panic 不是 exceptionpanic 是 Go 运行时控制流的硬中断。它会沿调用栈逐层 unwind直到遇到 recover 才停止。funcmain(){deferfunc(){ifr:recover();r!nil{fmt.Println(recovered:,r)}}()panic(boom)}输出recovered: boom进程继续运行。二、recover 的三大限制必须在 defer 函数里直接调用放进 subFunc 再调无效只能捕获当前 goroutine子协程 panic 不影响父协程反之亦然重复 recover 无副作用但 panic 后世界已变质DB 连接、文件描述符都可能已损坏三、生产级 panic 守护3.1 HTTP 服务的链路守护funcRecovery(next http.Handler)http.Handler{returnhttp.HandlerFunc(func(w http.ResponseWriter,r*http.Request){deferfunc(){ifrec:recover();rec!nil{log.Printf(panic on %s %s: %v\n%s,r.Method,r.URL.Path,rec,debug.Stack())w.WriteHeader(http.StatusInternalServerError)w.Write([]byte({code:500,msg:internal error}))}}()next.ServeHTTP(w,r)})}3.2 gRPC 拦截器funcUnaryRecovery()grpc.UnaryServerInterceptor{returnfunc(ctx context.Context,reqinterface{},info*grpc.UnaryServerInfo,handler grpc.UnaryHandler)(respinterface{},errerror){deferfunc(){ifr:recover();r!nil{errstatus.Errorf(codes.Internal,panic: %v,r)}}()returnhandler(ctx,req)}}3.3 Worker pool 容错forjob:rangech{func(){deferfunc(){ifr:recover();r!nil{metrics.IncPanic()}}()process(job)}()}每个 job 用独立函数包裹避免单个任务挂掉整条 worker。四、panic 传播与runtime.Goexitruntime.Goexit()让 goroutine 优雅退出。它会触发 defer但不会触发 panic。如果你在 goroutine 里手动调用 Goexitdefer 仍按顺序跑。如果想区分这两种情况可以这样deferfunc(){ifr:recover();r!nil{// 来自 panic}// 无法直接判断 Goexit要用 ctx.Err()}()五、调试技巧快速定位栈runtime/debug.Stack()能拿到整个 goroutine 的栈快照。建议写到日志里deferfunc(){ifr:recover();r!nil{log.Printf(panic: %v\n%s,r,debug.Stack())sentry.CaptureException(fmt.Errorf(%v,r))}}()能解决 90% 线上排障问题。六、避坑清单不要用 panic 写业务流error 才是正道recover 只兜底层在通用入口HTTP/gRPC兜业务函数不要写 recoverpanic 后慎用全局 state状态可能脏掉不要依赖 panic 值类型调试期 stringify 即可七、总结与展望panic/recover 是一把双刃剑。底层 goroutine 出错时它们守住最后一道防线业务层又应视为错误源而非捷径。未来 Go 1.24 可能引入 structured panic带 trace 信息便于 APM 直接捕获。八、参考文献Go 标准库 runtime/debugThe Go Blog: Defer, Panic, and RecoverSentinel / slog 文档