Go 的并发功能强大而且容易使用,但这种便利有时也会让经验丰富的开发者犯错。好在 Go 生态提供了实用的调试工具,例如数据竞争检测器。不过,现有工具也可能遗漏某些并发错误,本文讨论的 goroutine 泄漏就是其中之一。
Goroutine 通过共享的并发原语同步或交换信息,例如通道、锁和等待组。在通信过程中,goroutine 经常会阻塞在这些原语上,等待某个条件满足。常见情况包括等待获取一把已被持有的互斥锁,或等待从通道接收消息。Goroutine 也可能阻塞在操作系统操作上,例如读取网络套接字或文件。
如果一个 goroutine 已经阻塞,而解除阻塞所需的条件永远不可能满足,就可以认为它发生了泄漏。随着时间推移,泄漏的 goroutine 不断积累,会消耗过多内存,包括 goroutine 自身以及它们引用的内存;垃圾回收也会消耗更多 CPU,使用 GOMEMLIMIT 时尤其如此,最终导致性能下降。
Goroutine 泄漏往往很难检测。单元测试领域的重要进展之一是开源库 goleak:它可以接入单个测试,把测试结束后仍未终止的 goroutine 标记为可疑对象。
类似地,Go 1.25 在标准库中引入了 synctest 包。它让开发者更充分地控制并发事件的先后顺序,从而可靠地测试难以复现的场景,显著改善并发代码的单元测试质量。
但这两种方法都无法直接检查生产系统中的 goroutine 泄漏。大规模生产系统的行为尤其可能超出测试覆盖范围。普通 goroutine 剖析报告可以用于查找使过多 goroutine 阻塞的操作,或分析数量增长趋势;但它无法区分真正泄漏的 goroutine 与按设计暂时大量阻塞的 goroutine,例如微服务流量增长时出现的等待。
同样,数量很少的泄漏也可能多年都未被发现。
Go 1.27 引入了 goroutine 泄漏剖析器。这是一种灵活、轻量的机制,能够查找正在运行的 Go 程序中的 goroutine 泄漏,也适用于生产系统。相比需要人工分析的既有方法,它的检测更精确,几乎不产生误报。代价是检测范围有限:它只覆盖永久阻塞在通道或 sync 包原语上的一部分 goroutine 泄漏。
不过,这个范围已经涵盖了大量常见泄漏,后面的示例会展示这一点。
下面先介绍如何使用这项功能,再给出更多能够检测的泄漏示例,并说明其底层实现与权衡。
示例:并发工作协程
考虑一个并发处理工作项的函数:
type result struct {
res workResult
err error
}
func processWorkItems(ws []workItem) ([]workResult, error) {
// Process work items in parallel, aggregating results in ch.
ch := make(chan result)
for _, w := range ws {
go func() {
res, err := processWorkItem(w)
ch <- result{res, err}
}()
}
// Collect the results from ch, or return an error if one is found.
var results []workResult
for range len(ws) {
r := <-ch
if r.err != nil {
// This early return may cause goroutine leaks.
return nil, r.err
}
results = append(results, r.res)
}
return results, nil
}
由于 ch 是无缓冲通道,每个工作 goroutine 发送结果时都会阻塞,直到主 goroutine 从通道接收结果。如果 processWorkItems 因错误提前返回,接收循环就会结束,剩余所有发送方 goroutine 都会永远阻塞。
这个例子代表了一种在真实 Go 程序中发现的常见错误,Uber 的生产服务也曾出现过。下面看看如何使用新的 goroutine 泄漏剖析器找到这些泄漏。
使用 goroutine 泄漏剖析器调试
可以通过 runtime/pprof 包获取类型为 goroutineleak 的剖析报告,也可以安装 net/http/pprof 包定义的处理器。如果服务已经配置了 net/http/pprof,就无需额外接入:在处理器所监听的主机和端口上,报告会自动通过 /debug/pprof/goroutineleak 端点供采集。
把这个并发错误放进一个完整程序,并配置 net/http/pprof,就可以自行尝试:
package main
import (
"errors"
"log"
"net/http"
_ "net/http/pprof"
"time"
)
type workItem int
type workResult int
func processWorkItem(w workItem) (workResult, error) {
time.Sleep(10 * time.Millisecond)
if w == 5 {
return 0, errors.New("simulated error")
}
return workResult(w * 2), nil
}
type result struct {
res workResult
err error
}
func processWorkItems(ws []workItem) ([]workResult, error) {
ch := make(chan result)
for _, w := range ws {
go func() {
res, err := processWorkItem(w)
ch <- result{res, err}
}()
}
var results []workResult
for range len(ws) {
r := <-ch
if r.err != nil {
return nil, r.err
}
results = append(results, r.res)
}
return results, nil
}
func main() {
// Start pprof server
go func() {
log.Println(http.ListenAndServe("localhost:6060", nil))
}()
// Repeatedly trigger the leak
for {
items := []workItem{1, 2, 3, 4, 5, 6, 7, 8, 9, 10}
_, err := processWorkItems(items)
if err != nil {
log.Printf("Error processing items: %v", err)
}
time.Sleep(time.Second)
}
}
先构建上面的程序,再运行它:
$ go build -o leaky
$ ./leaky
采集剖析报告
程序运行不久后就会开始积累泄漏。可以通过 http://localhost:6060/debug/pprof 的网页界面查看。
也可以用 curl 采集 goroutine 泄漏报告,再用 go tool pprof 检查。以下命令与终端内容来自原文示例,报告中的时间、数量和行号均为示例值:
$ curl http://localhost:6060/debug/pprof/goroutineleak > leak.prof
$ go tool pprof leak.prof
Type: goroutineleak
Time: 2026-03-01 13:19:49 UTC
Entering interactive mode (type "help" for commands, "o" for options)
(pprof) list processWorkItems
Total: 116
ROUTINE ======================== main.processWorkItems.func1 in .../main.go
0 116 (flat, cum) 100% of Total
. . 31: go func() {
. . 32: res, err := processWorkItem(w)
. 116 33: ch <- result{res, err}
. . 34: }()
这份报告指出,goroutine 泄漏发生在 ch <- result{res, err},即示例报告的第 33 行,直接定位到了导致泄漏的操作。程序运行得越久,泄漏的 goroutine 通常就会越多。
修复泄漏
为 ch 配置缓冲区,就可以修复这个泄漏:
ch := make(chan result, len(ws))
这样,即使 processWorkItems 提前返回,所有处理工作项的 goroutine 仍然可以发送消息,而不会阻塞。
补充示例一节还列出了更多真实项目中的案例。
实现原理
本节介绍 goroutine 泄漏剖析器如何在底层检测泄漏。如果只关心局限与性能开销,可以直接阅读检测局限及其后面的性能部分。
核心概念
先观察一个简单情况:如果某个 goroutine 阻塞在一个并发原语上,而其他 goroutine 都无法通过内存引用访问这个原语,它显然已经泄漏。这个观察提供了重要线索,也可以进一步推广,用来定义 goroutine 何时没有泄漏。我们把这个性质称为活性(liveness),并按归纳方式定义:
一个 goroutine 具有活性,当且仅当满足以下任一条件:
- 它没有阻塞在并发原语上;
- 使它阻塞的并发原语中,至少有一个被另一个具有活性的 goroutine 引用。
在基本情形中,没有阻塞的 goroutine 显然没有泄漏。在归纳情形中,底层假设是:任何没有泄漏的 goroutine 都可能在将来使用它所引用的并发原语,解除阻塞在这些原语上的其他 goroutine。
要找出所有具有活性的 goroutine,可以从明显具有活性的、未阻塞的 goroutine 开始,沿着它们持有的引用追踪,例如局部变量中的引用,从而找到它们能够访问的并发原语。接着,把阻塞在这些原语上的 goroutine 也纳入具有活性的集合,再重复这个过程,直到不再发现新的具有活性的 goroutine。
Go 运行时的垃圾回收器(GC)已经会计算内存可达性,因此下一步就是调整 GC,让它服务于泄漏检测。下面两张原图可以帮助比较常规 GC 与修改后的 GC:
常规 GC 的标记方式,图片来自原文。
修改后的 GC 标记方式,图片来自原文。
不必彻底重写 GC。Go 运行时使用并发的三色标记清除垃圾回收器,如今还有 Green Tea 变体,其工作方式已经与我们的目标较为一致。只需要做几项关键修改:
-
在初始阶段,常规 GC 把所有 goroutine 和全局变量都视为可达对象,让它们不会被当作垃圾;它们也就是标记根。修改后的方式只把未阻塞的 goroutine 纳入最初的 goroutine 标记根,因为这些 goroutine 可以保证具有活性。
-
随后进入标记阶段:GC 追踪标记根直接或间接引用的对象,把它们标记为仍在使用的内存。虽然这个阶段本身没有直接修改,但第 1 步的变化使 GC 能先标记具有活性的 goroutine 所引用的内存。
-
标记阶段收尾时,检查第 1 步没有纳入标记根的所有阻塞 goroutine。如果使某个 goroutine 阻塞的并发原语中,至少有一个已经在第 2 步被标记,就把这个 goroutine 加为标记根,再从第 2 步继续标记。这与活性定义中的归纳步骤相对应。
-
找到所有具有活性的 goroutine 后,把仍未加入标记根的 goroutine 标记为已泄漏。
-
最后,把所有已泄漏的 goroutine 也加入标记根,再进行一轮标记,让 GC 仍然标记到常规回收周期本来应当标记的全部内存。
GC 周期完成后,goroutine 泄漏剖析器按照类似普通 goroutine 剖析的方式采集信息,只筛选出明确标记为已泄漏的 goroutine。
检测局限
上面的例子说明 goroutine 泄漏剖析报告很有用。不过,基于垃圾回收的检测也存在一些局限,可能导致漏报:
-
内存可达范围过宽。 如果一个并发原语始终能通过全局变量或可运行的 goroutine 访问,那么阻塞在它上面的 goroutine 永远不会被报告为泄漏,即使今后实际上再也不会使用这个原语。
更严格地管理对并发原语引用的访问,并更清楚地划定它们的生命周期,可以缓解这个问题。
-
非标准阻塞。 为保证检测正确,goroutine 泄漏检测严格限定于 Go 的原生并发原语:通道发送与接收,包括在
nil通道上的操作;会阻塞的select语句,也就是没有default分支的select,包括没有任何分支的select;以及sync包中的Mutex、RWMutex、WaitGroup和Cond。因其他原因阻塞的 goroutine,例如文件或网络 I/O、直接系统调用,都不会被这个机制认定为泄漏。用户自定义的并发机制,例如自旋锁,也同样不在检测范围内,除非其底层实现依赖上述原语。
-
非确定性。 检测只能在泄漏发生后发现它,无法预测尚未发生的泄漏,因此在行为不稳定的程序中复现和诊断泄漏仍然困难。为获得更好的效果,可以结合多种方法:在包括生产环境在内的不同层次使用 goroutine 泄漏剖析,同时维护全面的测试套件,并接入
goleak与synctest。
性能影响
Goroutine 泄漏检测经过谨慎设计,尽量降低性能影响,但仍然存在成本。
内存开销很小,只需少量额外空间记录状态;不过,泄漏检测可能比常规 GC 更慢。一个称为“雏菊链”(daisy-chain)的极端情形最能说明这一点:
雏菊链结构,图片来自原文。
这个例子没有泄漏:可运行的 goroutine G₀ 引用了并发原语 P₁,P₁ 使 G₁ 阻塞,后面依此类推。
因此,要证明 Pᵢ₊₁ 对应的活性,必须先证明 Pᵢ 对应的活性。这带来两项成本:
-
GC 标记阶段在 goroutine 的扫描顺序上实际上被串行化。必须先标记完从某个 Pᵢ 可达的全部内存,才能把 Pᵢ₊₁ 对应的 goroutine 加为根。
-
当前实现在每一轮标记结束时检查所有阻塞的 goroutine,因此一次 GC 周期最坏需要 O(n²) 步,其中 n 是 goroutine 的总数。
第二点将来可以优化,而第一点是泄漏检测本身固有的局限,无法绕开。
不过,除非通过运行时选项另行配置,GC 仍然与用户代码并发运行。此外,如果在执行过程的某个时刻能够观察到 goroutine 泄漏,那么同一次执行的任何后续时刻也都能够观察到它。因此,定期采集报告的基础设施可以调整采样频率,例如每 4 小时一次,以降低开销,而几乎不牺牲检测泄漏的能力。
致谢
Goroutine 泄漏检测源于奥胡斯大学、圣路易斯华盛顿大学与 Uber 的研究合作。相关工作发表在论文 “Dynamic Partial Deadlock Detection and Recovery via Garbage Collection” 中,作者为 Saioc 等人,发表于 ASPLOS 2025。
从学术原型转变为实际 Go 功能,得益于 Google Go 团队的 Michael Knyszek、Michael Pratt,以及 PJ Malloy(@thepudds)的指导。
补充示例
以下导致泄漏的编码模式来自工业级代码库与开源项目,按复杂程度递增排列。
可以在 Go Playground 中快速尝试这些例子,也可以构造自己的泄漏实验。下面的所有 pprof 报告均保留原文示例输出。
示例:发送两次
最简单的一类泄漏,是在通道上发送的消息数量超过预期。下面的 goroutine 本应通过无缓冲通道向主 goroutine 发送一条消息,但错误分支发送消息后漏写了 return。因此,每次出现错误时,发送方都会尝试发送两条消息,从而导致泄漏。
func DoubleSend() {
ch := make(chan any)
go func(err error) {
if err != nil {
// In case of an error, send nil.
ch <- nil
// Return statement is missing.
}
// Otherwise, continue with normal behaviour.
// This send is still executed, which causes a leak in the error case.
ch <- struct{}{}
}(fmt.Errorf("error"))
// Receive only one message.
<-ch
}
报告不会直接指出缺失的 return 就是根因,但它会突出显示发生泄漏的发送操作,至少能把我们引向出错的函数。
(pprof) list DoubleSend
Total: 1
ROUTINE ======================== main.DoubleSend.func1 in .../main.go
0 1 (flat, cum) 100% of Total
. . 118: go func(err error) {
. . 119: if err != nil {
. . 121: ch <- nil
. . 123: }
. 1 126: ch <- struct{}{}
. . 127: }(fmt.Errorf("error"))
. . 129: <-ch
只需要在错误分支的发送操作后补上 return,就可以修复这个泄漏。
示例:提前返回
相反的情况也很常见:接收方在某些控制流路径中跳过了通信。下面实际上是开篇示例的简化版本。
// Incoming error simulates an error produced internally.
func EarlyReturn(err error) {
ch := make(chan any)
// Create a worker goroutine.
go func() {
// Send something to the channel.
// Leaks if the parent goroutine terminates early.
ch <- struct{}{}
}()
if err != nil {
// The parent goroutine quits too early in case of an error.
// Sender leaks.
return
}
// Receive is only executed if there is no error.
<-ch
}
报告会暴露这个 goroutine 泄漏:
ROUTINE ======================== main.EarlyReturn.func1 in .../main.go
0 1 (flat, cum) 100% of Total
. . 140: go func() {
. 1 143: ch <- struct{}{}
. . 144: }()
. . 145:
. . 146: if err != nil {
为 ch 配置容量为 1 的缓冲区,就可以修复该泄漏。
示例:超时
上面的提前返回模式还有一个变体,涉及上下文(context)与非确定性选择,即 select 语句:
func Timeout(ctx context.Context) {
// An unbuffered channel is used to coordinate
// a worker and parent thread
ch := make(chan any)
// Create worker goroutine
go func() {
// Perform some work then signal to the parent thread.
ch <- struct{}{}
}()
// Wait for message from worker or context
// to be cancelled or timed out.
select {
case <-ch: // Receive message from worker
case <-ctx.Done():
// Sender leaks because there is no
// future rendezvous over the channel.
}
}
如果上下文在发送方与父 goroutine 同步之前被取消,发送方就会泄漏:
(pprof) list Timeout
Total: 10
ROUTINE ======================== main.Timeout.func1.1 in .../main.go
0 10 (flat, cum) 100% of Total
. . 198: go func() {
. 10 201: ch <- struct{}{}
. . 202: }()
和前一个例子一样,修复方式是为通道设置容量为 1 的缓冲区。
示例:遍历通道后没有关闭
使用 range 遍历通道,可以在循环中反复从通道接收值。通道关闭后,待其缓冲区内排队的所有值都被接收,循环才会退出。
如果通道始终没有关闭,range 循环就会让执行它的 goroutine 永远阻塞。漏掉 close 操作是常见错误,例如:
// Incoming list of items and the number of workers.
func noCloseRange(list []any, workers int) {
// Create a channel that distributes work items.
ch := make(chan any)
// Create the worker goroutines.
for i := 0; i < workers; i++ {
go func() {
// Each worker pulls items from the channel
// and then processes it.
for item := range ch {
// Process each item
_ = item
}
}()
}
// Queue items to the workers by using the channel.
for _, item := range list {
// The parent leaks by sending an item if workers == 0
// or if all the workers panic, but the panic is recovered.
ch <- item
}
// Otherwise, the channel is never closed, so workers
// leak once there are no more items left to process.
}
...
go noCloseRange([]any{1, 2, 3}, 3) // Leaks all 3 workers
这个程序的 goroutine 泄漏报告会包含如下内容:
Type: goroutineleak
(pprof) list noCloseRange.func1
Total: 4
ROUTINE ======================== main.noCloseRange.func1 in .../main.go
0 3 (flat, cum) 75.00% of Total
. . 82: go func() {
. 3 84: for item := range ch {
. . 86: _ = item
. . 87: }
. . 88: }()
可以看到,三个工作 goroutine 都阻塞在 range ch 上,这已经给出了明确线索。发送完全部消息后关闭通道,就可以修复这个泄漏:
for _, item := range list {
ch <- item
}
// All items have been sent. It is now safe to close.
close(ch)
细心的读者可能还发现了另一个潜在泄漏:如果误把工作 goroutine 数量设置为零,父 goroutine 的发送操作也会泄漏:
go noCloseRange([]any{1, 2, 3}, 0) // Sender leaks with 0 workers
报告同样可以捕获这一情况:
(pprof) list noCloseRange$
Total: 4
ROUTINE ======================== main.noCloseRange in .../main.go
0 1 (flat, cum) 25.00% of Total
. . 76:func noCloseRange(list []any, workers int) {
...
. . 92: for _, item := range list {
. 1 95: ch <- item
. . 96: }
真实生产系统通常可以假定 workers > 0 成立,但 goroutine 泄漏报告仍可监测偶然违反这一前提的情况,而不必仅为这种监测目的在代码中加入保守的 workers <= 0 检查。
示例:违反方法调用约定
前面几个模式的词法作用域都比较集中。但当功能分散在多个函数、方法和包中,且接口隐藏了实现细节时,人工发现泄漏会困难得多。
下面的自定义 worker 类型就展示了这种情况。它有两个通道字段 ch 和 done。Start 方法创建一个循环运行的 goroutine,通过 select 从这两个通道读取。这个 goroutine 只有从 done 通道接收时才会终止,而 done 由 Stop 方法关闭。
Start 可以调用任意多次;但只要至少调用过一次,最终就应当调用 Stop。
因此,Start 与 Stop 构成了一项隐含约定,限定这两个方法的调用顺序。违反约定会导致不良行为,在这个例子中就是 goroutine 泄漏:
func MethodContractViolation() {
items := make([]any, 10)
// Create a new worker
w := NewWorker()
// Start worker
w.Start()
// Operate on worker
for _, item := range items {
w.AddToQueue(item)
}
// Exits without calling ’Stop’.
}
type worker struct {
ch chan any
done chan any
}
type Worker interface {
Start()
Stop()
AddToQueue(item any)
}
func NewWorker() Worker {
return &worker{
ch: make(chan any),
done: make(chan any),
}
}
// Start spawns a background goroutine that extracts items pushed to the queue.
func (w *worker) Start() {
go func() {
for {
select {
case <-w.ch: // Normal workflow
case <-w.done:
return // Shut down
}
}
}()
}
func (w *worker) Stop() {
// Allows goroutine created by Start to terminate
close(w.done)
}
func (w *worker) AddToQueue(item any) {
w.ch <- item
}
实践中问题可能更严重,因为这类自定义类型常常只以接口形式对外暴露。这里的 Worker 接口名称并没有揭示底层实现。调用者甚至可能不知道内部启动了 goroutine,因此无意间违反了这项隐含约定。
好在采集 goroutine 泄漏报告能够暴露这个缺陷:
(pprof) list Start
Total: 1
ROUTINE ======================== main.(*worker).Start.func1 in .../main.go
0 1 (flat, cum) 100% of Total
. . 266: go func() {
. . 267: for {
. 1 268: select {
. . 269: case <-w.ch:
. . 270: case <-w.done:
. . 271: return
修复时需要追踪到调用 Start 的位置,并补上对 Stop 的调用。
示例(CockroachDB):忘记解锁
下面的案例来自 CockroachDB。代码在循环中获取和释放锁,但在执行 break 前忘记了解锁:
type Gossip struct {
mu sync.Mutex
closed bool
}
func (g *Gossip) bootstrap() {
for {
g.mu.Lock()
if g.closed {
// Missing g.mu.Unlock
break
}
g.mu.Unlock()
}
}
func Cockroach584() {
g := &Gossip{
closed: true,
}
// ...
g.bootstrap()
g.bootstrap() // Causes a leak
}
这种情况下,goroutine 再次尝试获取锁时会阻塞并泄漏。
(pprof) list Gossip
Total: 1
ROUTINE ======================== main.(*Gossip).bootstrap in .../main.go
0 1 (flat, cum) 100% of Total
. . 165:func (g *Gossip) bootstrap() {
. . 166: for {
. 1 167: g.mu.Lock()
. . 168: if g.closed {
. . 170: break
. . 171: }
. . 172: g.mu.Unlock()
在 break 前加入一次 Unlock 调用,就可以解决问题。
示例(etcd):通道操作顺序超出预期
etcd 中的这个案例展示了通道操作的意外执行顺序如何导致 goroutine 泄漏:
type node struct {
status chan chan struct{}
stop chan struct{}
done chan struct{}
}
func (n *node) Status() struct{} {
c := make(chan struct{})
n.status <- c
return <-c
}
func (n *node) run() {
for {
select {
case c := <-n.status:
c <- struct{}{}
case <-n.stop:
close(n.done)
return
}
}
}
func (n *node) Stop() {
select {
case n.stop <- struct{}{}:
case <-n.done:
return
}
<-n.done
}
func Etcd6857() {
n := &node{
status: make(chan chan struct{}),
stop: make(chan struct{}),
done: make(chan struct{}),
}
go n.run()
go n.Status()
go n.Stop()
}
run 方法启动一个循环,反复从 status 通道接收由 Status 方法发送的消息。同时,它也可能从 stop 通道接收由 Stop 方法发送的一条消息;一旦收到,就关闭 done 通道并退出。Stop 方法随后会等待从 done 接收,而关闭 done 会解除这个等待。
如果 run、Status 和 Stop 并发执行,就可能发生泄漏。Stop 与 run 对应的 goroutine 可能先完成同步并退出,尚未接收 Status 发送的消息,导致 Status 永久阻塞。
(pprof) list Status
Total: 8
ROUTINE ======================== main.(*node).Status in .../main.go
0 8 (flat, cum) 100% of Total
. . 16:func (n *node) Status() struct{} {
. . 17: c := make(chan struct{})
. 8 18: n.status <- c
. . 19: return <-c
. . 20:}
把向 status 的发送操作放进一个 select,并让另一个 case 尝试从 done 接收,就能让执行 Status 的 goroutine 在输给 Stop 调用的竞速时正常退出。
示例(Kubernetes):通道与互斥锁相互阻塞
Kubernetes 中的这个案例由混用通道和锁引起:
type Connection struct {
closeChan chan bool
}
type idleAwareFramer struct {
resetChan chan bool
writeLock sync.Mutex
conn *Connection
}
func (i *idleAwareFramer) monitor() {
var resetChan = i.resetChan
for range i.conn.closeChan {
i.writeLock.Lock()
close(resetChan)
i.resetChan = nil
i.writeLock.Unlock()
break
}
}
func (i *idleAwareFramer) WriteFrame() {
i.writeLock.Lock()
defer i.writeLock.Unlock()
if i.resetChan == nil {
return
}
i.resetChan <- true
}
func NewIdleAwareFramer() *idleAwareFramer {
return &idleAwareFramer{
resetChan: make(chan bool),
conn: &Connection{
closeChan: make(chan bool),
},
}
}
func Kubernetes6632() {
i := NewIdleAwareFramer()
go func() {
i.conn.closeChan <- true
}()
go i.monitor()
go i.WriteFrame()
}
执行 WriteFrame 的 goroutine 可能先获得 framer 的写锁,然后向 resetChan 发送消息;同时,monitor 对应的 goroutine 正在等待从 closeChan 接收消息。收到消息后,monitor 会尝试获取同一把写锁。但没有 goroutine 从 resetChan 接收,发送就会永久阻塞:WriteFrame 一直持有写锁,无法执行延迟的解锁操作,monitor 则一直等待这把锁。
因此,两个 goroutine 都会泄漏。
(pprof) list AwareFramer
Total: 200
ROUTINE ======================== main.(*idleAwareFramer).WriteFrame in .../main.go
0 100 (flat, cum) 50.00% of Total
. . 32:func (i *idleAwareFramer) WriteFrame() {
. . 33: i.writeLock.Lock()
. . 34: defer i.writeLock.Unlock()
. . 35: if i.resetChan == nil {
. . 36: return
. . 37: }
. 100 38: i.resetChan <- true
. . 39:}
ROUTINE ======================== main.(*idleAwareFramer).monitor in .../main.go
0 100 (flat, cum) 50.00% of Total
. . 21:func (i *idleAwareFramer) monitor() {
. . 22: var resetChan = i.resetChan
. . 23: for range i.conn.closeChan {
. 100 24: i.writeLock.Lock()
. . 25: close(resetChan)
修复方法是:在 monitor 从 closeChan 收到消息后,启动一个独立的 goroutine 从 resetChan 取走消息,再尝试获取写锁。
示例(Moby):误用 sync.WaitGroup
type Manager struct {
plugins []int
}
func (pm *Manager) init() {
var group sync.WaitGroup
group.Add(len(pm.plugins))
for _, p := range pm.plugins {
go func(p int) {
defer group.Done()
}(p)
group.Wait() // Block here
}
}
func Moby25384() {
pm := &Manager{
plugins: []int{1, 2},
}
go pm.init()
}
group 等待组先根据插件管理器 pm 持有的插件数量增加计数,再遍历插件,为每个插件启动一个 goroutine。每个 goroutine 完成任务后,用 Done 把计数减一。然而,代码错误地在循环体内调用 Wait,而不是在循环结束后调用。只要管理器有多个插件,执行 init 的 goroutine 就会发生泄漏。
(pprof) list init
Total: 1
ROUTINE ======================== main.(*Manager).init in .../main.go
0 1 (flat, cum) 100% of Total
. . 17: group.Add(len(pm.plugins))
. . 18: for _, p := range pm.plugins {
. . 19: go func(p int) {
. . 20: defer group.Done()
. . 21: }(p)
. 1 22: group.Wait() // Block here
. . 23: }
把 Wait 移到循环外,就可以解决问题。
示例(Moby):通道与互斥锁相互阻塞
type (
State struct {
Health *Health
}
Container struct {
sync.Mutex
State *State
}
Store struct {
ctr *Container
}
Daemon struct {
containers Store
}
Health struct {
stop chan struct{}
}
)
func (d *Daemon) StateChanged() {
c := d.containers.ctr
c.Lock()
d.updateHealthMonitorElseBranch(c)
defer c.Unlock()
}
func (d *Daemon) updateHealthMonitorElseBranch(c *Container) {
c.State.Health.CloseMonitorChannel()
}
func (s *Health) CloseMonitorChannel() {
if s.stop != nil {
s.stop <- struct{}{}
}
}
func monitor(c *Container, stop chan struct{}) {
for {
select {
case <-stop:
return
default:
handleProbeResult(c)
}
}
}
func handleProbeResult(c *Container) {
c.Lock()
defer c.Unlock()
// Additional work...
}
func NewDaemonAndContainer() (*Daemon, *Container) {
c := &Container{
State: &State{&Health{
stop: make(chan struct{}),
}},
}
d := &Daemon{Store{c}}
return d, c
}
func Moby28462() {
d, c := NewDaemonAndContainer()
go monitor(c, c.State.Health.stop)
go d.StateChanged()
}
调用 StateChanged 的 goroutine 可能先获得 daemon 保存的容器锁,再调用 updateHealthMonitorElseBranch;后者会尝试向容器的 stop 通道发送消息。另一方面,如果执行 monitor 的 goroutine 在检查 select 时没有可立即接收的 stop 消息,就会选择 default 分支,进入 handleProbeResult。
handleProbeResult 会尝试获取同一把容器锁,而锁已经被 StateChanged 对应的 goroutine 持有。后者又等待 stop 的接收方,双方因此都无法继续执行,最终泄漏。在原代码中,defer c.Unlock() 还位于可能阻塞的 updateHealthMonitorElseBranch(c) 调用之后;如果该调用一直不返回,这个 defer 也尚未注册。
(pprof) list .CloseMonitorChannel
Total: 2
ROUTINE ======================== main.(*Health).CloseMonitorChannel in .../main.go
0 1 (flat, cum) 50.00% of Total
. . 66:func (s *Health) CloseMonitorChannel() {
. . 67: if s.stop != nil {
. 1 68: s.stop <- struct{}{}
. . 69: }
. . 70:}
(pprof) list main.handleProbeResult
Total: 2
ROUTINE ======================== main.handleProbeResult in .../main.go
0 1 (flat, cum) 50.00% of Total
. . 83:func handleProbeResult(c *Container) {
. 1 84: c.Lock()
. . 85: // Additional work...
. . 86: defer c.Unlock()
. . 87:}
修复方式是关闭 stop 通道,而不是向它发送消息。关闭通道不会阻塞,所以 StateChanged 对应的 goroutine 能继续执行并释放锁。这样,monitor 就能继续运行,在下一次循环的 select 中选择已经就绪的 <-stop 分支并终止。
来源与许可
原文:Goroutine Leak Profiles,作者 Vlad Saioc,发表于 2026 年 9 月 2 日。正文及原文结构图按 CC BY 4.0 使用;中文正文为翻译,技术说明中补充了两个持锁阻塞位置的澄清,代码与原文示例输出保持不变。许可依据:Go 网站版权说明。
示例代码按 Go BSD 许可 使用,保留以下版权声明、条件与免责声明:
Copyright 2009 The Go Authors.
Redistribution and use in source and binary forms, with or without modification, are permitted provided that the following conditions are met:
- Redistributions of source code must retain the above copyright notice, this list of conditions and the following disclaimer.
- Redistributions in binary form must reproduce the above copyright notice, this list of conditions and the following disclaimer in the documentation and/or other materials provided with the distribution.
- Neither the name of Google LLC nor the names of its contributors may be used to endorse or promote products derived from this software without specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS “AS IS” AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT OWNER OR CONTRIBUTORS BE LIABLE FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.











暂无评论内容