SKILL·C0EF20

go-troubleshooting

eduardo-sl
Aktualisiert 27 days ago
6 Ansichten
70
9
70
Auf GitHub ansehen
Testenaitestingdesign

Über

Diese Fähigkeit diagnostiziert Laufzeitprobleme in Go-Programmen, einschließlich Panics, Deadlocks, Goroutine-/Speicherlecks und OOM-Kills. Sie hilft Entwicklern, Stack-Traces und Race-Reports zu interpretieren sowie Tools wie delve und pprof zu nutzen. Verwenden Sie sie zur Fehlerbehebung bei aktiven Ausfällen, jedoch nicht für Leistungsoptimierung, das Schreiben neuer nebenläufiger Code oder Testdesign.

Schnellinstallation

Claude Code

Empfohlen
Primär
npx skills add eduardo-sl/go-agent-skills -a claude-code
Plugin-BefehlAlternativ
/plugin add https://github.com/eduardo-sl/go-agent-skills
Git CloneAlternativ
git clone https://github.com/eduardo-sl/go-agent-skills.git ~/.claude/skills/go-troubleshooting

Kopieren Sie diesen Befehl und fügen Sie ihn in Claude Code ein, um diese Fähigkeit zu installieren

Dokumentation

Go Troubleshooting

Diagnosis before fixes. Reproduce, observe, localize, then change code. Never "fix" a symptom you haven't explained — the bug will move.

1. Pick the Procedure by Symptom

SymptomProcedure
Crash with stack trace§2 Read the panic
Program hangs / requests stall§3 Dump goroutines, find the block
fatal error: all goroutines are asleep§3 — Go detected total deadlock
Memory grows until OOM§4 Heap profile diff
Goroutine count grows§5 Goroutine profile diff
Intermittent corrupt data / weird values§6 Race detector
Need to inspect state interactively§7 Delve

2. Reading a Panic

panic: runtime error: invalid memory address or nil pointer dereference
[signal SIGSEGV: segmentation violation code=0x1 addr=0x0 pc=0x6bb0e4]

goroutine 43 [running]:
myapp/internal/service.(*UserService).Notify(0x0, {0xc000123456?, ...})
        /app/internal/service/user.go:87 +0x24
myapp/internal/handler.(*Handler).Create(0xc0001a2000, ...)
        /app/internal/handler/user.go:41 +0x1c5

Read it mechanically:

  1. First line: what kind of panic. nil pointer dereference + addr=0x0 means a nil receiver, nil field, or nil map/pointer argument.
  2. Top frame in YOUR code: user.go:87 — go there.
  3. Receiver value in the frame: (*UserService).Notify(0x0, ...) — the 0x0 first argument IS the receiver: the service itself was nil. Trace where it was constructed (or wasn't).
  4. goroutine 43 — if it's not goroutine 1, find who spawned it and whether a recover boundary should exist there.

3. Hangs and Deadlocks

Get a goroutine dump from the hanging process:

kill -QUIT <pid>      # dumps all goroutine stacks to stderr, then exits
# or, if net/http/pprof is mounted (see §4):
curl 'localhost:6060/debug/pprof/goroutine?debug=2'

Then classify the stacks:

  • [semacquire] on sync.(*Mutex).Lock — find which goroutine HOLDS the mutex: look for another stack inside the critical section. Two goroutines each holding one of two locks = lock-order inversion.
  • [chan send] / [chan receive] — the other side is gone. Find who should be receiving/sending and why it exited (or was never started).
  • [select] with a ctx.Done() case missing — blocked call that ignores cancellation.
  • Hundreds of identical stacks — that's a leak (§5), not a deadlock.

4. Memory Leaks

Mount pprof in long-running services (private port only, never public):

import _ "net/http/pprof"

go func() {
    log.Println(http.ListenAndServe("localhost:6060", nil))
}()

Diff heap profiles over time — a leak is growth that never returns:

curl -s localhost:6060/debug/pprof/heap > heap1.pb.gz
sleep 300   # let the leak accumulate
curl -s localhost:6060/debug/pprof/heap > heap2.pb.gz
go tool pprof -base heap1.pb.gz heap2.pb.gz
(pprof) top          # biggest positive delta = the leak
(pprof) list FuncName

Usual suspects: unbounded caches/maps without eviction, subslices pinning large arrays, time.Ticker never stopped, response bodies not closed, growing global slices, forgotten goroutines holding buffers.

5. Goroutine Leaks

curl -s localhost:6060/debug/pprof/goroutine > g1.pb.gz
sleep 300
curl -s localhost:6060/debug/pprof/goroutine > g2.pb.gz
go tool pprof -base g1.pb.gz g2.pb.gz
(pprof) top    # the growing stack is your leak site

The leaking stack tells you which go statement never terminates. Fix the termination path (context, channel close) — patterns in the concurrency skill. In tests, goleak (uber-go/goleak) fails a test that leaves goroutines behind.

6. Race Detector

go test -race ./...        # in CI, always
go build -race ./cmd/api   # staging binaries under real traffic

A report shows two stacks: the write and the concurrent read/write, each with the goroutine's creation site. The fix is never "add a sleep" — protect the state (mutex), transfer ownership (channel), or make it immutable. -race only reports races that actually executed: a clean run proves nothing about untested paths.

7. Delve

dlv test ./internal/service -- -test.run TestTransfer   # debug a test
dlv attach <pid>                                        # running process
dlv core ./api core.1234                                # post-mortem

(dlv) break user.go:87
(dlv) continue
(dlv) print svc.repo          # inspect exact values
(dlv) goroutines -t           # all goroutines with stacks
(dlv) goroutine 43 bt         # switch and backtrace

Use delve when you need actual values or goroutine states, not just locations. For quick localizations, a focused t.Logf or slog.Debug plus one test run is often faster.

8. Diagnostic Environment Variables

GOTRACEBACK=all ./api        # panic dumps ALL goroutines, not just one
GODEBUG=gctrace=1 ./api      # GC cycles: pacing, heap goal, pause times
GOMEMLIMIT=512MiB ./api      # soft memory limit — mitigates OOM while
                             # you find the real leak

Verification Checklist

  1. Symptom reproduced (or captured via dump/profile) before any code change
  2. Root cause explained: you can say WHY the failure happened at that site
  3. Panic fixes address the nil/bounds source, not a wrapper recover
  4. Deadlock fixes establish a single lock order or remove the shared lock
  5. Leak fixes verified: goroutine/heap profile flat after the fix
  6. go test -race ./... passes after concurrency-related fixes
  7. A regression test now fails without the fix
  8. pprof endpoints bound to localhost/private interfaces only

GitHub Repository

eduardo-sl/go-agent-skills
Pfad: skills/(safety)/go-troubleshooting
0
FAQ

Häufig gestellte Fragen

Was ist der Skill go-troubleshooting?

go-troubleshooting ist ein Claude Skill von eduardo-sl. Skills bündeln Anweisungen und Ressourcen, die Claude bei Bedarf lädt, um Aufgaben rund um go-troubleshooting ohne zusätzliche Eingaben auszuführen.

Wie installiere ich go-troubleshooting?

Verwende die Installationsbefehle auf dieser Seite: Füge go-troubleshooting als Plugin zu Claude Code hinzu oder klone das Repository in dein Skills-Verzeichnis. Starte Claude danach neu, damit der Skill geladen wird.

Zu welcher Kategorie gehört go-troubleshooting?

go-troubleshooting gehört zur Kategorie Testen.

Kann ich go-troubleshooting kostenlos nutzen?

Ja. go-troubleshooting ist auf AIMCP gelistet und kann kostenlos installiert werden.

Verwandte Skills

evaluating-llms-harness
Testen

Diese Claude Skill führt den lm-evaluation-harness aus, um LLMs über 60+ standardisierte akademische Aufgaben wie MMLU und GSM8K zu benchmarken. Sie wurde für Entwickler entwickelt, um Modellqualität zu vergleichen, Trainingsfortschritt zu verfolgen oder akademische Ergebnisse zu berichten. Das Tool unterstützt verschiedene Backends, einschließlich HuggingFace- und vLLM-Modelle.

Skill ansehen
cloudflare-cron-triggers
Testen

Diese Fähigkeit bietet umfassendes Wissen zur Implementierung von Cloudflare Cron Triggers, um Workers mithilfe von Cron-Ausdrücken zu planen. Sie behandelt das Einrichten periodischer Aufgaben, Wartungsjobs und automatisierter Workflows, während häufige Probleme wie ungültige Cron-Ausdrücke und Zeitzonenprobleme behandelt werden. Entwickler können sie zum Konfigurieren geplanter Handler, zum Testen von Cron-Triggers und zur Integration mit Workflows und Green Compute verwenden.

Skill ansehen
webapp-testing
Testen

Diese Claude Skill bietet ein Playwright-basiertes Toolkit zum Testen lokaler Webanwendungen durch Python-Skripte. Es ermöglicht Frontend-Verifizierung, UI-Debugging, Screenshot-Aufnahme und Log-Einblick bei gleichzeitiger Verwaltung von Server-Lebenszyklen. Nutzen Sie es für Browser-Automatisierungsaufgaben, führen Sie Skripte jedoch direkt aus, anstatt deren Quellcode zu lesen, um Kontextverschmutzung zu vermeiden.

Skill ansehen
finishing-a-development-branch
Testen

Diese Fähigkeit unterstützt Entwickler dabei, abgeschlossene Arbeiten zu finalisieren, indem sie testet, ob Tests bestehen, und dann strukturierte Integrationsoptionen präsentiert. Sie leitet den Workflow für das Zusammenführen von Code, das Erstellen von PRs oder das Bereinigen von Branches nach Abschluss der Implementierung. Nutzen Sie sie, wenn Ihr Code bereit und getestet ist, um den Entwicklungsprozess systematisch abzuschließen.

Skill ansehen