deploying-monitoring-stacks

jeremylongshore

Updated Today

16 views

712

Metadesigndata

About

This skill generates production-ready configurations for deploying monitoring stacks like Prometheus, Grafana, and Datadog. Use it when you need to set up metric collection, visualization dashboards, and alerting rules. It provides infrastructure-aware configurations for Kubernetes, Docker, or bare metal environments.

Quick Install

Claude Code

Recommended

Plugin CommandRecommended

/plugin add https://github.com/jeremylongshore/claude-code-plugins-plus

Git CloneAlternative

git clone https://github.com/jeremylongshore/claude-code-plugins-plus.git ~/.claude/skills/deploying-monitoring-stacks

Copy and paste this command in Claude Code to install this skill

Documentation

Prerequisites

Before using this skill, ensure:

Target infrastructure is identified (Kubernetes, Docker, bare metal)
Metric endpoints are accessible from monitoring platform
Storage backend is configured for time-series data
Alert notification channels are defined (email, Slack, PagerDuty)
Resource requirements are calculated based on scale

Instructions

Select Platform: Choose Prometheus/Grafana, Datadog, or hybrid approach
Deploy Collectors: Install exporters and agents on monitored systems
Configure Scraping: Define metric collection endpoints and intervals
Set Up Storage: Configure retention policies and data compaction
Create Dashboards: Build visualization panels for key metrics
Define Alerts: Create alerting rules with appropriate thresholds
Test Monitoring: Verify metrics flow and alert triggering

Output

Prometheus + Grafana (Kubernetes):

# {baseDir}/monitoring/prometheus.yaml
apiVersion: v1
kind: ConfigMap
metadata:
  name: prometheus-config
data:
  prometheus.yml: |
    global:
      scrape_interval: 15s
      evaluation_interval: 15s
    scrape_configs:
      - job_name: 'kubernetes-pods'
        kubernetes_sd_configs:
          - role: pod
        relabel_configs:
          - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape]
            action: keep
            regex: true
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: prometheus
spec:
  replicas: 1
  template:
    spec:
      containers:
      - name: prometheus
        image: prom/prometheus:latest
        args:
          - '--config.file=/etc/prometheus/prometheus.yml'
          - '--storage.tsdb.retention.time=30d'
        ports:
        - containerPort: 9090

Grafana Dashboard Configuration:

{
  "dashboard": {
    "title": "Application Metrics",
    "panels": [
      {
        "title": "CPU Usage",
        "type": "graph",
        "targets": [
          {
            "expr": "rate(container_cpu_usage_seconds_total[5m])"
          }
        ]
      }
    ]
  }
}

Error Handling

Metrics Not Appearing

Error: "No data points"
Solution: Verify scrape targets are accessible and returning metrics

High Cardinality

Error: "Too many time series"
Solution: Reduce label combinations or increase Prometheus resources

Alert Not Firing

Error: "Alert condition met but no notification"
Solution: Check Alertmanager configuration and notification channels

Dashboard Load Failure

Error: "Failed to load dashboard"
Solution: Verify Grafana datasource configuration and permissions

Resources

Prometheus documentation: https://prometheus.io/docs/
Grafana documentation: https://grafana.com/docs/
Example dashboards in {baseDir}/monitoring-examples/

GitHub Repository

jeremylongshore/claude-code-plugins-plus

Path: plugins/devops/monitoring-stack-deployer/skills/monitoring-stack-deployer

aiautomationclaude-codedevopsmarketplacemcp

Related Skills

langchain

Algorithmic Art Generation

webapp-testing

Testing

This Claude Skill provides a Playwright-based toolkit for testing local web applications through Python scripts. It enables frontend verification, UI debugging, screenshot capture, and log viewing while managing server lifecycles. Use it for browser automation tasks but run scripts directly rather than reading their source code to avoid context pollution.

View skill

requesting-code-review

Design

This skill dispatches a code-reviewer subagent to analyze code changes against requirements before proceeding. It should be used after completing tasks, implementing major features, or before merging to main. The review helps catch issues early by comparing the current implementation with the original plan.

View skill