Monitoring Expert — Setup Guide¶
Source: jeffallan/claude-skills (Community)
Skill: monitoring-expert · Installs: 3,900+ · Category: DevOps & Infrastructure
Platform: Linux, macOS, Windows
Monitoring Expert is an observability and performance specialist skill that implements comprehensive monitoring, alerting, tracing, and performance testing systems. It covers Prometheus/Grafana stacks, structured logging pipelines, distributed tracing instrumentation, load testing with k6/Artillery, and capacity planning. Ideal for Hermes agents managing infrastructure or debugging production issues.
Installation¶
npx skills add jeffallan/claude-skills@monitoring-expert
What It Does¶
The skill follows a four-stage observability workflow:
- Assess — Identify what needs monitoring (SLIs, critical paths, business metrics)
- Instrument — Add logging, metrics, and traces to applications
- Collect — Configure aggregation and storage (Prometheus, log shippers, OTLP)
- Visualize — Build dashboards using RED (Rate/Errors/Duration) or USE (Utilization/Saturation/Errors) methods
Prerequisites¶
| Requirement | Details |
|---|---|
| Hermes Agent | v1.0+ |
| Target application | Access to instrument code |
| Monitoring stack | Prometheus, Grafana, or Datadog (optional) |
| Load testing tools | k6 or Artillery (optional) |
Usage with Hermes¶
Trigger the skill for any observability task:
"Set up application monitoring for our Node.js API"
"Add structured logging to the Python service"
"Create a Grafana dashboard for API latency"
"Run a load test against production with k6"
"Profile CPU/memory bottlenecks in the worker process"
"Define alerting rules for 5xx error rates"
Example: Setting Up API Monitoring¶
"Configure Prometheus metrics for our Express.js API — track request rate, error rate, and p95 latency"
The skill instruments the application with appropriate metrics libraries, configures scrape endpoints, builds Grafana dashboards, and defines alerting rules.
Core Capabilities¶
| Capability | Tools/Frameworks |
|---|---|
| Metrics collection | Prometheus, OpenTelemetry |
| Visualization | Grafana, Datadog |
| Logging pipelines | Structured logging, log shippers |
| Distributed tracing | OpenTelemetry, Jaeger |
| Load testing | k6, Artillery |
| Profiling | CPU/memory profilers |
| Capacity planning | Trend analysis, forecasting |
Guardrails¶
- Always verify data arrives before building dashboards
- Use RED method for services, USE method for resources
- Alert on symptoms, not causes
- Never alert on raw metrics without aggregation windows
Related Skills¶
- AWS Agent Toolkit Setup — AWS infrastructure monitoring
- Sentry AI Monitoring Setup — Error tracking
- HashiCorp Agent Skills Setup — Infrastructure management
Source¶
- skills.sh: jeffallan/claude-skills@monitoring-expert
- GitHub: github.com/jeffallan/claude-skills