dt-obs-services
Service performance monitoring with RED metrics (Rate, Errors, Duration) and runtime-specific telemetry for Java, .NET, Node.js, Python, PHP, and Go. Use when analyzing service health, SLA compliance, or runtime issues. Trigger: "service response time", "error rate", "throughput", "SLA compliance", "service mesh overhead", "JVM GC", "Java heap", "Node.js event loop", ".NET CLR", "Python threads", "PHP OPcache", "Go goroutines", "service performance", "p95 latency", "request failures", "database
Security Assessment
About dt-obs-services
This skill provides guidance for monitoring application-service performance and health in Dynatrace using DQL (Dynatrace Query Language). It centers on RED metrics (Rate, Errors, Duration) and runtime-specific telemetry, helping engineers answer questions about service response times, error rates, throughput, SLA compliance, messaging health, and service-mesh overhead without hand-writing queries from scratch.
The skill documents the key service metrics (request response_time, count, and failure_count in microseconds/counts) and supplies ready-to-adapt DQL query templates for response-time percentiles (p50/p95/p99), error-rate calculation, failure spike detection, failures by HTTP status, throughput and peak-traffic detection, traffic growth, performance-degradation detection against a baseline, and Kubernetes-context breakdowns by workload, namespace, and cluster. Additional sections cover span-based advanced analysis for custom SLA thresholds and health scoring, messaging metrics for publish/receive/process rates and consumer-lag detection, and service-mesh ingress monitoring including mesh-vs-direct overhead. The skill is explicitly scoped: it defers infrastructure metrics, logs, and tracing to sibling dt-obs skills.
Target users are SRE and platform teams, application-performance engineers, and on-call responders who operate services instrumented by Dynatrace, particularly across Java, .NET, Node.js, Python, PHP, and Go runtimes and Kubernetes environments. Typical use cases include diagnosing latency regressions, tracking SLA compliance, spotting error or failure spikes, comparing multi-cluster performance, and analyzing queue/topic messaging health. All operations are read-only observability queries.
FAQ
What kind of queries does this skill generate?
Read-only DQL queries over Dynatrace service metrics and spans: response-time percentiles, error rates, throughput, failure spikes, degradation detection, messaging health, and service-mesh overhead, often grouped by service, endpoint, workload, namespace, or cluster.
Which runtimes and environments are supported?
It targets services on Java, .NET, Node.js, Python, PHP, and Go, and includes Kubernetes context queries by workload, namespace, and cluster.
When should I NOT use this skill?
The skill defers to sibling skills for other domains: use dt-obs-hosts for infrastructure metrics, dt-obs-logs for log analysis, and dt-obs-tracing for distributed tracing. It also isn't meant for explaining existing queries or product-documentation questions.
How is SLA compliance measured?
Via span-based queries that flag each root span as meeting an SLA (for example not failed and duration under a threshold) and then summarize compliant versus total requests per service into a compliance percentage.
Does it modify any systems?
No. It only reads telemetry through DQL timeseries and span queries; it performs no writes or destructive actions.
Install dt-obs-services
Quick Setup:
- Copy the skill folder to
.claude/skills/ - Claude will automatically detect and use the skill
Repository
dynatrace/dynatrace-for-ai