Skip to main content

Observability

CPU Observability is the operational view for clusters activated under the API Router Platform. It provides a real-time dashboard that helps you monitor CPU usage, API request volume, latency, error rates, and other performance metrics of your reseller infrastructure.

Before you begin​

  • The cluster must be in the Activated state.
  • The API Router Platform is fully deployed.

Open CPU Observability​

  1. In the Console, open Activate Cluster.
  2. Find a Router-mode activation with the Activated status.
  3. In Actions, select the CPU Observability icon.

Control the view​

  • Backend / Gateway — switch the scope between the reseller backend service and the gateway.
  • Time window — choose 15m, 1h, 6h, 24h, or 7d. Longer windows use coarser sampling steps so charts stay readable.
  • Auto refresh — set to Off, 30s, 1m, or 5m. The system automatically extends the interval when the window's sampling step is slower, to avoid redundant refetches. The effective interval is shown in the selector.
  • Refresh now — reload all panels immediately.
  • Refreshed N ago — shows when data was last fetched.

Read the health strip​

The health strip summarizes key indicators as cards. Each card shows the current value, the data-collection time, and a warning hint when a threshold is crossed. The page shows Low traffic for metrics with no traffic in the window.

Backend scope​

IndicatorWarns when
ProcessProcess is down.
Heap after GCHeap usage after GC exceeds 70%.
GC pause p99Pause exceeds 1 second.
GC CPU overheadGarbage collection exceeds 5% of CPU.
File descriptorsDescriptor usage exceeds 80%.
Request rateNo threshold; watch for anomalies.
Error rateError ratio exceeds 5%.
Direct memoryNo threshold; watch for growth.
Live threadsNo threshold; watch for growth.
Process CPUCPU usage exceeds 85%.
ERROR log rateMore than one ERROR log event per second.
DB pool usageConnection pool usage exceeds 80%.
Redis pool waitersAny thread is waiting for a Redis connection.
Redis worst borrow waitNo threshold; watch for growth.
UptimeProcess uptime since start.

Backend

Gateway scope​

IndicatorWarns when
Metrics agentThe agent is down.
ScrapeThe gateway scrape target is down.
Request rateNo threshold; watch for anomalies.
Upstream error rateUpstream error ratio exceeds 5%.
Timeout rateTimeout ratio exceeds 10%.
503 rate503 response rate exceeds 10%.
Live workersNo live workers.
Data ageSnapshot is older than 60 seconds.

Gateway

Read time-series charts​

The page groups charts by concern. Hover over a chart for exact values; each chart includes a hint that explains how to read it. Typical groups include heap and non-heap memory by pool, GC pause and allocation rates, request rate and latency by domain, top endpoints by request rate, error rates by status code, thread pools, database pool usage, and gateway upstream metrics.

Use the time window selector to align all charts with the period you are investigating.

Run diagnostics​

Select Diagnostics to open a drawer with live Spring Boot actuator data for the backend:

ViewWhat it shows
HealthApplication health endpoints and their status.
EnvEnvironment properties with effective and configured values.
MappingsHTTP request mappings registered by the backend.
Scheduled TasksScheduled jobs and their last execution.
ThreadsLive thread list and stack traces.
LoggersLoggers with their levels, searchable by name.

Diagnostics

Troubleshooting​

The entry is not available​

The CPU Observability entry appears only for Activated Router-mode clusters. For Activating records, select Manage and complete the activation first. When the cluster is disconnected, the entry is disabled until connectivity is restored.

A card shows Fetch failed​

The panel could not load its metric source. The page separates fetch failures from genuinely missing data, so a failed card does not zero out the rest of the strip. Select Refresh now to retry.

Data looks stale​

Check the Data age indicator on the gateway scope and the Refreshed N ago label. If the page was in a background tab, auto refresh was paused; returning to the tab triggers an immediate refresh.

Next Steps​

Activate Cluster

Review activation records, connection status, and service management.

Review Usage

Trace reseller traffic in the API-based (Router) and History views.