All templates

Automated Rancher Cluster Resource Threshold Alerts

Every few minutes, WebRun checks Rancher for nodes whose CPU or memory usage crosses your configured threshold, posts a Slack alert naming the cluster and node, and sends a Telegram escalation if the same node is still over threshold 30 minutes later, so real capacity problems get pushed further than routine spikes.

Runs on WebRun · Strict Lockdown policy
Checks every 5 minutes, day and night WebRunorchestrates each step
1 Rancher check cluster resource usage
2 Slack post the threshold alert
3 Telegram escalate if it persists
In short

How do I get alerted when a Rancher cluster is running low on resources?

WebRun checks Rancher every few minutes for nodes crossing your CPU or memory threshold, posting a Slack alert with the cluster and node the moment it happens. If the same node is still over threshold 30 minutes later, it sends a Telegram escalation, so a brief spike stays routine while a real capacity problem gets pushed further.

  • A resource threshold breach gets a Slack alert within minutes
  • Sustained pressure gets escalated to Telegram automatically
  • Brief spikes stay routine instead of paging the whole team

Built for Platform engineering teams · Kubernetes administrators · infrastructure teams · site reliability engineers

Step by step

What does WebRun do on every run?

The exact actions WebRun takes, in order - in plain language, so you can adjust anything.

  1. WebRun signs in and gets to work

    Opens rancher.com in a real browser with your saved login - no setup, no API keys.

  2. 1
    Rancher - check cluster resource usage
    • Open Rancher and review CPU and memory usage across nodes in managed clusters
    • Compare each node's usage against your configured threshold
    • Flag any node currently over threshold and note how long it has been over

    Done when Every node currently over its resource threshold is identified with its cluster and duration.

  3. 2
    Slack - post the threshold alert
    slack.com
    WebRun in Slack: post the threshold alert
    WebRun opens Slack to post the threshold alert.
    • Post an alert to the infrastructure Slack channel for each new threshold breach
    • Include the cluster, node, and current CPU or memory usage
    • Note whether this is a new breach or one still ongoing

    Done when The team has a Slack alert for every current resource threshold breach.

  4. 3
    Telegram - escalate if it persists
    telegram.org
    WebRun in Telegram: escalate if it persists
    WebRun opens Telegram to escalate if it persists.
    • Send a Telegram escalation for any node still over threshold after 30 minutes
    • Include the cluster, node, and total time over threshold
    • Mark it clearly as an escalation

    Done when Any breach still ongoing after 30 minutes has been escalated on Telegram.

Run settings

How is each run configured?

Starting pageWhere Chrome opens at the start of each run
rancher.com
ScheduleRuns automatically on this cadence
Checks every 5 minutes, day and night
DeliveryHow each run's result reaches you
Resource threshold alert · Slack + Telegram
OutputWhat each run produces - A Slack alert for every new resource threshold breach and a Telegram escalation for anything lasting 30 minutes.
Alert
Setup & safety

Secure by default

Connect once, stays signed in

WebRun signs in once and keeps each session in a persistent environment, so every run picks up right where it left off.

Your credentials stay in your own private environment - WebRun never stores your passwords.
Strict Lockdown

Every action is checked against this policy before it runs.

Domains ALLOWLIST
Typed input ALLOW
Shell command BLOCK
File uploads BLOCK
Runs in a contained environment More on policies
Good to know

Questions, answered

Will WebRun scale the cluster or evict workloads itself?

No. It only reports the threshold breach. Scaling a cluster, adding nodes, or evicting workloads stays a decision made by your infrastructure team in Rancher.

Where do the thresholds come from?

From the CPU and memory limits your team configures for WebRun to watch, so the alerts match what your infrastructure team already considers a problem.

Why wait 30 minutes before escalating?

So a brief resource spike from a normal workload burst does not trigger an unnecessary Telegram escalation. Sustained pressure for 30 minutes usually means a real capacity issue.

Put this on autopilot.

Turn it on in minutes - or have our team set it up for you.