FeaturedDev & IT Ops#AI Observability#Agent SLO#Error Budget#Tool Calls#OpenTelemetry#Reliability

AI Agent Trace SLO Burn Alert

Calculate an agent tool-call error budget from trace data and alert on fast SLO burn without sending raw prompts.

Workflow at a glance

The full canvas, before you import it

Click any node to see its config.

#AI Observability#Agent SLO#Error Budget#Tool Calls#OpenTelemetry#Reliability

Click a node to select it — same as the Heym editor; the panel shows its settings.

8 nodes · Free & source-available

AI Agent Trace SLO Burn Alert

Monitor operational reliability for an AI agent with an error-budget calculation instead of a single noisy threshold. The workflow reads bounded trace metadata, calculates tool-call failure and latency burn, and alerts only when the recent window threatens the configured SLO.

What this workflow does

  1. AgentSloSchedule runs every ten minutes
  2. FetchAgentTraceWindow retrieves trace metadata from your observability endpoint
  3. CalculateAgentSloBurn computes availability, latency, and error-budget burn
  4. FastBurnGate detects a fast-burn condition
  5. PostAgentSloAlert publishes a concise Discord alert
  6. LogAgentSloWindow records every evaluated window
  7. Alert and healthy paths return separate receipts

Use cases

  • AI agent observability
  • Tool-call error budget monitoring
  • LLM workflow SLO alerts
  • Agent latency and reliability operations

Setup

Point the HTTP node at an endpoint that returns bounded trace metadata without raw prompts or tool arguments. Adjust the 99 percent availability target and 2.0 burn threshold in the code. Connect Discord and DataTable.

How to import this template

  1. 1Click Import → Copy JSON on this page.
  2. 2Open your Heym and navigate to a workflow canvas.
  3. 3PressCmd+V/Ctrl+V— nodes appear instantly.
  4. 4Add your API keys in the node config panels and click Run.
More workflow templates
View all templates
Heym
incident analysis · production AI
Observed across 100s of AI rollouts

AI workflows don't fail because of prompts.
They fail because of orchestration.

symptom · glue code01
5 tools
Scripts, vector DB, approval bot, tracing, browser runner — none of them talk.
symptom · visibility02
~0%
Observable behavior across the stack. Debugging is guesswork.
with heym · one runtime
1 canvas
Agents, RAG, HITL, MCP, traces & evals. Self-hosted. Observable.
AI-Native RuntimeProduction-Grade
github.com/heymrun/heym