FeaturedFinance Ops#GPU FinOps#Inference Cost#Cost per Token#LLM Operations#AI Infrastructure#BigQuery

GPU Inference Cost per Token Anomaly Monitor

Calculate GPU inference cost per million tokens, detect efficiency regressions, and send a FinOps investigation brief.

Workflow at a glance

The full canvas, before you import it

Click any node to see its config.

#GPU FinOps#Inference Cost#Cost per Token#LLM Operations#AI Infrastructure#BigQuery

Click a node to select it — same as the Heym editor; the panel shows its settings.

9 nodes · Free & source-available

GPU Inference Cost per Token Anomaly Monitor

Watch AI infrastructure efficiency as traffic, models, and GPU prices change. The workflow queries hourly usage and cost data, calculates cost per million tokens against a baseline, and explains only material regressions.

What this workflow does

  1. GpuFinOpsSchedule starts the check every four hours
  2. QueryGpuInferenceUsage reads cost, token, model, and GPU-hour data from BigQuery
  3. CalculateInferenceUnitCost derives current and baseline unit economics
  4. GpuCostAnomalyGate detects regressions above the configured threshold
  5. ExplainGpuCostAnomaly identifies likely demand, model, batching, and utilization drivers
  6. EmailAiFinOpsOwners sends the investigation brief
  7. AppendGpuCostHistory records the observation in Google Sheets
  8. Separate receipts cover alert and normal states

Use cases

  • GPU FinOps monitoring
  • LLM inference cost per token tracking
  • AI infrastructure unit economics
  • Model serving efficiency alerts

Setup

Connect BigQuery, an LLM credential, email, and Google Sheets. Update the table and field names in the query to match your billing export and inference telemetry. Validate the cost allocation model with finance and platform teams before using the metric for chargeback.

How to import this template

  1. 1Click Import → Copy JSON on this page.
  2. 2Open your Heym and navigate to a workflow canvas.
  3. 3PressCmd+V/Ctrl+V— nodes appear instantly.
  4. 4Add your API keys in the node config panels and click Run.
More workflow templates
View all templates
Heym
incident analysis · production AI
Observed across 100s of AI rollouts

AI workflows don't fail because of prompts.
They fail because of orchestration.

symptom · glue code01
5 tools
Scripts, vector DB, approval bot, tracing, browser runner — none of them talk.
symptom · visibility02
~0%
Observable behavior across the stack. Debugging is guesswork.
with heym · one runtime
1 canvas
Agents, RAG, HITL, MCP, traces & evals. Self-hosted. Observable.
AI-Native RuntimeProduction-Grade
github.com/heymrun/heym