Skip to content
Technology agents
Enterprise OperationsTechnologyPlatform & Site Reliability (SRE)

Capacity & Cost Agent

Right-sizes infrastructure for performance and spend, continuously.

Watches utilization, forecasts demand, and applies scaling and instance-mix changes within bounds. It backtests each change against historical load and acts on its own forecasts, holding headroom for peak while trimming idle capacity.

Authority

Act within policy

Team role

Provides specialist analysis

Handoffs

Named collaborators

The role

What it owns and where its authority ends

Desk

Platform & Site Reliability (SRE)

Desk workflow

Signal, then alert, then agentic triage, then diagnosis, remediation and post-mortem. Known incidents run runbook autopilot (restart, scale, roll back); novel outages route to a responder agent that reasons them end to end.

Collaboration

Works within a defined desk workflow

Decision boundary

Acts only within an explicit policy, permission and escalation boundary.

Systems and capabilities involved

  • Cloud billing + usage API

  • Autoscaler config

    apply within bounds

  • Demand-forecast sandbox

Handoffs

What this role gives and receives

Capabilities offered

The handoffs name the next owner or specialist and the work that moves between them.

Context

What the role needs to do the work

Current work
Current utilization snapshot and the change under evaluation.
Prior interactions
Past scaling events and their cost/performance outcomes.
Policies and reference
Service SLOs, instance pricing, reservation commitments.
Working method
Not specified for this role.

Checks and boundaries

What must be tested or reviewed

  1. 01Changes above a spend/risk threshold require an oversight-agent gate before commit.
  2. 02Backtest every right-sizing change against historical load before applying.
  3. 03SLO-protection guardrail: never scale below the headroom needed for peak.

Human authority

Acts only within an explicit policy, permission and escalation boundary.

Keep exploring