Capacity & Cost Agent
Right-sizes infrastructure for performance and spend, continuously.
Watches utilization, forecasts demand, and applies scaling and instance-mix changes within bounds. It backtests each change against historical load and acts on its own forecasts, holding headroom for peak while trimming idle capacity.
Authority
Act within policy
Team role
Provides specialist analysis
Handoffs
Named collaborators
The role
What it owns and where its authority ends
Desk
Platform & Site Reliability (SRE)
Desk workflow
Signal, then alert, then agentic triage, then diagnosis, remediation and post-mortem. Known incidents run runbook autopilot (restart, scale, roll back); novel outages route to a responder agent that reasons them end to end.
Collaboration
Works within a defined desk workflow
Decision boundary
Acts only within an explicit policy, permission and escalation boundary.
Systems and capabilities involved
Cloud billing + usage API
Autoscaler config
apply within bounds
Demand-forecast sandbox
Handoffs
What this role gives and receives
Capabilities offered
The handoffs name the next owner or specialist and the work that moves between them.
Receives from
Context
What the role needs to do the work
- Current work
- Current utilization snapshot and the change under evaluation.
- Prior interactions
- Past scaling events and their cost/performance outcomes.
- Policies and reference
- Service SLOs, instance pricing, reservation commitments.
- Working method
- Not specified for this role.
Checks and boundaries
What must be tested or reviewed
- 01Changes above a spend/risk threshold require an oversight-agent gate before commit.
- 02Backtest every right-sizing change against historical load before applying.
- 03SLO-protection guardrail: never scale below the headroom needed for peak.
Human authority
Acts only within an explicit policy, permission and escalation boundary.
Keep exploring