AGENT hub
Health, cost & trust
in one view.
Every agent scored on what it delivers, what it costs and whether it’s still calibrated.
Acceptance and MTTR delta, token economics down to cache hits, hallucination
and PII evaluators, drift thresholds — with recommendations to tune, demote or
pause.
◉Agent Perflive
ACCEPTANCEEFFECTIVENESS
89%
▲ 3 pts vs last week
MTTR DELTAEFFECTIVENESS
−39%
39m vs 1h 4m
P95 LATENCYOPERATIONAL
2.55s
2.4% errors
SPEND (MONTH)COST
$717
forecast $760
CALIBRATIONDRIFT
86%
1.66% hallucination
RUNS / DAYVOLUME
493
3,213 auto-sent
8 healthy 3 watch 1 degraded 2 guardrail violations · 5.0% approval reject
⊙Slow Endpoint Detective acceptance dropped 22 pointsOpen agent · review last 10 dismissed runs ›

Last week 83% → this week 61%. Coincides with the Datadog v7 ingest schema change.

⚠AWS Spend Anomaly calibration driftingOpen agent · refresh prompt examples ›

Predicted confidence vs observed acceptance gap widened from 4 → 14 points over 30 days.

⚠Pod CrashLoop Diagnoser p95 latency over budgetOpen agent · inspect run trace ›

p95 at 3.6s against a 3.0s budget. Most of the climb is in the RAG-retrieval stage.

⚖ LLM-AS-JUDGE EVALUATORSOpen all ›
HALL
1.66%
BIAS
0.76%
TOX
0.23%
PII
0.45%
SENT
0.73
every run scored · rubric v4 · 0.9s after completion
⬡ TOKEN ECONOMICS119.60M/mo · cache 40%
in62.12M
out10.06M
cache read41.41M
cache write6.21M
◈ AUTONOMY POSTURE
78%
active · 25 / 33 · 3,213 auto-sent
Disagreement rate 2.4%
Kill-switch activations 2
Audit completeness 100%
AUTONOMY

Sort any column · Pause trips that agent's kill switch

Cert Auditor
haiku · automate
Acceptance 99%
MTTR 2m
P95 1.2s
Spend/mo $96
Calibration 97%
AWS Investigation
opus · automate
Acceptance 74%
MTTR 18m
P95 4.1s
Spend/mo $864
Calibration 72%
Slow Endpoint Detective
opus · discover
Acceptance 61%
MTTR 26m
P95 3.9s
Spend/mo $598
Calibration 64%

Same task class, three agents — the cheapest one is also the best.

SPEND BY AGENT · MONTH
COST DRIVERS
Cost per resolved incident $1.46
Cache hit rate 40%
Saved by caching $318/mo
Forecast this month $760
◈Model right-sizing available

3 agents run opus on structured tasks. Estimated saving $412/mo with no measured quality loss.

FOUR VIEWS
One dashboard, four questions.
Each view answers a different one — and they all read from the same run data.
Is everything OK?
The whole estate at a glance — effectiveness, cost, drift and what needs attention today.
Which one is the problem?
Sort by acceptance, latency, tokens or spend to find the outlier fast.
Which one should we keep?
Put agents doing similar work side by side and let the numbers decide. time.
What is this costing?
Spend per agent and per resolved incident, cache savings and month-end forecast. time.
CALIBRATION & DRIFT DETECTION
Catch an agent going
stale.
Agents don’t fail loudly. They drift — staying confident while getting less right. Lumi tracks stated confidence against measured acceptance and flags the gap.

1

Confidence vs reality
A calibrated agent that says 80% should be right about 80% of the time.

2

Drift thresholds
Cross the band and it’s flagged, demoted or paused automatically.

3

Root the cause
Stale few-shot examples, a changed schema, a model swap — the dashboard names it.
AWS SPEND ANOMALY · CALIBRATION 30D⚠ gap 4 → 14 pts
stated confidence observed acceptance
stated confidenceobserved acceptancecalibrated band
AI-DRIVEN RECOMMENDATIONS
It tells you what to do about it.
Four active right now — each with the evidence and the exact setting to change.
⏻
Pause Slow Endpoint Detective — SLO compliance below 30%

Acceptance dropped 83% → 61% over 14 days, coinciding with the Datadog v7 schema change. 18 drafts dismissed last week alone.

Settings → Agent governance → pause Slow Endpoint Detective
⚠
Refresh AWS Spend Anomaly prompt examples

Calibration gap widened from 4 to 14 points over 30 days. Newer SKUs are missing from the few-shot set.

Open agent → Versions → promote v4 (drafted)
◎
Trim Pod CrashLoop Diagnoser RAG retrieval

p95 at 3.6s against a 3.0s budget, with most of the climb in the retrieval stage rather than generation.

Open agent → inspect run trace → retrieval depth
◈
Right-size 3 agents to a smaller model

All three run opus on highly structured tasks at 97%+ acceptance. Estimated $412/mo saving with no measured quality loss.

FinOps → model mix → apply recommendation
CAPABILITIES
Everything on the dashboard.
📈
Effectiveness metrics
Acceptance rate and MTTR delta against the human baseline.
$
Cost tracking & forecasting
Month-to-date spend per agent with an end-of-month projection.
◉
Health status overview
Healthy, watch or degraded — plus guardrail violations and reject rate. time.
◈
Autonomy posture
Percent active, disagreement rate, kill-switch activations, audit completeness.
⚖
LLM-as-judge evaluators
Hallucination, bias, toxicity, PII and sentiment scored on every run.
✦
AI recommendations
Proactive calls to pause, refresh, inspect or right-size — with the setting path.
⏻
Kill-switch controls
Stop any agent instantly; activations are counted and audited.
⬡
Token economics
In, out, cache read and cache write — with the cache hit rate that drives cost.
〰
Calibration & drift
Stated confidence tracked against observed acceptance, with alert thresholds.

WHY IT MATTERS

Autonomy you can defend in a
budget review.
Running agents without this is paying for trust you can’t evidence. Here the spend
becomes a line you can justify — or cut.
AI SPEND VISIBILITY
one invoice
$1.46 / incident
Cost attributed to the work it actually closed.
TOKEN WASTE
invisible
40% cached
Cache hits and model right-sizing are the fastest savings.
TIME TO CATCH DRIFT
a bad quarter
days
Calibration alerts fire before users lose confidence.
QUALITY REVIEW
spot-checks
every run scored
Continuous evaluation instead of quarterly sampling.
👤
For the ITOps team
DAY TO DAY
✓
Know which agents to trust
Acceptance and calibration tell you when to take the recommendation and when to check the work.
✓
Kill it in one click
An agent misbehaving at 2am is a button, not an incident bridge.
✓
No silent degradation
A schema change upstream shows up as a drift alert, not as angry stakeholders.
✓
Evidence for promotion
Move an agent to automate because the numbers earned it.
◈
For the organization
ON THE BALANCE SHEET
✓
AI spend stops being a black box
Cost per resolved incident is a metric finance recognises — and can hold you to.
✓
Cut waste without cutting capability
Caching and right-sizing trim spend while acceptance holds.
✓
Governable autonomy
Kill switches, disagreement rates and 100% audit completeness are what let risk teams approve automation.
✓
Scale on evidence
Expand automation where the data justifies it, not on a hunch.
Dashboard figures are illustrative. Results depend on your model mix, task profile and current autonomy posture.
Get started
See what your
agents actually cost
Connect a week of runs and we’ll show you the health, spend and drift picture for your estate.
Autonomous ITOps platform powered by an AI Coworker, SRE Orchestrator, and Agent Builder
© 2026 – 2027 LumiOps.AI. All rights reserved.

© 2026 – 2027 LumiOps.AI. All rights reserved.