Part of Foresite’s From Prototype to Production series, exploring building multi-agent systems in Google Cloud SecOps. |
Why MTTT fails autonomous agents, how Weight of Evidence transforms triage calibration, and the engineering of a dual-loop SOC.
In Part 1 we tackled the metric nobody in AI security likes to publish: the abstention rate, or how often an autonomous agent has the discipline to say "I don’t know." We echoed a warning from Google Cloud's Chris Martin:
"The moment you cross the autonomy threshold, you trade a human-in-the-loop for a metric-in-the-loop."
When engineering leaders first hear this, the instinct is to reach for traditional security operations metrics: Mean Time to Triage (MTTT), Mean Time to Investigate (MTTI), or autonomous closure rate. That instinct is a trap.
In a human SOC, cutting MTTI from 45 minutes to 15 is a genuine operational triumph. In an agentic SOC, MTTT is the easiest metric in the world to game. An AI agent can ingest a complex, multi-stage alert and hallucinate a confident, beautifully formatted verdict in 3.2 seconds. The dashboard lights up green: MTTT has plummeted 90%! The catch is that the model just stamped a stealthy credential-dumping campaign as BENIGN because no known signature was present in the endpoint cache.
Speed without calibrated certainty is reckless. Optimize for MTTT without a mathematical floor and you have not automated triage. You have automated alert abandonment.
So we rethought the relationship between metrics, models, and human analysts. We did not swap the human-in-the-loop for a metric-in-the-loop. We built a Dual-Loop Architecture, where metrics guide human evaluation and human evaluation continuously recalibrates the metrics.
Why not just ask the AI agent: "How confident are you in this verdict, 0 to 1?"
We tried. Early in our engineering prototypes we allowed specialist agents self-report confidence alongside their verdicts. The empirical reality was brutal:
| Core principle: a model must never grade its own homework. Large language models excel at semantic synthesis, narrative reconstruction, and contextual pattern recognition. They are poor at objective probability estimation. |
To solve the overconfidence trap we made an architectural separation. AI agents produce structured semantic observations. Deterministic mathematics produces the score.
At the heart of our Stage-2 aggregation engine is a Weight of Evidence (WoE) framework.
Weight of Evidence: reasoning decoupled from scoring
The underlying formulation is proprietary Foresite intellectual property, but the operational concept is straightforward and mathematically rigorous:
τ = ln(CFP / CFN)
where CFP and CFN are the cost of a false positive and a false negative. Raise the cost of a miss and the threshold moves to catch more.
This matters more in a managed practice than in a single enterprise. Foresite runs security operations for hundreds of organizations across healthcare, technology, financial services and public sector, each on their own Google SecOps instance. A missed intrusion at a hospital and a missed intrusion at a software company carry different costs, and those customers are not served by one global confidence cutoff. The threshold is set per tenant, against that customer's risk posture, and the same WoE framework produces defensible verdicts under all of them.
With scoring stripped away from the model and placed in a calibrated WoE framework, verdicts become reproducible, explainable, and provably bounded.
Enterprise reliability does not come from removing humans. It comes from orchestrating how human judgment and statistical models interact.
The dual loop
When an alert reaches an analyst, the metric does not hand over a raw score. It provides contextual navigation:
When an analyst makes a determination on an escalated case, that decision does not vanish into an archived ticket. The analyst's judgment, together with the raw payload and the intermediate agent outputs, is captured as a structured event. Those events feed the evaluation pipeline:
Deploy a change to a traditional microservice and CI/CD runs unit and integration tests. Green, ship.
In an agentic system, every prompt tweak, model update, or framework patch is an architectural risk. A change intended to fix a phishing edge case can quietly degrade the identity compromise specialist.
To govern this, we engineered a two-tier continuous validation bench.
Zero credentials, zero model calls, instant. When we modify aggregation logic, voter weights, or severity thresholds, we do not spend thousands of dollars re-running inference. We maintain a frozen ledger of hundreds of SOC-verified production incidents, and a deterministic replay harness re-runs the mathematical aggregation engine over the historical vote arrays in milliseconds.
If a mathematical change recovers four false negatives but introduces even a single new false positive across 500 held-out cases, the build fails automatically.
For reforms touching specialist instructions, ADK agent behaviors, or prompt recognizers, deterministic replay is not enough. We run live end-to-end evaluations on true held-out data:
In software engineering you monitor CPU, memory, and HTTP 500 errors. In AI engineering, an agent can be failing catastrophically while returning HTTP 200 with zero latency.
We monitor three vital signs to detect production drift before it ever impacts an analyst:
Three drift vital signs
strict → repaired → partial → salvaged → empty. When the share of repaired or salvaged outputs starts rising, the upstream foundation model has drifted in its formatting habits, and we know before anything crashes.stop_reason, native output tokens, and thinking budgets on every call. We do not guess whether a specialist timed out or ran out of token headroom. The telemetry explicitly documents it.If you are evaluating AI triage capability, change the questions you ask your vendors and your own teams.
| Stop asking | Start asking |
|---|---|
| What is your agent's Mean Time to Triage? | How do you stop the agent optimizing for speed at the expense of false negatives? |
| How confident is the AI in its decision? | Is confidence self-reported by the model, or derived through a decoupled, calibrated evidence framework? |
| Do you have a human in the loop? | Does your human-in-the-loop actually close the feedback loop and recalibrate the production models? |
In a managed service the damage compounds. A hallucinated
BENIGNdoes not just cost one analyst's trust in one tool. It costs a customer's trust in their security provider, and that is far harder to win back than a tuning cycle.
Replace unconstrained model confidence with a disciplined Weight of Evidence framework and generative AI stops being an unpredictable black box. It becomes a defensible, auditable force multiplier for the SOC.
Now that we’ve examined how mathematical calibration and the dual loop govern verdict accuracy, how do you handle the actual reasoning?
In Part 3: "Why We Don't Let the Model Score Itself: Two-Stage Architecture & The Recursive Count-Ratio Identity," we’ll dive deep into our ADK pipeline design, exploring how we separate per-entity investigation from alert synthesis, and how a single recursive identity runs across eight distinct operational scales.
Stay tuned.
How do you know an AI triage agent is still performing correctly? Through continuous evaluation against a live holdout set, not a one-off accuracy figure. Foresite runs two tiers: deterministic replay of SOC-verified historical cases in milliseconds, and live end-to-end runs with multiple stochastic draws per case to test verdict stability. Evaluation detects drift. It does not establish accountability, which is why named practitioners remain in the loop.
Why is Mean Time to Triage a bad metric for an agentic SOC? Because a model can produce a confident, well-formatted, wrong verdict in seconds. MTTT rewards exactly that. A system optimizing for triage speed without a calibrated mathematical floor has not automated triage, it has automated alert abandonment.
Can a language model accurately report its own confidence? No. In Foresite's testing, self-reported confidence on identical cases swung between 0.20 and 0.85 on minor prompt differences, and agents inflated confidence when threat intelligence returned empty results. Models are strong at semantic synthesis and weak at probability estimation, so scoring belongs in a separate deterministic framework.
What is agent drift and how do you detect it? Drift is degradation that returns no error. The agent answers successfully while its quality falls. Three signals catch it: a rising share of repaired or salvaged output parses, execution telemetry showing timeouts or token exhaustion, and a falling abstention rate, which indicates the agent has started guessing rather than escalating.
What should a CISO ask a vendor about AI triage evaluation? Three questions. How do you stop the agent trading false negatives for speed? Is confidence self-reported or independently calibrated? And does your human-in-the-loop actually feed back into model recalibration, or is it just a review step?