Security Blog

Metric-in-the-Loop Without Giving Up Human-in-the-Loop

Written by Shishir Suresh | October 9, 2026
FORESITE EDITORIAL SERIES

Part of Foresite’s From Prototype to Production series, exploring building multi-agent systems in Google Cloud SecOps.

Why MTTT fails autonomous agents, how Weight of Evidence transforms triage calibration, and the engineering of a dual-loop SOC.

In Part 1 we tackled the metric nobody in AI security likes to publish: the abstention rate, or how often an autonomous agent has the discipline to say "I don’t know." We echoed a warning from Google Cloud's Chris Martin:

"The moment you cross the autonomy threshold, you trade a human-in-the-loop for a metric-in-the-loop."

When engineering leaders first hear this, the instinct is to reach for traditional security operations metrics: Mean Time to Triage (MTTT), Mean Time to Investigate (MTTI), or autonomous closure rate. That instinct is a trap.

In a human SOC, cutting MTTI from 45 minutes to 15 is a genuine operational triumph. In an agentic SOC, MTTT is the easiest metric in the world to game. An AI agent can ingest a complex, multi-stage alert and hallucinate a confident, beautifully formatted verdict in 3.2 seconds. The dashboard lights up green: MTTT has plummeted 90%! The catch is that the model just stamped a stealthy credential-dumping campaign as BENIGN because no known signature was present in the endpoint cache.

Speed without calibrated certainty is reckless. Optimize for MTTT without a mathematical floor and you have not automated triage. You have automated alert abandonment.

So we rethought the relationship between metrics, models, and human analysts. We did not swap the human-in-the-loop for a metric-in-the-loop. We built a Dual-Loop Architecture, where metrics guide human evaluation and human evaluation continuously recalibrates the metrics.

 

The fatal flaw of self-scoring agents

Why not just ask the AI agent: "How confident are you in this verdict, 0 to 1?"

We tried. Early in our engineering prototypes we allowed specialist agents self-report confidence alongside their verdicts. The empirical reality was brutal:

  • Confidence flux. Re-running the same captured case produced identical technical observations, but self-reported confidence swung between 0.20 and 0.85 on minor prompt tokenization differences.
  • The "Absence as Evidence" Fallacy. When threat intelligence returned empty, (no malicious hits on VirusTotal or Mandiant), the agent routinely inflated its confidence that the entity was safe. Silence was mistaken for innocence.
  • Cognitive Dissonance. Asked to vote, specialist agents fell into sycophancy, agreeing with each other's conclusions for conflicting and incompatible reasons.


Core principle: a model must never grade its own homework. Large language models excel at semantic synthesis, narrative reconstruction, and contextual pattern recognition. They are poor at objective probability estimation.

 

Weight of Evidence: decoupling reasoning from scoring

To solve the overconfidence trap we made an architectural separation. AI agents produce structured semantic observations. Deterministic mathematics produces the score.

At the heart of our Stage-2 aggregation engine is a Weight of Evidence (WoE) framework.


Weight of Evidence: reasoning decoupled from scoring

The underlying formulation is proprietary Foresite intellectual property, but the operational concept is straightforward and mathematically rigorous:

  1. Information-theoretic calibration. Instead of treating every agent's opinion equally or using arbitrary scalar weights (e.g., $w=0.5$), our framework evaluates the actual empirical information value contributed by distinct categories of evidence.
  2. Monotone constraints. By construction, benign evidence cannot arbitrarily mask or suppress a distinct threat signal.
  3. Earned-clean enforcement. If an alert shows zero malicious IOCs because the endpoint is offline or the indicator is zero-day, the framework treats that as zero information, not as negative evidence. The system scores at the global prior, which is what prevents the false negative.
  4. Bayes minimum-cost thresholding. Instead of a hardcoded confidence cutoff, the decision threshold 𝞃 is tied to the organization's risk posture and the operational cost ratio between false positives and false negatives.

τ = ln(CFP / CFN)

where CFP and CFN are the cost of a false positive and a false negative. Raise the cost of a miss and the threshold moves to catch more.

This matters more in a managed practice than in a single enterprise. Foresite runs security operations for hundreds of organizations across healthcare, technology, financial services and public sector, each on their own Google SecOps instance. A missed intrusion at a hospital and a missed intrusion at a software company carry different costs, and those customers are not served by one global confidence cutoff. The threshold is set per tenant, against that customer's risk posture, and the same WoE framework produces defensible verdicts under all of them.

With scoring stripped away from the model and placed in a calibrated WoE framework, verdicts become reproducible, explainable, and provably bounded.

 

The dual loop: how metrics and humans actually coexist

Enterprise reliability does not come from removing humans. It comes from orchestrating how human judgment and statistical models interact.


The dual loop

Loop 1: metric-in-the-loop → informs the human reviewer

When an alert reaches an analyst, the metric does not hand over a raw score. It provides contextual navigation:

  • Which evidence dimensions carried the highest Weight of Evidence.
  • Exactly which telemetry was missing, for example "endpoint offline; no process ancestry available"
  • Precisely where the model hit its reasoning frontier, turning a 30-minute forensic hunt into a 2-minute targeted verification.


Loop 2: human-in-the-loop → auto-regressively optimizes the metrics

When an analyst makes a determination on an escalated case, that decision does not vanish into an archived ticket. The analyst's judgment, together with the raw payload and the intermediate agent outputs, is captured as a structured event. Those events feed the evaluation pipeline:

  1. They populate our Live Holdout Ledger of real-world captures.
  2. They re-fit our WoE tables, adjusting the weight of evolving threat techniques.
  3. They become permanent fixtures in our Continuous Regression Benches, so the system never repeats a solved mistake.


Engineering the continuous regression bench

Deploy a change to a traditional microservice and CI/CD runs unit and integration tests. Green, ship.

In an agentic system, every prompt tweak, model update, or framework patch is an architectural risk. A change intended to fix a phishing edge case can quietly degrade the identity compromise specialist.

To govern this, we engineered a two-tier continuous validation bench.

Tier 1: deterministic replay

Zero credentials, zero model calls, instant. When we modify aggregation logic, voter weights, or severity thresholds, we do not spend thousands of dollars re-running inference. We maintain a frozen ledger of hundreds of SOC-verified production incidents, and a deterministic replay harness re-runs the mathematical aggregation engine over the historical vote arrays in milliseconds.

If a mathematical change recovers four false negatives but introduces even a single new false positive across 500 held-out cases, the build fails automatically.

Tier 2: the end-to-end live harness

For reforms touching specialist instructions, ADK agent behaviors, or prompt recognizers, deterministic replay is not enough. We run live end-to-end evaluations on true held-out data:

  • Fresh re-execution. We discard intermediate state and run full multi-stage triage from raw JSON to final analyst report.
  • Distributional verification. Because AI agents are non-deterministic, a single pass proves nothing. Our harness runs multiple stochastic draws per case and evaluates the stability of the verdict distribution.
  • Seam validation. We inspect every integration seam, ensuring the typed Pydantic contracts between Stage 1 and Stage 2 remain intact.


Monitoring production agent drift

In software engineering you monitor CPU, memory, and HTTP 500 errors. In AI engineering, an agent can be failing catastrophically while returning HTTP 200 with zero latency.

We monitor three vital signs to detect production drift before it ever impacts an analyst:


Three drift vital signs

  1. Parse-level gradients. Every agent output is classified along a spectrum: strict → repaired → partial → salvaged → empty. When the share of repaired or salvaged outputs starts rising, the upstream foundation model has drifted in its formatting habits, and we know before anything crashes.
  2. Execution telemetry. We capture stop_reason, native output tokens, and thinking budgets on every call. We do not guess whether a specialist timed out or ran out of token headroom. The telemetry explicitly documents it.
  3. Abstention velocity. If the production abstention rate suddenly drops from its healthy ~12% baseline to 2%, we do not celebrate. A collapsing abstention rate is the diagnostic signature of an agent that has started to over-generalize and guess.


What this means for security leadership

If you are evaluating AI triage capability, change the questions you ask your vendors and your own teams.

Stop asking Start asking
What is your agent's Mean Time to Triage? How do you stop the agent optimizing for speed at the expense of false negatives?
How confident is the AI in its decision? Is confidence self-reported by the model, or derived through a decoupled, calibrated evidence framework?
Do you have a human in the loop? Does your human-in-the-loop actually close the feedback loop and recalibrate the production models?

 

In a managed service the damage compounds. A hallucinated BENIGN does not just cost one analyst's trust in one tool. It costs a customer's trust in their security provider, and that is far harder to win back than a tuning cycle.

Replace unconstrained model confidence with a disciplined Weight of Evidence framework and generative AI stops being an unpredictable black box. It becomes a defensible, auditable force multiplier for the SOC.

 

Up next in Part 3

Now that we’ve examined how mathematical calibration and the dual loop govern verdict accuracy, how do you handle the actual reasoning?

In Part 3: "Why We Don't Let the Model Score Itself: Two-Stage Architecture & The Recursive Count-Ratio Identity," we’ll dive deep into our ADK pipeline design, exploring how we separate per-entity investigation from alert synthesis, and how a single recursive identity runs across eight distinct operational scales.

Stay tuned.

 

FAQ

How do you know an AI triage agent is still performing correctly? Through continuous evaluation against a live holdout set, not a one-off accuracy figure. Foresite runs two tiers: deterministic replay of SOC-verified historical cases in milliseconds, and live end-to-end runs with multiple stochastic draws per case to test verdict stability. Evaluation detects drift. It does not establish accountability, which is why named practitioners remain in the loop.

Why is Mean Time to Triage a bad metric for an agentic SOC? Because a model can produce a confident, well-formatted, wrong verdict in seconds. MTTT rewards exactly that. A system optimizing for triage speed without a calibrated mathematical floor has not automated triage, it has automated alert abandonment.

Can a language model accurately report its own confidence? No. In Foresite's testing, self-reported confidence on identical cases swung between 0.20 and 0.85 on minor prompt differences, and agents inflated confidence when threat intelligence returned empty results. Models are strong at semantic synthesis and weak at probability estimation, so scoring belongs in a separate deterministic framework.

What is agent drift and how do you detect it? Drift is degradation that returns no error. The agent answers successfully while its quality falls. Three signals catch it: a rising share of repaired or salvaged output parses, execution telemetry showing timeouts or token exhaustion, and a falling abstention rate, which indicates the agent has started guessing rather than escalating.

What should a CISO ask a vendor about AI triage evaluation? Three questions. How do you stop the agent trading false negatives for speed? Is confidence self-reported or independently calibrated? And does your human-in-the-loop actually feed back into model recalibration, or is it just a review step?