# Pilot Evidence Report

Illustrative example based on a hypothetical pilot. Every count, date, identity and outcome below is sample data, not a real customer result.

## Evaluation details

- Organization: hypothetical five-developer B2B SaaS engineering team (not a named company).
- Evaluation window: 2026-08-10 through 2026-08-23 inclusive, UTC; 14 days.
- Registered agents: 5 synthetic identities, sample-agent-01 through sample-agent-05.
- Reviewed calls: 1240; one initial authorization decision per unique synthetic request.
- Workflow owner: platform engineer (scenario role). Reviewer: security engineer (scenario role).
- Deployment: isolated HTTP gateway and local stdio fixture; teaching model v1.
- Success criteria: every sample request has a decision, blocked requests do not execute, and approval outcomes reconcile.

## Workflow summary

A coding agent reads repository context, proposes database changes, requests refunds and opens an internal review application. All tools use test data. Production database writes are denied, consistent with the restricted-production pack. Risk scores remain illustrative.

## Initial decision distribution

| Decision | Calls | Share |
|---|---:|---:|
| ALLOW | 1080 | 87.10% |
| DENY | 60 | 4.84% |
| APPROVAL_REQUIRED | 100 | 8.06% |
| Total | 1240 | 100.00% |

## High-risk tool-call types

| Type | Calls | Risk / initial outcome | Sample evidence range |
|---|---:|---|---|
| Destructive shell execution | 60 | 100 / DENY | sample-shell-001–060 |
| Sensitive read or export | 70 | 70 / APPROVAL_REQUIRED | sample-db-001–070 |
| Refund request | 20 | 70 / APPROVAL_REQUIRED | sample-refund-001–020 |
| Internal browser access | 10 | 60 / APPROVAL_REQUIRED | sample-browser-001–010 |

High-risk total = 60 + 70 + 20 + 10 = 160; 12.90% of all calls. These evidence identifiers identify hypothetical records, not downloadable customer evidence.

## Applied teaching policies

| Tool | Policy | Calculated score | Decision |
|---|---|---:|---|
| shell.exec | destructive-commands-denied | 100 | DENY |
| postgres.query | restricted-production | 80 | DENY |
| github.read_file | read-only-repository-access | 10 | ALLOW |
| stripe.refund | financial-actions-require-approval | 70 | APPROVAL_REQUIRED |
| browser.navigate | internal-app-access-requires-approval | 60 | APPROVAL_REQUIRED |

## Approval actions and outcomes

| Outcome | Requests | Share of approval requests |
|---|---:|---:|
| Approved | 70 | 70.00% |
| Denied by reviewer | 20 | 20.00% |
| Expired without approval | 10 | 10.00% |
| Total | 100 | 100.00% |

Synthetic request sample-db-001: requested 2026-08-10T09:00:00Z, reviewed by the scenario security role at 09:02:00Z, approved, one exact retry succeeds, replay is refused. Approval alone does not prove execution: the scenario assumes a successful fixture response for the 70 approved retries. Initial decisions and review outcomes are separate dimensions and must not be summed as unique calls.

## Operational outcome

1150 permitted fixture executions = 1080 initial allows + 70 approved retries. 90 requests remain blocked = 60 initial denials + 20 reviewer denials + 10 expirations. Together they account for 1240 unique requests. No incident-prevention or financial-savings claim is inferred.

## Recommended rollout model

Use a hosted HTTP gateway for remote tools and Local Connector for command-mode tools; keep real production writes denied under the current built-in pack. Start with one fixture, inspect capability mapping, assign approval ownership and force traffic through the gateway. Customer-operated deployments own backup, retention and incident response. The scenario decision is a limited read-only rollout, not authorization for a real customer deployment.

Scenario review date: 2026-08-24. Scenario reviewers: platform owner and security owner; no real customer signature is represented.
