EVE Proof & verification

Verifying an EVE Red Team Report

Example status: illustrative, not runnable as written. The Python examples on this page are built on an embedded in-repo facade imported as core.eve_sdk (an EVE(...) object with govern_* methods). That module is not present in the current repository and is not published on any registry, so these snippets raise ModuleNotFoundError if run. They are kept because they show the intended governance flow. For code that runs today, use the published client: pip install eve-coreguard, then from eve_coreguard import CoreGuardClient and call client.evaluate(...), with offline evidence checking via client.verify_evidence(...) or the standalone verify_decision_record. See the CoreGuard integration guide and the Python SDK reference.

A red team report (eve.redteam.report.v1) is a signed record of running a packaged adversarial regression suite — single-step cases and multi-step chains — against your own configuration. For every attack that was blocked, the report records zero unauthorized side effects, measured by an independent recording target. The report lets a reviewer confirm the results are authentic, complete, and unaltered — without trusting the EVE server.

Readiness: the customer-runnable red-team harness + signed report is PILOT_READY (customer-pilot ready at supported boundaries). The suite is representative, not exhaustive, and standards mappings are descriptive, not certification.

What it proves

  • Integrity + authenticity — the results (per-case outcomes, totals, side-effect counts) were signed by the EVE signing key (ECDSA P-384 (AWS KMS) in the production configuration; HMAC-SHA256 fallback is symmetric and not independently verifiable) over jcs-1 canonical bytes.
  • Completeness (no selective omission) — the results_digest binds the full result set, so quietly dropping a failed case is detected. Strict verifiers detect tampering and selective omission.
  • Zero-side-effect proof — for every blocked_before_execution case, an independent RecordingTarget counted external calls / file writes / budget / session / approvals and the count stayed 0; positive controls prove that permitted actions still execute. In the built-in suite, 16 cases plus 2 multi-step chains passed with 0 unauthorized side effects.

What it does NOT prove

  • It does not certify anything. The suite is representative, not exhaustive; customers add fixtures. Passing the suite is not a certification, a penetration-test pass, or a compliance guarantee.
  • Standards mappings (MITRE ATLAS / OWASP-LLM / NIST AI RMF) are descriptive links from case ids to framework controls — they do not imply certification against those frameworks.
  • A green report describes the cases that were run against your configuration; it does not prove the absence of an attack that is not in the suite. Deterministic detectors are pattern-based and novel obfuscations may evade them until a detector is added.

Required keys

eve.redteam.report.v1 binds at least:

Field Meaning
schema version eve.redteam.report.v1 identifier
suite / scanner / policy / app versions Versioned components under test
run_id Run identifier
results_digest Digest of the full result set (omission binding)
totals Passed / failed / skipped / unsupported counts
side-effect counts Unauthorized side effects (expected 0 for blocked cases)
outcome taxonomy blocked / failed / unknown / executed / executed-with-findings
kid, canon, signature Signer key, jcs-1, signature

Verification fails closed if a required key is absent.

Canonicalization (jcs-1)

The report is signed over jcs-1, a constrained RFC 8785 profile with proven byte-for-byte parity across Python, TypeScript, and the browser for the supported value domain. jcs-1 fails closed on non-integer floats / NaN / Infinity. A verifier reading a report whose canon is not jcs-1 reports unsupported canonicalization.

Verification profiles: strict vs standard

  • Standard — verifies canonicalization, content hash, signature, and schema.
  • Strict — additionally requires an independently verifiable ECDSA P-384 signature, rejects HMAC-fallback evidence, and enforces the results_digest binding so a partial (selectively omitted) result set is rejected. Use strict for third-party review. The Python strict verifier is verify_redteam_report; the Node strict verifier is verify_redteam_report.mjs.

Linked evidence

Verify each linked artifact independently.

Selective-disclosure / omission checks

Selective omission is a first-class check here: the results_digest binds the full set of case results, so removing a case (for example, a failed one) is detected by the strict verifier. Removing or altering any signature-covered field invalidates the content hash and the signature. Verify the full report before reading totals; do not re-canonicalize a reduced object.

Failure categories

Category Meaning
invalid signature Signature does not match the canonical content under the given key
unknown key kid / public key not recognized or not supplied for a strict check
unsupported schema Report schema version is not understood by this verifier
unsupported canonicalization canon is not jcs-1
digest mismatch results_digest or content hash does not match the recompute
selective omission A case result was removed from the bound result set
stale Report is older than your freshness window

Freshness

A red team report is pinned to specific suite / scanner / policy / app versions. It does not expire on its own; enforce a freshness policy so you do not accept a report produced under an older configuration than the one you are running, and treat over-age reports as stale. A new suite, scanner, or policy version warrants a fresh run.

Replay behavior

The harness is safe by default: local recording targets, no destructive tests, dry-run, secret redaction, bounded concurrency/runtime, and remote testing requires an explicit allowlist. A report is a record, not an authorization — re-presenting it does not re-run the suite and causes no side effect. Re-running the suite over the same configuration is reproducible via the run_id and versioned components.

Verify it yourself

Python

# The unified verifier ships with the EVE service; the first argument is the
# evidence KIND ("redteam"). Synthetic fixture — no production credentials.
import json
from core.eve_sdk import verify_evidence

with open("redteam_report.json", "r", encoding="utf-8") as fh:
    report = json.load(fh)

# strict=True enforces the results_digest binding (rejects a selectively-omitted set).
result = verify_evidence("redteam", report, strict=True)
print(result["valid"], result["reason"])  # True, "verified"

# Pip-wheel path for content-hash + ECDSA P-384 signature only:
#   from eve_coreguard import verify_decision_record
#   r = verify_decision_record(report, public_key_pem=public_key_pem)
#   assert r.valid and r.independently_verifiable

Node (TypeScript / ESM, Node >= 18)

import { verifyDecisionRecord } from "eve-ai-governance";
import { readFileSync } from "node:fs";

const publicKeyPem = readFileSync("public_key.pem", "utf8");
const report = JSON.parse(readFileSync("redteam_report.json", "utf8"));

const result = await verifyDecisionRecord(report, { publicKeyPem });
console.log(result.valid ? "VERIFIED" : "FAILED");

CLI

# The `eve` CLI ships with the EVE service/repo; it also runs the harness
# (`python -m core.eve_sdk.cli redteam`). Verify a report with:
python -m core.eve_sdk.cli verify redteam redteam_report.json --strict

Readiness

  • The customer-runnable red-team harness + signed report is PILOT_READY (customer-pilot ready at supported boundaries).
  • Offline verification (Python / Node / browser for the report) is SUPPORTED; jcs-1 is SUPPORTED (constrained RFC 8785 profile).
  • The harness runs in embedded (service) / sidecar mode; the Python eve-coreguard client (0.2.10) reaches the service via hosted mode. TypeScript verifies the report offline and never signs.
  • Independent verification requires the ECDSA P-384 public key; the HMAC fallback is symmetric and not independently verifiable.
  • The clients are published on public registries: pip install eve-coreguard (PyPI) and npm install eve-ai-governance (npm).

Limitations

  • The suite is representative, not exhaustive; customers add their own fixtures. Standards mappings are descriptive, not certification.
  • Zero-unauthorized-side-effects is proven for the cases in the suite via an independent recording target; it is not a claim about attacks outside the suite.
  • Deterministic detectors are pattern-based; novel obfuscations may evade them until a detector is added (findings are versioned so coverage changes are auditable).
  • In the pilot, harness state is in-process (bounded, tenant-partitioned); durability is a deployment concern.
  • HMAC-fallback signatures are not independently verifiable; use the ECDSA P-384 configuration and the strict profile for external review.
Part of the EVE AI Core control plane Deterministic AI Governance Control Plane → Policy decisions that return the same result for the same input every time, before execution.