EVE Core/Docs/Limitations

Getting started

Limitations

The credibility of a proof system depends on stating plainly what it does not do. This page is authoritative; where it and marketing copy disagree, this page wins.

Preview, not shipping

  • External witnessing & transparency-log inclusion. Described on the Trust Services page as a preview. There is no implementation of independent witnessing or transparency-log submission in this codebase today. A certificate is currently self-signed by EVE only.
  • Cryptographic custody manifests / multi-party co-signing. Not implemented. All signing is single-party (ECDSA P-384 by default).

Verification caveats

  • HMAC mode is not independently verifiable. A symmetric signature requires the shared secret. Only ECDSA P-384 gives public-key, third-party verification.
  • Key attribution is out-of-band. Verifying a signature proves who signed (a key), not that the key is EVE's. Pin EVE's fingerprint through a trusted channel.
  • Replay id is not signature-bound in v3. Anti-replay and deterministic replay exist (core/coreguard/replay_guard.py, CLI/replay_verify.py), but the decision-chain coordinates are unsigned metadata rather than fields inside the signed record.

Claims we will not make without evidence

  • Latency ("sub-millisecond") is scope- and percentile-dependent. Harnesses ship in benchmarks/coreguard/; recorded output is data/bench/decision_latency.json (schema eve.bench.decision_latency.v1). Three distinct scopes, all on a single warm host (Windows Server 2019, AMD64, CPython 3.12.8):
    • Policy evaluation only (compiled governance VM; no signing, no evidence write, no network): p50 0.006 ms, p99 0.015 ms, n=20,000.
    • Verdict derivation including binding, over five n=20,000 runs: a single PASS verdict measured p50 0.09–0.12 ms, p95 0.17–0.33 ms. A 10-violation verdict is the worst case and is not reliably sub-millisecond: p50 0.49–0.57 ms, p95 0.73–1.19 ms (exceeding 1 ms in two of five runs), p99 1.3–4.7 ms. The spread is load, not measurement error — the same metric measured 0.73 ms and 1.19 ms on the same host under different concurrent load. Sub-millisecond describes a typical single-verdict rule evaluation, not a many-violation worst case and not the tail.
    • Full governed decision (adds hash-chain anchoring and evidence persistence): p50 17.6 ms, p99 20.7 ms, n=300.
    The served HTTP path is not measured. Figures will differ under production CPU and concurrent load. Treat any sub-millisecond number quoted without both a scope and a percentile as unverified.
  • The benchmark harness is not yet independently runnable. governance_benchmark_public.py is named "public" but requires a signing key (EVE_VERDICT_SIGNING_KEY or JWT_SECRET_KEY) and imports core/coreguard/, so a third party cannot today reproduce these figures from a released artifact. Until that ships, the numbers above are ours, not independently verified.
  • No customer counts, certifications, or independent-audit claims are made here.

Scope of the quickstart evidence

The certificates in examples/quickstart/evidence/ are signed with a local development key. They are genuine and independently verifiable against that dev key, but they are not production evidence and are not externally witnessed. See examples/quickstart/evidence/PROVENANCE.md.

Part of the EVE AI Core control plane Deterministic AI Governance Control Plane → Policy decisions that return the same result for the same input every time, before execution.