EVE Core › Docs › AI Policy Enforcement Architecture

Concepts

AI Policy Enforcement Architecture: A Technical Guide

Deploying an AI governance enforcement layer in a production LLM pipeline is an architectural problem, not just a policy problem. The system must intercept the right events, evaluate them against the right policies at the right point in the request lifecycle, return a verdict before the output reaches the end user, and produce an audit record that satisfies compliance requirements — all in under one millisecond. Get any one of these properties wrong and you either create a compliance gap or degrade the user experience to the point where the enforcement layer is bypassed. This guide covers the core architectural patterns borrowed from XACML and attribute-based access control, how to adapt them for LLM pipelines, the latency budget you need to hit, the three major integration patterns (in-process SDK, sidecar proxy, and API gateway), policy versioning and rollback procedures, and the audit log schema requirements for regulated industries. The EVE CoreGuard architecture is used throughout as a concrete reference implementation.

The Policy Enforcement Point Pattern from XACML

The conceptual foundation for AI policy enforcement architecture is the Policy Enforcement Point (PEP) pattern, originally defined in the eXtensible Access Control Markup Language (XACML) specification and widely adopted in attribute-based access control (ABAC) systems. The pattern separates enforcement into three distinct roles:

  • Policy Enforcement Point (PEP): The component on the critical path that intercepts requests and enforces verdicts. It has no policy logic of its own — it delegates to the PDP and enforces whatever verdict comes back.
  • Policy Decision Point (PDP): The component that evaluates a request against the applicable policy and returns a verdict. This is where the governance rules live.
  • Policy Administration Point (PAP): The management interface for authoring, versioning, and publishing policies. In AI governance, this is the policy editor and deployment pipeline.

In a monolithic implementation like EVE CoreGuard, the PEP and PDP are co-located for performance, but the logical separation remains important: the PEP code handles request interception and verdict enforcement, while the PDP code handles rule evaluation. This separation enables testing each in isolation and enables future distribution (e.g., a remote PDP for centralized policy management in multi-application environments).

Applying the PEP Pattern to LLM Pipelines

In a standard LLM application, the request lifecycle looks like this: the application constructs a prompt, calls the LLM API, receives the model response, and delivers it to the user or downstream system. Inserting a PEP produces this modified lifecycle:

Application constructs prompt
→ Normal application logic
LLM API call
→ Inference: 200ms–2000ms
PEP: EVE CoreGuard evaluation
→ Enforcement: <1ms | ALLOW / BLOCK / MODIFY
Output delivery
→ Only compliant output reaches the user
Audit log write
→ Async — does not block delivery

The critical design decision is placement: the PEP must be inserted between the LLM response receipt and the output delivery, not before the LLM call. Pre-inference enforcement (evaluating the prompt before sending it) is valuable for additional security properties but does not substitute for post-inference enforcement, because the model may generate non-compliant content regardless of prompt design. The authoritative enforcement gate is always on the output side.

Architecture principle: The PEP must be on the output path, not the input path. Prompt filtering is a useful secondary control. It is not a substitute for output enforcement, because you cannot predict what a stochastic model will generate from a compliant prompt.

Latency Budget: Keeping Enforcement Under 1ms

LLM inference typically takes 200ms to 2,000ms depending on model size, provider, and response length. The enforcement layer's latency budget is the remaining headroom before users perceive a slowdown — in practice, anything under 5ms is invisible, and anything under 1ms is architecturally negligible.

Achieving sub-millisecond enforcement requires four engineering constraints:

  1. No LLM inference in the critical path. A secondary model call to evaluate the output would add 200ms minimum — unacceptable. All rule evaluation must be deterministic computation: pattern matching, threshold comparisons, structured field extraction.
  2. Pre-compiled rule sets. Policy packs are parsed, validated, and compiled into an in-memory evaluation structure at startup. No file I/O or JSON parsing occurs during request evaluation. The compiled rule set is a lookup-optimized data structure that supports O(1) or O(log n) access to rules by domain and severity.
  3. Early exit on CRITICAL violations. Rule evaluation stops as soon as a CRITICAL-severity rule is triggered. A CRITICAL violation always produces BLOCK regardless of other rules, so there is no need to evaluate remaining rules once one CRITICAL fires. In most violation scenarios, this means evaluation terminates in the first few microseconds.
  4. In-process or co-located deployment. Network round-trips add 1–50ms of latency depending on topology. The enforcement engine must run either in the same process as the application (SDK pattern) or in a co-located sidecar on the same host (sidecar pattern). Cross-region or cross-datacenter enforcement calls violate the latency budget for synchronous enforcement.

Observed latency: EVE CoreGuard's evaluation engine completes in under 1ms for standard policy packs (lending_v1, healthcare_v1, legal_practice_v1) on commodity compute. Certificate signing and the audit-chain append are measured separately. Rule-evaluation overhead: under 1ms in all measured configurations. Signing and hash-chaining are measured separately.

Integration Patterns

Three integration patterns cover the full range of enterprise deployment scenarios. The right choice depends on the ownership boundary between the team that owns the AI application and the team that owns the enforcement policy.

01

Client SDK

The PEP is a function call in the application code: the SDK sends the proposed action to the hosted CoreGuard decision service and acts on the signed verdict it returns. One HTTPS round trip; no policy engine runs in your process.

Best when: application team owns policy
02

Sidecar Proxy

The enforcement engine runs as a separate container on the same Kubernetes pod or VM. The application sends all LLM requests through the proxy. The proxy enforces before passing compliant responses through.

Best when: centralized enforcement team
03

API Gateway Integration

Your API gateway calls the CoreGuard HTTP API (POST /v1/decisions/evaluate) from its own policy or callout step. There is no shipped Kong, Apigee or AWS API Gateway plugin; for Envoy-based gateways EVE provides an ext_authz backend.

Best when: many AI services, one policy team

Pattern 1: Client SDK

The SDK pattern is the simplest to implement. The application installs the EVE CoreGuard Python SDK (pip install eve-coreguard), creates a CoreGuardClient with its API key at startup, and calls evaluate() before delivering any LLM output. The SDK is a thin client: evaluation runs in the hosted CoreGuard service, which chooses the governing policy pack from your organization and proposed_action["type"] and returns the signed decision. No policy pack or signing key is loaded into your process.

SDK Integration — Python
import os
from eve_coreguard import CoreGuardClient, CoreGuardError

# Create once at startup and reuse
guard = CoreGuardClient(api_key=os.environ["COREGUARD_API_KEY"])

# In the request handler — after the LLM response, before delivery
llm_output = await llm_client.generate(prompt)

try:
    # chat_response is bound server-side to the chat_safety_v1 content-safety pack
    result = guard.evaluate(
        tenant_id=org_id,
        user={"id": user_id, "role": "loan_officer"},
        proposed_action={"type": "chat_response"},
        context={"response": llm_output, "session_id": session_id},
        # explicit: eve-coreguard 0.2.10 sends "lending_v1" when this is omitted
        policy_set="chat_safety_v1",
    )
except CoreGuardError:
    return policy_safe_fallback_response([])  # no decision obtained: fail closed

if result.blocked:
    return policy_safe_fallback_response(result.policy_violations)
elif result.verdict == "MODIFIED":
    return apply_modifications(llm_output, result.modifications)
else:
    return llm_output  # ALLOWED — unmodified

Pattern 2: Sidecar Proxy

The sidecar pattern is preferred when the AI application team does not own the enforcement policy, or when multiple AI services share a policy configuration. The EVE sidecar (private preview) runs in the same pod as an egress governance forward proxy and Envoy ext_authz backend (HTTP on port 9001, gRPC on 9002). It evaluates each request with CoreGuard and signs a decision certificate for it, failing closed when it cannot decide.

Because routing is configured in the proxy or service mesh rather than in application code, application teams do not call the sidecar directly or adopt the SDK. This is the deployment model for organizations that want centralized policy management at the network layer. Contact us for the sidecar build and Helm chart.

Pattern 3: API Gateway Integration

For organizations with many AI services that all route through a central API gateway, enforcement can sit at the gateway instead of in each service. This is an integration pattern built on the HTTP API, not a shipped plugin: EVE does not ship a Kong, Apigee or AWS API Gateway plugin. The gateway’s own extension point (a custom plugin, policy callout or request/response hook that you configure) sends the proposed action to POST /v1/decisions/evaluate and enforces the returned ALLOWED / BLOCKED / MODIFIED verdict, failing closed if no decision comes back.

For Envoy and Envoy-based gateways there is a concrete option: EVE’s ext_authz authorization backend (HTTP on port 9001, gRPC on 9002), with a reference Envoy configuration. It authorizes each request before Envoy proxies it upstream; a denied request never reaches the backend. The trade-off of any gateway pattern is an extra network round trip to the decision service on every governed call.

Policy Versioning and Rollback

Policy versioning is one of the most underspecified aspects of AI governance architecture. Signed decision records bind the policy version that evaluated them on by default — best-effort, so a record whose frozen version was unavailable is emitted empty and unpinned. The fail-closed decision certificate carries no policy_version field. If policy versions are not managed correctly, the audit trail is incomplete and historical decisions cannot be reproduced.

The versioning model for EVE CoreGuard policy packs follows semantic versioning with strict immutability rules:

  • MAJOR version increment (e.g., v1 → v2): A rule's blocking behavior changes in a way that affects existing decisions. Existing decisions evaluated under v1 are not invalidated — they remain valid under v1. New decisions use v2.
  • MINOR version increment (e.g., v1.3 → v1.4): New rules are added without modifying existing rule behavior. Decisions under v1.3 remain valid; new decisions benefit from additional coverage.
  • PATCH version increment (e.g., v1.4.1 → v1.4.2): Documentation, metadata, or regulatory citation updates. No change to rule logic. Certificates under either version are equivalent.

Immutability rule: Once a policy version is deployed to production, it must not be modified. Bug fixes in rule logic require a new version increment. This ensures that a regulator asking "what policy was active on date X?" always gets a precise, reproducible answer.

Rollback Procedure

Policy rollback means activating an earlier version for new decisions — not retroactively changing past decisions. The rollback procedure is:

  1. Identify the last known-good policy version.
  2. Update the enforcement configuration to reference that version.
  3. Restart or hot-reload the enforcement engine.
  4. Record the rollback event in the audit log with the reason and the version transition.
  5. Verify that new decision certificates reference the prior version.

Because policy versions are immutable, rollback does not require any modification to policy files — the prior version already exists in the policy repository. Rollback is a configuration change, not a deployment.

Audit Log Schema Requirements

The audit log is the forensic record of every enforcement decision. Its schema must satisfy the documentation requirements of the applicable regulatory frameworks — and it must be designed for query efficiency, because compliance teams will need to retrieve decisions by date range, user, policy version, verdict, and violated rule.

EVE CoreGuard Audit Log Entry — Complete Schema
{
  "decision_id": "cg_7f3a9c2e4b1d",       // Globally unique
  "request_id": "req_a1b2c3d4",          // Correlates with app logs
  "timestamp": "2026-05-05T14:23:11.042Z",  // ISO-8601 millisecond precision

  "subject": {
    "user_id": "u_123",
    "role": "loan_officer",
    "session_id": "sess_abc"
  },

  "policy": {
    "set": "fair_lending_v1",
    "version": "1.0.0",
    "version_hash": "sha256:3c9f2a..."  // Hash of policy pack artifact
  },

  "decision": {
    "status": "BLOCKED",
    "risk_level": "HIGH",
    "risk_score": 0.87,
    "evaluation_ms": 0.6
  },

  "violations": [
    {
      "rule_id": "FL-004",
      "severity": "MEDIUM",
      "citation": "ECOA 15 U.S.C. § 1691(a)",
      "triggered_by": "zip_code_proxy_variable"
    }
  ],

  "certificate": {
    "hmac": "sha256:a3f8d1e2b7c4...",
    "algorithm": "HMAC-SHA256",
    "key_id": "org_signing_key_v3"
  }
}

Recommended Indices for Compliance Queries

For regulatory response-time requirements, the audit log store should maintain composite indices on: (subject.user_id, timestamp) for per-user history; (policy.version, timestamp) for per-version audit trails; (decision.status, timestamp) for violation frequency reporting; and (violations[].rule_id, timestamp) for per-rule enforcement history. These four indices cover the vast majority of regulatory examination queries without requiring full-table scans.

Frequently Asked Questions

What is a Policy Enforcement Point (PEP) in AI systems?

A Policy Enforcement Point (PEP) is an architectural component borrowed from XACML and attribute-based access control that intercepts every request requiring a policy decision, forwards it to a Policy Decision Point (PDP), and enforces the resulting verdict. In AI systems, the PEP sits on the critical path between the LLM inference call and the output delivery channel. Every proposed action — a model output, a tool invocation, a data access request — passes through the PEP before reaching the end user or downstream system.

How do you keep AI policy enforcement under 1ms?

Sub-millisecond rule evaluation requires four engineering choices: (1) deterministic rule evaluation — no LLM calls in the critical path; (2) pre-compiled rule sets parsed at startup, not at evaluation time; (3) early exit on CRITICAL violations to avoid evaluating unnecessary rules; and (4) in-process or co-located deployment to avoid network round-trips. EVE CoreGuard's evaluation engine completes in under 1ms on commodity compute using these techniques.

What should an AI governance audit log record?

A complete AI governance audit log entry should record: a unique decision ID, the request ID for correlation with application logs, the user and session context, the policy set name and version, the complete list of evaluated rules, the triggered violations with severity and regulatory citations, the final verdict (ALLOW/BLOCK/MODIFY), the composite risk score, the evaluation latency in milliseconds, the HMAC signature, and an ISO-8601 timestamp with millisecond precision. If MODIFY was applied, both the original and modified outputs should be preserved.

How do you handle policy versioning in an AI enforcement system?

Policy versions must be immutable — a version, once published, never changes. New behavior requires a new version increment. A signed decision record binds the exact policy version that evaluated it — on by default, best-effort, enabling audit replay. Policy versions use semantic versioning where MAJOR changes modify rule behavior, MINOR changes add new rules without modifying existing ones, and PATCH changes fix documentation without changing rule logic.

Try the EVE CoreGuard Enforcement API

Explore all three integration patterns — SDK, sidecar proxy, and API gateway — against live lending, healthcare, and legal policy packs in the interactive demo.

Part of the EVE AI Core control plane Deterministic AI Governance Control Plane → Policy decisions that return the same result for the same input every time, before execution.