AI Implementation / Private and enterprise deployment

Private AI with measurable operating boundaries.

Deploy into customer-controlled cloud, private VPC, hybrid infrastructure or on-premises environments against frontier targets for availability, recovery, retention, inference performance and release quality.

Every target is bound to a named architecture, workload, policy matrix and acceptance test before production.

The frontier profile is fixed before production.

These are service specifications for a named deployment profile, not claims about historical YPAI results. The SOW names the target, architecture, prerequisites and measurement method. Production starts only after the acceptance evidence passes.

Area What is numbered
Network boundary

0 unapproved egress paths

Customer-specific cloud boundaries, private endpoints and disabled public network access are used where the selected services support them. Every required inbound and outbound path is named.

Inference retention

0-day prompt and response retention

Approved stateless inference paths retain no prompt or response content. Payload logging and caching are disabled, and incompatible stateful features are excluded unless separately approved.

Availability

99.99% monthly service availability

Measured from YPAI-managed ingress through application, retrieval, routing and inference to a protocol-valid response. Monthly availability is eligible minutes minus unavailable minutes, divided by eligible minutes. An unavailable minute begins after 60 seconds of continuous synthetic-transaction failure across the active path and every eligible fallback. Customer connectivity before ingress, customer-directed suspension and traffic beyond the accepted envelope are excluded; planned maintenance inside the boundary is not. Active-active application paths and a pre-approved policy-equivalent inference fallback are prerequisites; without one, traffic fails closed.

Recovery

RPO 0 critical writes · RTO <=60 sec serving path

Critical writes are acknowledged only after required synchronous replicas confirm them. Synchronously dual-written object classes use RPO 0; asynchronous object classes use a declared maximum of 15 minutes.

Private 70B inference

p95 TTFT <=1 s · p95 ITL <=50 ms

Llama-3.3-70b-Instruct on NVIDIA NIM 1.8.0, FP8 and four units of 2 x H200 141GB. Each unit runs 25 in-flight requests at 5,000 input and 500 output tokens. The aggregate target is 100 in-flight requests and at least 2,500 output tokens per second.

Fast 8B inference

p95 TTFT <=250 ms · p95 ITL <=10 ms

Llama-3.1-8b-Instruct on NVIDIA NIM 1.8.0, FP8 and 1 x H200 141GB, with 100 in-flight requests at 200 input and 200 output tokens. This is a separate profile and is not presented as the 70B target.

Release quality

100% required gates · 0 critical findings

Every required build, security, evaluation and deployment-verification gate passes. Releases permit 0 unresolved Critical vulnerabilities, 0 known actively exploited findings and 0 failures on critical control criteria.

Rollback

<=5 min to last known good version

Versioned application, prompt and model configuration return to 100% traffic within five minutes. State and schema migrations use a separately tested recovery path.

Audit and keys

100% agreed event coverage · customer-managed keys

YPAI-controlled persistent data uses customer-managed keys. Audit records cover every event in the agreed contract without prompt or response payload retention by default.

Environments

3 isolated environments

Development, staging and production are isolated. Promotion requires the release gates.

Actions

Fail closed outside approved authority

Maximum steps, tools or actions before human approval are set per agent and per irreversible action class.

Model control

Policy-equivalent targets only

Automatic model and provider failover is limited to pre-approved targets with equivalent retention, residency, security and accepted task quality.

Cost

Bound to accepted output

Budget limits, alert thresholds and maximum cost per task or period are set before production traffic starts.

The SOW binds the selected profile; production begins only after its prerequisites and acceptance tests pass.

Four patterns. Same envelope, different placement.

The pattern changes where the application, models and data sit. It does not remove the requirement to number egress, retention, availability, recovery and release quality.

Customer-controlled cloud

Customer account, subscription or project

The application and agreed services run in the customer's existing cloud control plane. Identity, networking, secrets and logs stay there. Managed models connect only through approved private or restricted paths that keep unapproved egress at 0.

Private or dedicated cloud

Dedicated tenant, account or network

Stronger isolation, restricted ingress and egress, and workload-specific access without moving the full operating burden on premises. Public endpoints, approved destinations and zones are still counted.

Hybrid deployment

Sensitive data stays inside; approved tasks may leave

Source documents, retrieval, state and actions remain in customer-controlled systems. Only approved context crosses to a managed or specialised model. An external inference hop uses zero prompt and response retention only when the selected endpoint, model, tools, cache and logging configuration all support the approved profile.

On-premises or disconnected

Customer infrastructure; disconnected only when proven

Licensable models and application services run on customer hardware. Disconnected operation is claimed only after inference, updates, telemetry and support access are shown to stay inside the boundary. Operating burden for serving, patching and incidents transfers with the pattern.

Each control carries a number.

Identity and permissions

Roles assigned before production

Users, services and agents connect through the approved identity provider. Least-privilege roles separate user, operator and administrator functions.

Network and egress

0 unapproved egress paths

Private connectivity, segmentation and allowlisted destinations. Every required inbound and outbound path is named. Undeclared paths stay closed.

Data handling

0-day payload retention on approved inference paths

Prompts and responses are not retained on the approved stateless inference path. Files, retrieval state, embeddings, traces and backups each receive their own processing location, storage location, key and deletion rule.

Models and routing

Policy-equivalent fallbacks only

Count, pin and restrict models by data class. A fallback is eligible only after it passes the same task evaluation and preserves the approved retention, geography, key and security controls.

Tools and actions

Approval limit assigned per agent

Agents receive only the systems, functions and credentials required for the approved task. Schema validation, scoped credentials, execution limits and fail-closed behaviour sit around irreversible actions.

Observability and operations

100% of events in the agreed audit contract

Capture identity, request, deployment, model, route, region, outcome, latency, usage and configuration events without prompt or response payloads by default. Monitor quality, failures and spend against the gates.

Every required release gate passes.

Production requires complete gate evidence, 0 unresolved Critical vulnerabilities, 0 known actively exploited findings, 0 failures on critical control criteria and a verified last-known-good rollback path.

100%

Required gates

Missing, interrupted or failing build, security, evaluation or deployment-verification evidence blocks release.

0

Critical findings

Critical vulnerabilities, known exploited findings and failures on critical privacy, authorisation, egress or financial-control criteria block production.

<=5 min

Rollback

Versioned application, prompt and model configuration return to the verified last-known-good version. The path is tested before production.

100% critical criteria

Evaluation

Non-critical task thresholds are set against representative held-out and adversarial cases, with no accepted regression on protected slices.

3

Environments

Development, staging and production stay isolated. Promotion follows the gates above.

Artefacts that carry the numbers

A private or enterprise deployment includes the records that make the envelope enforceable:

  • Data-flow map with public endpoints, approved egress and network zones
  • Model, endpoint, tool, cache and logging capability matrix for zero-retention paths
  • 99.99% active-active availability design and accepted inference-fallback policy
  • RPO 0 critical-write design and object-class recovery matrix
  • 60 second serving-path recovery and control-plane switch test
  • Named 70B or 8B p95, in-flight request and throughput acceptance profile
  • Customer-managed key and regional-processing matrix
  • Action-control matrix with approval limits
  • Evaluation suite with critical criteria, protected slices and acceptance report
  • 100% digest-bound signed provenance, SBOM, attestations and release evidence for production artifacts
  • Release and rollback procedure with the five minute target
  • Budget limits, alert thresholds and cost caps
  • Operating ownership and handover documentation

The SOW binds the selected profile, its prerequisites and its measurement method.

One numbered envelope across the systems YPAI builds

Private and enterprise deployment is the operating layer on assistants, document workflows and bounded agents. It is not a separate application type.

Document workflow automation

Sensitive documents move through intake, extraction, validation and system updates against assigned throughput and approval limits.

Explore document workflow automation

Bounded agents

Agents act through restricted tools, scoped credentials and a maximum number of steps before human approval.

Explore bounded agents

Customer and employee assistants

User-facing traffic is separated from internal systems, private data and administrative controls inside the same envelope.

Explore AI assistants

Zero-retention and regional boundaries are explicit profiles.

Processing, storage and failover can be confined to one agreed region or named geography. The actual model, endpoint, subprocessor and transfer path are recorded; a provider region label is not enough.

Eligible inference paths use 0-day prompt and response retention. Files, retrieval state, embeddings, traces, logs and backups are separate components with their own approved storage, key and deletion rules.

All YPAI-controlled persistent data can use customer-managed keys. Provider-side key coverage is confirmed for each selected service and feature.

Architecture can support legal and regulatory obligations. Architecture alone does not establish compliance.

AI Implementation / Private and enterprise deployment

Select the frontier profile. Then prove it.

Bring the use case, data classes, existing architecture, model requirements and failure boundaries.

YPAI binds the applicable frontier targets to an architecture, acceptance plan and controlled path through development, staging and production.

AI Implementation / Private and enterprise deployment

Scope the deployment

Tell us what the system must do, where it may run, which data it may access and what cannot leave your environment.

We use the information to assess the deployment scope and determine the right technical next step.

Private and enterprise deployment

Frequently asked questions

Can the complete system run inside our environment?

Yes, where the selected models, licences, infrastructure and dependencies support it. Inference, retrieval, application state, logging, updates and operational access are checked against the same envelope: 0 unapproved egress paths and an explicit retention rule per component. A system is not described as fully private while material processing uses an undeclared external service.

Can we use managed AI models without exposing our environment publicly?

Often, yes. Private connectivity, restricted public access, regional processing, zero-retention inference and customer-managed encryption keys are used where the selected service and feature support them. The path is accepted only when it keeps unapproved egress at 0 and retains no prompt or response content.

Can sensitive data remain inside our environment while an external model is used?

In some architectures. A hybrid design can keep source documents, retrieval, permissions and actions inside the customer environment and send only approved context to an external model. That inference hop can use 0-day prompt and response retention while stateful components keep their separately approved storage and deletion rules.

Which availability target do you commit to?

The frontier profile targets 99.99% monthly service availability. It requires active-active application architecture and a pre-approved inference fallback or redundant self-hosted path that preserves the same quality and control policy. Without those prerequisites, that profile is not offered.

Can YPAI deploy into our existing cloud or Kubernetes environment?

Yes, subject to the agreed scope, environment and access. The three isolated environments, identity, networking, secrets, logging and release gates are mapped onto the organisation's existing operating model.

Can we change models later?

Yes, when the approved model list, fallback routes and evaluation set are in place. Replacement models are tested against the same cases and must preserve the accepted retention, geography, security and task quality. A failing fallback is removed rather than silently weakening the profile.

Do you support on-premises or air-gapped environments?

Yes, where the model, licence, hardware and update process permit it. Disconnected or air-gapped is claimed only after inference, telemetry, package retrieval and support access are shown to stay inside the boundary, with unapproved egress at 0.

Who operates the system after launch?

YPAI can transfer the deployment to an internal owner or continue under an agreed monitoring and update scope. Ownership, access, the 99.99% profile prerequisites, RTO, RPO and escalation paths are defined before production. Delivery time in weeks is set per deployment class in the SOW.

Does private deployment make the system GDPR compliant?

No. Location, 0 unapproved egress paths and a TTL matrix can support applicable obligations. They do not establish compliance by themselves. Roles, purposes, transfers and subprocessors still belong in the DPA and organisational framework.